feat(p22): stage 0's unknown-code refusal NAMES the codes the project buys
P21 was the first ANCHORED stress round and it bought something real -- must_refuse 5/5, all five caught on stage0-baseline. It cost something measured just as clearly: 0 of 20 approaches validated, and 26 of 26 rejections (20 approach rows plus 6 own-proposals -- a LARGER population than the 20) read `unknown cost code '<invention>': not in project P's cost baseline (5 known codes)`. The model invented signalregulering_konstruksjon, VENTIL_IMP, RIGG01, baerelag_asfalt and 22 more, and it could not have done otherwise: the price schedule reaches the VALIDATOR and never the proposer, and the refusal stated the COUNT of known codes, not one name. Step 5 feeds that sentence verbatim into the next attempt -- and "you guessed wrong, there are five right answers" carries nothing to correct towards. The contrast already lived in the same stage: the MAGNITUDE half NAMES the baseline value, and that is the half that let the loop converge in session 94. This gives that property to the other half, in one place, and Step 5 carries it forward for free. The window is a FIXED COUNT of whole codes (20), never a share, and it counts codes rather than characters because a character cut can sever a code mid-name and hand the proposer an identifier that exists nowhere. The cut is announced; a schedule that fits is not marked truncated; the order is the schedule's own. Measured with denominators: every cost baseline in this repo or its measured corpora is at most six codes, and the largest real delivered price schedule measured is K2's prissammenstilling at 14 priced rows. Nothing measured reaches the window; it exists for the R761-style mengdebeskrivelse. Load-bearing MEASURED (tests/test_named_known_codes_loadbearing.py, 10 arms), eight mutations all red against the WHOLE suite + green control 1873/5 (from 1863/5, superset, 0 removed) and golden demo-transcript.stdout BYTE-UNCHANGED (shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f): A1 revert to the bare count (7 red) - A2 a share instead of a fixed window (6) - A3 sorted (1) - A4 silent cut (2) - A5 character slice (3) - A6 no bound at all (3) - A7 break the rejection_stage marker (2, one an OLDER independent witness) - A8 grow the magnitude half with a code list (2, one an OLDER independent witness). Honesty limits, stated: no paid run yet confirms this changes the outcome live (that is DEL D); the truncation branch is exercised only synthetically because nothing measured reaches the window; and the schedule still does not reach the prompt, so the first attempt guesses as before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
f751a66bbe
commit
0a81de2d76
4 changed files with 302 additions and 3 deletions
40
CLAUDE.md
40
CLAUDE.md
|
|
@ -2971,6 +2971,46 @@ Python ≥3.10. MAF (`agent-framework-core` 1.16.0, `-orchestrations` 1.1.1 —
|
|||
lesninger teller mot gulvet — regelen måler at kjøringen SÅ seg om, ikke at den forsto; ingen
|
||||
LEVENDE modell har møtt noen av de to nektene (structured-output-grensens klasse); og C2 er
|
||||
hjelpetekst, ikke en gate — den kan ikke gjøre en gjettet sti riktig, bare billigere å rette.
|
||||
- **Stadium 0s ukjent-kode-nekt NAVNGIR kodene prosjektet faktisk kjoeper, i et FAST vindu
|
||||
(P22 DEL A, 15.09):** P21 var den foerste FORANKREDE runden, og den kostet noe maalt like tydelig
|
||||
som den kjoepte: **0 av 20** tilnaerminger validerte (runde 4: 4 av 20), og **26 av 26**
|
||||
avvisninger — 20 approach-rader PLUSS 6 egne forslag, altsaa en STOERRE populasjon enn de 20 —
|
||||
leste `unknown cost code '<paafunn>': not in project P's cost baseline (5 known codes)`. Modellen
|
||||
fant paa `signalregulering_konstruksjon`, `VENTIL_IMP`, `RIGG01`, `baerelag_asfalt` og 22 til, og
|
||||
den KUNNE ikke gjort annet: prisskjemaet naar VALIDATOREN og aldri proposeren, og nekten oppga
|
||||
ANTALLET kjente koder, ikke ett eneste navn. Steg 5 mater setningen ORDRETT inn i neste forsoeks
|
||||
prompt, saa «du gjettet feil, det finnes fem riktige» baerer ingenting aa korrigere mot.
|
||||
**Kontrasten bodde allerede i samme stadium:** MAGNITUDE-halvdelen NAVNGIR baseline-verdien, og
|
||||
DET er halvdelen som lot loekka konvergere i oekt 94. `_known_codes_clause` er den ene
|
||||
renderingen, og magnitude-halvdelen faar den IKKE — der er setningen alt korrigerbar, og en
|
||||
kodeliste paa begge ville gjort de to halvdelene uskillbare for en leser (M A8 → 2 roede, hvorav
|
||||
ett ELDRE uavhengig vitne, `test_stage0_all_violations::test_a_single_violation_is_byte_identical`).
|
||||
**Vinduet er et FAST ANTALL, aldri en andel** (`_CATALOGUE_EXCERPT_CHARS`-regelen; M A2 → 6 roede,
|
||||
M A6 → 3), **og det teller KODER, ikke tegn**, av en maalt grunn: et tegn-kutt kan kappe en kode
|
||||
midt i navnet og gi proposeren en identifikator som finnes ingen steder — `_index_excerpt`-regelen
|
||||
invertert (M A5 → 3 roede paa hel-kode-armen). **Kuttet ANNONSERES** (`first N:`) mens nevneren
|
||||
staar i begge grener, og et skjema som PASSER merkes ikke avkortet (`index_truncated`s regel;
|
||||
M A4 → 2 roede). **Rekkefoelgen er skjemaets egen** — en sortering ville oppfunnet en rangering
|
||||
prosjektet aldri uttalte (M A3 → 1 roed, den armen ALENE). Vinduet er 20, valgt ved MAALING med
|
||||
nevner: hver kostbaseline i repoet eller dets maalte korpora er hoeyst SEKS koder (kontekstsettene
|
||||
5/5/5/5/6, de to `shared/examples` 1 hver, MAJOR-4s avledning av det syntetiske K2-prisskjemaet 3),
|
||||
og det stoerste EKTE leverte prisskjemaet maalt er K2s `prissammenstilling-sheet-1.md` med 14
|
||||
prisede rader av 118 linjer — ingenting maalt naar vinduet; det finnes for den umaalte
|
||||
R761-formede mengdebeskrivelsen, der korpuset erklaerer 2 727 `prosessnr`.
|
||||
**`rejection_stage` er koblingen P21s «a4 5/5 paa stage0-baseline» hviler paa** og noekler paa
|
||||
delstrengen `cost baseline (` — hadde den nye klausulen flyttet den, ville hver stadium-0-nekt
|
||||
blitt omdoept til `other` i stillhet (M A7 → 2 roede, hvorav ett ELDRE uavhengig vitne).
|
||||
**IKKE GJORT, og det er en beslutning:** skjemaet rendres ALDRI inn i proposer-prompten — en ny
|
||||
prompt-flate med egen kostnad per kjoering som dessuten ikke fjerner behovet for at nekten er
|
||||
korrigerbar. Load-bearing MAALT (`tests/test_named_known_codes_loadbearing.py`, 10 armer), **aatte
|
||||
mutasjoner alle roede mot HELE suiten** + groenn kontroll **1873/5** (fra 1863/5, supersett,
|
||||
0 fjernet) og golden `demo-transcript.stdout` BYTE-UENDRET (`shasum -a 1` av INNHOLDET =
|
||||
`ea8c534773acdbe41ae68f2c55724d69aaf8be4f`): A1 tilbake til det bare antallet (7) · A2 andel (6) ·
|
||||
A3 sortert (1) · A4 stille kutt (2) · A5 tegn-slice (3) · A6 ingen binding (3) · A7 bryt
|
||||
stage-markoeren (2) · A8 la magnitude-halvdelen vokse en kodeliste (2). **AErlighets-grenser,
|
||||
uttalt:** ingen betalt kjoering bekrefter ennaa at nekten endrer utfallet levende (DEL D er
|
||||
maalingen); vinduet er aldri naadd av noe maalt, saa avkortnings-grenen er oevd kun syntetisk; og
|
||||
skjemaet naar fortsatt ikke prompten, saa foerste forsoek gjetter som foer.
|
||||
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
||||
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
||||
|
||||
|
|
|
|||
|
|
@ -757,8 +757,12 @@ when the seam is detached, so the loop cannot silently degrade into theater.
|
|||
belongs to the project, and this is where you hand it over: FILE is a `cost-baseline.json` — the
|
||||
same `{"project_id": …, "items": {"<code>": {"quantity": …, "unit_cost": …}}}` shape a bundle may
|
||||
ship — and it is used INSTEAD of one inside the base. With it, a proposal naming a cost line the
|
||||
project does not buy is refused as a fabricated line, and one naming a real line with invented
|
||||
magnitudes is refused with the real ones named, so the next attempt can correct.
|
||||
project does not buy is refused as a fabricated line — and the refusal NAMES the codes the
|
||||
project does have, so the next attempt has something to correct towards. A real line with
|
||||
invented magnitudes is likewise refused with the real figures named. (Measured on a live tender
|
||||
round before the codes were named: 26 rejections, every one of them an invented code, and not
|
||||
one run recovered — a refusal that says only "there are five right answers" carries nothing to
|
||||
aim at. A long schedule is cut to a fixed number of codes and says so.)
|
||||
|
||||
In `--across-bundle` mode the same schedule anchors every base: one project, one price schedule.
|
||||
Mutually exclusive with `--derive-cost-baseline` (two sources for one baseline), and it satisfies
|
||||
|
|
|
|||
|
|
@ -154,6 +154,50 @@ def baseline_from_project(project: Project) -> CostBaseline:
|
|||
)
|
||||
|
||||
|
||||
#: P22 DEL A - how many of the project's OWN cost codes the stage-0 refusal names.
|
||||
#:
|
||||
#: A FIXED WINDOW, never a share of the schedule: a share scales with the corpus again, only with
|
||||
#: a smaller constant, which is the exact regression the catalogue excerpt's
|
||||
#: ``_CATALOGUE_EXCERPT_CHARS`` exists for. It counts CODES rather than characters for a measured
|
||||
#: reason - a character cut can sever a code mid-name and hand the proposer an identifier that
|
||||
#: exists nowhere, the ``_index_excerpt`` rule inverted (a path that never existed is worse than
|
||||
#: no path). A count window can only ever emit whole codes.
|
||||
#:
|
||||
#: MEASURED with denominators (okt 126): every cost baseline anywhere in this repo or its measured
|
||||
#: corpora is at most SIX codes - the five context sets carry 5/5/5/5/6, the two shipped
|
||||
#: ``shared/examples`` baselines carry 1 each, and MAJOR-4's derivation of the synthetic K2 price
|
||||
#: schedule yields 3 - while the largest REAL delivered price schedule measured is K2's
|
||||
#: ``prissammenstilling-sheet-1.md``, 14 priced rows of 118 lines. Nothing measured reaches this
|
||||
#: window. It exists for the unmeasured R761-style mengdebeskrivelse, where a Norwegian road
|
||||
#: contract is priced BY prosesskode and the corpus declares 2 727 of them.
|
||||
_KNOWN_CODE_WINDOW: Final = 20
|
||||
|
||||
|
||||
def _known_codes_clause(baseline: CostBaseline) -> str:
|
||||
"""The parenthetical of stage 0's unknown-code refusal: how many codes the project buys, and
|
||||
which ones (P22 DEL A).
|
||||
|
||||
MEASURED (P21 round 5): all 26 rejections across six paid runs named an invented code and
|
||||
answered with a COUNT - ``(5 known codes)`` - so the sentence Step 5 feeds verbatim into the
|
||||
next attempt carried nothing to correct towards, and 0 of 20 approaches validated. The
|
||||
magnitude half of this same stage NAMES the baseline value, and that is the half that let the
|
||||
loop converge in okt 94. This is that half's property, given to the other one.
|
||||
|
||||
Bound by ``_KNOWN_CODE_WINDOW`` - a fixed count of WHOLE codes, see the constant. The cut is
|
||||
ANNOUNCED (``first N:``) rather than left for the reader to subtract from the denominator,
|
||||
which stays in either branch; a schedule that FITS is not marked truncated and gets its whole
|
||||
list. Omission, never a lie in either direction (``index_truncated``'s rule).
|
||||
|
||||
The order is the SCHEDULE's own. A sort would invent a ranking the project never stated, and
|
||||
the window would then be an alphabetical accident rather than the head of the document the
|
||||
operator wrote."""
|
||||
codes = list(baseline.items)
|
||||
listed = ", ".join(repr(c) for c in codes[:_KNOWN_CODE_WINDOW])
|
||||
if len(codes) <= _KNOWN_CODE_WINDOW:
|
||||
return f"{len(codes)} known codes: {listed}"
|
||||
return f"{len(codes)} known codes, first {_KNOWN_CODE_WINDOW}: {listed}"
|
||||
|
||||
|
||||
def _reconcile_against_baseline(
|
||||
proposal: SavingsProposal, baseline: CostBaseline, tolerance: float
|
||||
) -> Rejection | None:
|
||||
|
|
@ -194,7 +238,7 @@ def _reconcile_against_baseline(
|
|||
if line is None:
|
||||
violations.append(
|
||||
f"unknown cost code {item.code!r}: not in project {baseline.project_id}'s "
|
||||
f"cost baseline ({len(baseline.items)} known codes)"
|
||||
f"cost baseline ({_known_codes_clause(baseline)})"
|
||||
)
|
||||
continue
|
||||
for field, claimed, actual in (
|
||||
|
|
|
|||
211
tests/test_named_known_codes_loadbearing.py
Normal file
211
tests/test_named_known_codes_loadbearing.py
Normal file
|
|
@ -0,0 +1,211 @@
|
|||
"""P22 DEL A — stage 0's unknown-code refusal NAMES the codes the project actually buys.
|
||||
|
||||
MEASURED (P21 stress round 5, re-measured at the head of økt 126): all 26 rejections across the
|
||||
six paid runs read ``unknown cost code '<invention>': not in project P's cost baseline (5 known
|
||||
codes)``, and **0 of 20** approaches validated (round 4: 4 of 20). The model invented
|
||||
``signalregulering_konstruksjon``, ``VENTIL_IMP``, ``RIGG01``, ``bærelag_asfalt`` and 22 more. It
|
||||
could not have done otherwise: the price schedule reaches the VALIDATOR and never the proposer,
|
||||
and the refusal states the COUNT of known codes, not one code's name. Step 5 feeds that sentence
|
||||
verbatim into the next attempt's prompt — and "you guessed wrong, there are five right answers"
|
||||
carries nothing to correct towards.
|
||||
|
||||
The contrast lives in this repo already. The MAGNITUDE half of the same stage NAMES the baseline
|
||||
value, and that is the half that let the loop converge in økt 94: the proposer found each correct
|
||||
figure because the refusal told it what the figure was.
|
||||
|
||||
**The window is a fixed COUNT, never a share** (the catalogue-excerpt rule, `_CATALOGUE_EXCERPT_
|
||||
CHARS`), and it counts CODES rather than characters for a measured reason: a character cut can
|
||||
sever a code mid-name and hand the proposer an identifier that exists nowhere — the
|
||||
``_index_excerpt`` rule inverted, where a path that never existed is worse than no path. A count
|
||||
window can only ever emit whole codes.
|
||||
|
||||
Measured sizes, with denominators (økt 126): every cost baseline anywhere in this repo or its
|
||||
measured corpora is at most SIX codes (the five context sets: 5 · 5 · 5 · 5 · 6; the two shipped
|
||||
``shared/examples`` baselines: 1 each; MAJOR-4's derivation of the synthetic K2 price schedule: 3),
|
||||
and the largest REAL delivered price schedule measured is K2's ``prissammenstilling-sheet-1.md`` at
|
||||
14 priced rows of 118 lines. Nothing measured reaches the window; it exists for the unmeasured
|
||||
R761-style mengdebeskrivelse, where a contract is priced BY prosesskode and the corpus declares
|
||||
2 727 of them.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Final
|
||||
|
||||
from portfolio_optimiser.generate import _build_messages
|
||||
from portfolio_optimiser.ir import (
|
||||
AffectedItem,
|
||||
CostBaseline,
|
||||
CostBaselineLine,
|
||||
SavingsProposal,
|
||||
)
|
||||
from portfolio_optimiser.reference_domain import Project
|
||||
from portfolio_optimiser.validator import (
|
||||
_KNOWN_CODE_WINDOW,
|
||||
Rejection,
|
||||
ValidatedProposal,
|
||||
rejection_stage,
|
||||
validate_proposal,
|
||||
)
|
||||
|
||||
_SMALL: Final = ("RIGG", "ASFALT", "MASSE", "FROST", "SKILT")
|
||||
|
||||
|
||||
def _baseline(codes: tuple[str, ...], project_id: str = "proj") -> CostBaseline:
|
||||
return CostBaseline(
|
||||
project_id=project_id,
|
||||
items={c: CostBaselineLine(code=c, quantity=100.0, unit_cost=1000.0) for c in codes},
|
||||
)
|
||||
|
||||
|
||||
def _proposal(code: str, *, quantity: float = 100.0, unit_cost: float = 1000.0) -> SavingsProposal:
|
||||
return SavingsProposal(
|
||||
project_id="proj",
|
||||
measure="Reduce scope",
|
||||
affected_items=[AffectedItem(code=code, quantity=quantity, unit_cost=unit_cost)],
|
||||
claimed_saving_nok=1000.0,
|
||||
assumptions={},
|
||||
)
|
||||
|
||||
|
||||
def _refusal(codes: tuple[str, ...], guess: str = "INVENTED") -> str:
|
||||
result = validate_proposal(_proposal(guess), baseline=_baseline(codes))
|
||||
assert isinstance(result, Rejection), "an invented code must never reach validated"
|
||||
return result.reason
|
||||
|
||||
|
||||
# --- (a) known-POSITIVE: the refusal names them --------------------------------------------------
|
||||
|
||||
|
||||
def test_the_refusal_names_every_code_a_small_schedule_carries() -> None:
|
||||
"""RED before DEL A: the refusal carried only ``(5 known codes)``. It now names all five, so
|
||||
the sentence Step 5 feeds forward is CORRECTABLE. The denominator stays — a reader must be
|
||||
able to tell how many exist from the same sentence that lists them."""
|
||||
reason = _refusal(_SMALL)
|
||||
for code in _SMALL:
|
||||
assert code in reason, f"{code!r} must be named: {reason}"
|
||||
assert "5 known codes" in reason, reason
|
||||
|
||||
|
||||
# --- (b) known-NEGATIVE: the gate can still FELL, and stays silent when it should -----------------
|
||||
|
||||
|
||||
def test_a_code_the_project_really_buys_still_validates() -> None:
|
||||
"""The control the order demands alongside the positive: a corrected gate can still be untrue
|
||||
for an independent reason. A proposal on a REAL line validates, so the naming above is caused
|
||||
by the code being absent — not by a stage that now rejects everything."""
|
||||
result = validate_proposal(_proposal("RIGG"), baseline=_baseline(_SMALL))
|
||||
assert isinstance(result, ValidatedProposal)
|
||||
|
||||
|
||||
def test_the_magnitude_half_does_not_grow_a_code_list() -> None:
|
||||
"""A REAL code at an invented magnitude falls on the OTHER half of this stage, which already
|
||||
names the baseline value. Listing the schedule there would be noise on a sentence that is
|
||||
already correctable — and would make the two halves indistinguishable to a reader."""
|
||||
result = validate_proposal(_proposal("RIGG", unit_cost=5000.0), baseline=_baseline(_SMALL))
|
||||
assert isinstance(result, Rejection)
|
||||
assert "tolerance around the baseline" in result.reason
|
||||
assert "known codes" not in result.reason, result.reason
|
||||
|
||||
|
||||
# --- (c)/(d) the window is BOUND, and it is a window and not a share ------------------------------
|
||||
|
||||
|
||||
def test_a_large_schedule_is_bound_by_a_fixed_window() -> None:
|
||||
"""A mengdebeskrivelse priced by prosesskode can carry thousands of lines. The list is bound at
|
||||
``_KNOWN_CODE_WINDOW`` codes, the denominator still states how many exist, and the cut is
|
||||
ANNOUNCED rather than left for the reader to subtract (``index_truncated``'s rule)."""
|
||||
many = tuple(f"P{i:04d}" for i in range(500))
|
||||
reason = _refusal(many)
|
||||
named = [c for c in many if c in reason]
|
||||
assert len(named) == _KNOWN_CODE_WINDOW, f"named {len(named)}, want {_KNOWN_CODE_WINDOW}"
|
||||
assert "500 known codes" in reason, reason
|
||||
assert f"first {_KNOWN_CODE_WINDOW}" in reason, reason
|
||||
|
||||
|
||||
def test_the_window_is_a_count_and_not_a_share_of_the_schedule() -> None:
|
||||
"""A share scales with the schedule again, only with a smaller constant — the exact regression
|
||||
the catalogue excerpt's fixed window exists for. Two schedules an order of magnitude apart name
|
||||
the SAME number of codes."""
|
||||
small_named = sum(
|
||||
1
|
||||
for c in tuple(f"P{i:04d}" for i in range(200))
|
||||
if c in _refusal(tuple(f"P{i:04d}" for i in range(200)))
|
||||
)
|
||||
big_named = sum(
|
||||
1
|
||||
for c in tuple(f"P{i:04d}" for i in range(2000))
|
||||
if c in _refusal(tuple(f"P{i:04d}" for i in range(2000)))
|
||||
)
|
||||
assert small_named == big_named == _KNOWN_CODE_WINDOW
|
||||
|
||||
|
||||
def test_a_schedule_exactly_at_the_window_is_not_announced_as_cut() -> None:
|
||||
"""Omission, never a lie in either direction: a schedule that FITS is not marked truncated and
|
||||
gets its whole list (``index_truncated``'s positive half)."""
|
||||
exact = tuple(f"P{i:04d}" for i in range(_KNOWN_CODE_WINDOW))
|
||||
reason = _refusal(exact)
|
||||
assert all(c in reason for c in exact)
|
||||
assert "first " not in reason, reason
|
||||
|
||||
|
||||
# --- (e)/(f) the codes it hands back are the project's own, whole and in its own order ------------
|
||||
|
||||
|
||||
def test_the_named_codes_are_the_schedules_own_order_never_sorted() -> None:
|
||||
"""The schedule's order is the PROJECT's order. A sort would invent a ranking the project never
|
||||
stated, and the first window would then be an alphabetical accident rather than the head of the
|
||||
document the operator wrote."""
|
||||
unsorted = ("ZZ-last", "AA-first", "MM-middle") + tuple(f"P{i:03d}" for i in range(50))
|
||||
reason = _refusal(unsorted)
|
||||
named = [c for c in unsorted if c in reason]
|
||||
assert named[:3] == ["ZZ-last", "AA-first", "MM-middle"], named[:3]
|
||||
|
||||
|
||||
def test_every_named_token_resolves_to_a_real_baseline_line() -> None:
|
||||
"""The ``_index_excerpt`` property, measured rather than assumed: each code the refusal hands
|
||||
back is a WHOLE key of the baseline. A character-sliced window could emit ``'P004`` and send
|
||||
the proposer after an identifier that exists nowhere."""
|
||||
many = tuple(f"code-{i:03d}-long-enough-to-slice" for i in range(100))
|
||||
reason = _refusal(many)
|
||||
listed = reason.split("first ")[1].split(": ", 1)[1].rstrip(")")
|
||||
tokens = [t.strip().strip("'") for t in listed.split(", ")]
|
||||
assert len(tokens) == _KNOWN_CODE_WINDOW, tokens
|
||||
for token in tokens:
|
||||
assert token in many, f"{token!r} is not a baseline code"
|
||||
|
||||
|
||||
# --- (g) the coupling P21's "a4 5/5 on stage0-baseline" rests on ----------------------------------
|
||||
|
||||
|
||||
def test_the_longer_sentence_is_still_labelled_stage0() -> None:
|
||||
"""``rejection_stage`` keys on ``"cost baseline ("``. Had the new clause moved that substring,
|
||||
every stage-0 refusal would have been relabelled ``"other"`` in silence, and P21's headline
|
||||
measurement (``must_refuse`` 5/5, all on ``stage0-baseline``) would have read as a regression
|
||||
caused by the fix meant to strengthen it."""
|
||||
assert rejection_stage(_refusal(_SMALL)) == "stage0-baseline"
|
||||
|
||||
|
||||
# --- (h) Step 5 carries it to the place it has to reach ------------------------------------------
|
||||
|
||||
|
||||
def test_the_named_codes_reach_the_next_attempts_prompt() -> None:
|
||||
"""The whole point: the refusal is only correctable if it reaches the proposer. Step 5 feeds
|
||||
the reason VERBATIM into the next attempt, so the codes ride for free — this arm measures that
|
||||
they arrive, rather than trusting that they do."""
|
||||
reason = _refusal(_SMALL)
|
||||
messages = _build_messages(
|
||||
Project(
|
||||
id="proj",
|
||||
name="Proj",
|
||||
description="",
|
||||
currency="NOK",
|
||||
cost_items=(),
|
||||
docs_dir="",
|
||||
),
|
||||
"context",
|
||||
Rejection(proposal=_proposal("INVENTED"), reason=reason),
|
||||
)
|
||||
prompt = "\n".join(m.text for m in messages)
|
||||
for code in _SMALL:
|
||||
assert code in prompt, f"{code!r} never reached the prompt"
|
||||
Loading…
Add table
Add a link
Reference in a new issue