Commit graph

315 commits

Author SHA1 Message Date
ac0bfdba27 feat(p21): a declaration that must have LOOKED, and a refusal that names the neighbours
C1. Round 4 produced 13 declarations over six runs and NOT ONE named a fasit concept.
The distinct documents opened before each were 1,1,1,1,1,1,1,2,5,5,6,13,13: seven
declared the base's FIRST requirement after opening exactly ONE document.

The order offered two rules and asked which discriminates. Replayed against the real
listings: "the declared document must have come back from a read_dir filtered on a word
from the approach's label" refuses 13 of 13 -- including Soraasen's 12.11, the closest
any run came -- because ZERO of the 13 were reached through a filtered listing at all.
A gate that refuses every measured case, right and wrong alike, cannot discriminate.
"fewer than k distinct documents opened" at k=3 refuses 8 of 13 and keeps the five that
navigated. k=3, 4 and 5 refuse the SAME eight -- the distribution has a gap between 2
and 5 -- so the threshold is not on a cliff, and 3 is the lowest of that plateau.
DISTINCT paths, not calls, and capped by the base's own size so a small base stays
declarable.

C2. Over the same traces 18 of 143 path-bearing calls named a path the base does not
hold, ELEVEN of them one run walking R761/4-3, 4.3, 4-2, 4-1, 4-0, 4-5, 4-6 while the
real names are R761/4, R761/41, R761/42. The refusal already named the nearest listable
ancestor; now it also names up to five of that rung's own subdirectories, ranked by
longest common prefix with the segment that failed. ONE copy shared by both refusal
sites, built from context_files through in_dimension, so every name handed back resolves
and the verdict layer can never be advertised in an apology. MEASURED after: 16 of 18.

Load-bearing MEASURED, four mutations all red against the WHOLE suite, green control
1863/5, golden byte-unchanged: C3(i) the declaration gate detached (2 red) . C3(ii) the
neighbour list empty (6) . C3(iii) built from files (1, the verdict arm alone) . C3(iv)
count CALLS instead of distinct documents (1, the repetition arm alone).

Three existing arms REWRITTEN, not weakened: all three read one document and declared,
which is the measured failure class exactly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 11:38:57 +02:00
7b4f85d77c feat(p21): the PROJECT carries the price, so a run against a road normal can be anchored
Four paid stress rounds ran entirely UN-ANCHORED, all of them, because the one file
loader reads cost-baseline.json out of the BUNDLE and no vegnormal ships one: N100,
N200, N500 and R761 are knowledge, and knowledge carries requirements, never amounts.
The validator's stage 0 -- the one stage that tells an invented cost line from a line
this project actually buys -- was skipped in every single run, so "validated" could not
mean what it says. P20 G1/G2 measured real R761 process numbers (12.11 three times on
Soraasen, 1.1.1 on Lindaas) validating with amounts nobody had anywhere.

--cost-baseline FILE is PM decision (e), taken over the three alternatives P20 wrote
down. A LOADED object, never a path (prepass_payload's rule): the CLI owns the file and
loads it ONCE, so the notice, the stamp and every base of an --across-bundle pass all
descend from one read. ONE parse, two doors -- load_cost_baseline delegates to
load_cost_baseline_file -- while safe_resolve stays on the bundle door alone, because a
project's own schedule is legitimately outside every base. No tolerant twin: this path
exists only because an operator NAMED a file.

DEL B: five anchored context sets, a1-a3 with their line and a4 with none, so stage 0 is
what catches the falsification arm. THE ORDER'S OWN ARM (h) WAS FELLED BY MEASUREMENT:
"no baseline code is a requirement number the base declares" is measured 0 of 4 on the
project-coded sets and 5 of 5 on kontrakt-sorasen -- which is what R761 Prosesskoden IS,
a bill of quantities priced BY process code. The complement keeps both, and the order's
own mutation still bites.

DEL B3: the judge reports anchored (off the run's own stamp), priced per row, and WHICH
falsifier caught the falsification arm.

Load-bearing MEASURED, five mutations all red against the WHOLE suite, green control
1850/5 (from 1809/5, superset, 0 removed), golden byte-unchanged:
A3(i) the flag is read but the baseline is unused (3 red) . A3(ii) only the first base
gets it (1) . A3(iii) report_forbidden drops it (1) . B2(i) a4 gets a line (1, arm (g)
alone) . B2(ii) a code swapped to 12.11 (2, arms (f) and (h)).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:49:10 +02:00
e513bc97ad fix(p20): code_forms follows ITS OWN approach, and the announcement has a witness [skip-docs]
Two defects the mutation battery and the paid round found, both measured before
being touched.

(1) code_forms described the WRONG candidate. Every per-approach artefact copied
the run's stamp and overrode only validator_decision, so an artefact about
approach 2 reported approach 1's codes. Measured in BOTH round 3 and round 4 --
and stress.py, which reads this field before re-deriving, then produced an EMPTY
prose_codes for every approach but the first, which is what round 3's table was
built on. The field's own comment already says it is stamped "off the proposal
being stamped"; run-level was the drift, not the intent. Model, citations and
token usage stay the run's, because they are the run's.

(2) The C2 announcement seam had no witness. Mutation C-iii reverted the call
site to `args.project_id or "the portfolio"` and the WHOLE suite stayed green
(1808/5): all three arms drove announced_subject directly. The missing arm drives
main() on a free dry run and reads the announcement off STDOUT, where an operator
reads it, and is red against exactly that mutation.

Sixteen mutations, ALL red against the whole suite. Green control 1809/5 (from
1781, +28, 0 removed), golden demo-transcript.stdout BYTE-UNCHANGED
(shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 08:28:44 +02:00
c8f0c8f7c4 feat(p20): the requirement that is RIGHT, and a clause number that is not a price
Three seams, one commit: A, B and C touch the same four modules (run.py carries
the debate task, the grounding composition and the announcement; okf.py carries
one reference-number vocabulary read by both A and B), so splitting them into
three commits would have meant hunk-level staging of entangled files. Stated
rather than silently restructured.

A — the declaration answers with the DOCUMENT's own words. Measured: 13
declarations over round 3 and P17b, not one naming a fasit concept, while the
tool answered {"declared": true, ...} by echoing the caller's own arguments. It
now returns the document's title and req_number, read off Bundle.context_files
(so the type: verdict layer can never be named back), plus the sentence saying
what the declaration binds. A path the base carries as no concept answers with
empty strings rather than refusing. The commission's success_criteria now reach
the DEBATE task through mandate.criteria_block, the one renderer, empty when
there are none — which is what keeps every un-commissioned prompt, and the
golden, byte-identical.

B — a clause number is not a price. THE ORDER'S OWN RULE WAS FELLED BY
MEASUREMENT: it asks to refuse a code that IS declared req_number/prosessnr,
and neither of its two known positives is. n500 declares seksjon 10.4.1..10.4.4
but never the bare 10.4; r761 declares 2727 prosessnr and 2753 seksjon, none of
them 1.10.4, which occurs once, as prose ("iht. vegnormal N200 kap. 1.10.4").
The COMPLEMENT fires on both and closes the hole _ground_against_input already
admits in writing -- "it fails OPEN on a coincidental match". Unanchored run +
requirement-shaped code + the base declares a vocabulary + the code is not in
it -> refused, naming the denominator. All five of kontrakt-sorasen's real
process codes ARE declared and pass, which is what keeps the one context set
built on real codes measurable. Replayed over all 24 codes of round 3 + P17b:
exactly the two known positives flip validated -> rejected, 22 unchanged.

C — a parse failure no longer burns the round ledger blind. _fetch_parsed takes
a BUILDER instead of a finished message list, so the retry carries the parse
reason; measured, kontrakt-sorasen-04 spent 11 of 12 rounds re-asking the same
question. And announced_subject names the routed bases instead of saying "the
portfolio" for a two-base commission.

Suite 1807/5 (from 1781, +26, 0 removed), golden demo-transcript.stdout
BYTE-UNCHANGED (shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f),
ruff and mypy clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 06:02:46 +02:00
c4e88003e2 fix(p17b): close the flags multi-base mode neither carried nor refused [skip-docs]
Measured after the paid run, not before it: the across-bundle door honoured
fourteen flags and refused five, which left eight accepted and then dropped. The
worst of them was ``--mcp-config`` -- configured egress with nothing printed,
which this repo forbids outright -- and ``PROJECT_ID``/``--docs-dir``, which
would LOOK honoured while the dispatch read each base's project from that base's
own IR projection and used each base as its own docs dir.

The two anchoring flags are WIRED rather than refused. They are bundle concerns
and this dispatch hands ``run_project`` one bundle at a time, so they compose
exactly -- and ``--require-cost-baseline`` is the named remedy for the defect
this session's own paid run measured (``1.10.4``, a requirement number accepted
as a cost code on a base with no schedule: P19 F1, now reproduced on a second
base). Wiring the free drill too, so the dry run and the paid run cannot
disagree about what the run will do.

The two ``requires --bundle-dir`` guards no longer answer for this mode: falling
through would tell an operator to add the one flag this mode also refuses, which
is the repo's own standing objection to that pattern.

Mutation (v) -- the requirement reaches the dispatch and never the per-base runs
-- is red on its own arm against the whole suite. Suite 1781/5, golden unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 05:02:05 +02:00
da0ccd0489 feat(p17b): a context set that spans TWO bases, and a judge told which one [skip-docs]
``contexts/dekke-og-kontrakt-lindaas-2027`` is the first set whose approaches
route at more than one knowledge base: a1/a2 at n200-2024 (material requirements)
and a3/a4 at r761-2025 (the rig, and the falsification arm). That is the whole
reason it exists -- P17b measures that ONE commission can be run across several.

``bundle.txt`` grows a block per base; a set naming one base is one block, so the
four pre-P17b files parse byte-identically. The reader now has ONE home
(``stress.read_bundle_declarations``): it used to be a private copy in the P14
gate and a second, looser one inside ``stress.main``, and the multi-base form is
exactly the change that would have let them drift.

Rule U becomes the UNION of every declared base, and that is not a formality.
MEASURED 15.09: ``enhetspris`` is absent from n200-2024 and carried by 70 of
r761-2025's 2 756 concepts, so anchors admitted per base would have admitted a
question the pass as a whole CAN ground. It was dropped from the fifth set's
anchors for that reason.

``score_context_set(bundle_id=...)`` restricts the judgement to the approaches
routed at THIS base. Without it, judging the n200 outbox reports the r761
approach as ``not_evaluated``/``absent`` -- a false finding, because that
approach WAS evaluated, against the other base, under the other run_id. That
defect is pinned by its own arm. The judge's CLI refuses to guess when a set
declares several bases, with an rc-0 control on ``--bundle``.

Arm (d) gained a second half: every DECLARED base must be named by some
approach, because a base no approach names is never run.

The P19/B2 fasit denominator moved 26 -> 32 and is asserted, not dropped: six new
references, two of them bare ``prosessnr`` (12.11, 12.12), so B1's
punctuation-and-digits form is now exercised by a fasit and not only by a
known-positive.

Suite 1774/5, golden byte-unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 04:44:37 +02:00
5e4c497a84 feat(p17b): ONE commission, SEVERAL bases -- reachable from the command line
``run_mandate_across_bundles`` has existed since session 58, reachable from FIVE
test files and from NO command line (measured: ``grep -n across-bundle run.py``
= 0 hits). ``--across-bundle <dir>``, repeated once per base, is that door.

The engine takes a CALLBACK rather than an outbox directory. Its own docstring
has always said N runs need N ``run_id``s and that minting them there would
default a key this repo requires a caller to supply -- so ``outbox_for`` is that
contract KEPT, not relaxed, and the operator-chosen ``<run-id>-<bundle_id>``
rule lives in ``main()`` where the decision was made. The order's alternative (a
caller running ``run_project`` itself over ``route_by_bundle``'s sub-mandates)
would be a second copy of the loop's id reconciliation, shared store, per-base
project resolution, collision accounting and both budget teeth.

``resolve_bundle_routing`` is ONE resolution shared by the engine and the
dry-run arm: a free trip answering with a different project id, or tolerating a
duplicate id the paid dispatch refuses, would rehearse a different run.

``{run-id}-multibase.json`` is written from a ``finally`` and every row is built
from the resolution plus disk, so the pass a cap cut short still leaves the
record. ``completed`` is a required field for ``ExplorationTrace.completed``'s
reason. ``stop_reason`` is read BACK from each base's own coverage artefact.

Load-bearing MEASURED (17 arms), four mutations all red against the WHOLE suite,
green control 1761/5 (from 1744/5, superset, 0 removed), golden byte-unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 04:24:48 +02:00
4c6084e5df feat(p19): the trace says HOW, and a run says what it spent and why it stopped
DEL C. P18 gave read_dir a window (filter/offset/limit) and then measured its
own paid round without being able to see it used: five of 31 documents read
lay outside the default window, so the window HAD been widened and the trace
could not say with which knob. ToolCall now carries the three arguments,
always present and empty/zero when not passed -- an absent key and "not
narrowed" must not read the same -- and the judge counts filter_calls and
paged_calls. _number_argument is a SIBLING of _string_argument, not a widening
of it: a model may send limit as 10 or as "10", and a reader that knew one
shape would report a paged call as unpaged.

DEL D. P18's finding 4 was WRONG AS WRITTEN. provenance.token_usage has been
stamped on every proposal artefact since S3.4 and stands in every one of round
2's; what was missing is a READER. The judge reads it now (round 2 measured:
289 054 tokens against round 1's 2 679 305, -89 %), and the P18 report gets a
dated correction UNDER its original paragraph rather than instead of it.

What was genuinely absent is {run_id}-coverage.json. settle prints the
coverage report and ApproachOutcome has carried not_evaluated since Trekk A3,
but neither ever reached a file, so a judge could see an approach had no
artefact and could not tell a budget stop from an approach nobody ordered.
Written from the finally IFF a mandate was given. stop_reason comes from a
CALLER-OWNED sink rather than from in_flight, and that is a measurement:
_evaluate_mandate SWALLOWS BudgetExceeded once something has been produced, so
run_project's own in_flight never sees it.

Load-bearing measured (10 arms), four mutations all red against the whole
suite, green control 1744/5 and the golden byte-unchanged. D-i stood GREEN
first -- the vacuous-gate class, 25th time: the arm called write_coverage
itself and therefore chose the reason it then asserted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 03:05:52 +02:00
d74f32dc1c feat(p19): a cost code must have a FORM where the input offers forms
P18's round 2 ended with two VALIDATED proposals whose affected_item codes
were ordinary words from a road standard's prose -- impulsventilator (4 of
270 N500 documents) and bituminoest baerelag (4 of 1133 N200). Both are
grounded in P7's sense and neither is inert in P18/B1's sense; they are simply
not identifiers of a cost line, and the gate had no stage that could say so.
Known positive MEASURED, not asserted: replayed offline against the bases
those runs were given, both come back Rejection naming the denominator.

IDENTIFIER_FORMS moved from generate.py to validator.py: they now drive both
P8's report and this gate, and two copies of "what an identifier looks like"
would let the two disagree about one run's own input.

B1 -- two new forms, transcribed from measurement. R761's requirement numbers
are bare dotted numbers and all six refs in kontrakt-sorasen's fasit are of
that shape, which neither pre-P19 form matched: r761's whole offer was 3
identifiers over 6.5 MB, and is now 2332. The FIRST form was widened in the
same pass because B2 made these forms decide prose vs identifier, and this
repo's own ENERGI-TOTAL-EL matched none of them -- a gate may only be wrong in
the direction that admits too much.

Three things keep the gate from being a rule about shapes: the generality
guard (it fires only where the input offers forms), the baseline exemption
(stage 0 has already ruled that code real), and full-matching.

Honesty limit, measured and given its OWN arm: a decimal and an R761 process
number are typographically identical, so the form counts both. P8's existing
"bare numbers" arm is narrowed to bare INTEGERS accordingly.

Measured over all nine round-1+2 outboxes: 26 of 36 codes are prose.

Load-bearing measured (22 arms), four mutations all red against the whole
suite (16 / 10 / 1 / 2), green control 1734/5 and the golden byte-unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 02:11:29 +02:00
c84e8bf6f1 feat(p19): a direction must NAME the requirement that binds it, and have READ it
Two paid rounds scored 0 of 26 fasit concepts opened -- the same number twice.
P18 closed the navigation side (a listing is a window, an invented path is
refused by name) and it did not move, which makes it a ROLE question: nothing
in the loop ever asked the model to say what requirement binds the direction it
committed to, so opening one was never on the critical path to an answer.

A PREMISE OF THE ORDER WAS FELLED BEFORE ANYTHING WAS BUILT ON IT. A1 places
the demand in _INSTRUCTIONS[HYPOTHESISER_ROLE] alone. Measured: the stress
command sends --mandate and NOT --explore, the two are refused together by
name, and none of the nine round-1/2 outboxes holds a {run_id}-exploration.json
-- the hypothesiser never runs in a stress round, so A3 would have been
unreachable in exactly the paid runs this order commissions.

A2's own sentence resolves it: the refusal goes to the model "som en tur den
kan rette (samme mekanisme som quick_validate's nekt), ikke som en raise" --
and quick_validate IS a tool. declare_requirement therefore lives in
navigator_tools, held by BOTH roles that navigate (the exploration, and since
S2c the debate). It EXISTS only when the caller offers both sinks, which keeps
every pre-P19 call site byte-identical; one sink without the other is refused
at construction. 'opened' is the SAME list ExplorationToolRecorder fills, so
the refusal reads the run's own read trace.

The marked hypothesis carries 'requirement' as a REQUIRED key: omitted is a
hard error, explicit null is legal and needs 'why_none', a half-named one is
refused. A minted approach carries it; a seed never acquires one. The proposer
prompt names it only when the field exists, and the judge counts a hit against
THIS approach's fasit concepts, never against the base.

Load-bearing measured (12 arms), four mutations all red against the whole
suite, green control 1711/5 and demo-transcript.stdout byte-unchanged.
A-iii's predicted signature was FALSIFIED: the golden stays green because the
demo runs without a mandate, so _build_messages' approach branch is never
taken there. A-iv was GREEN first -- the repo's vacuous-gate class, 24th time:
the arm drove _attributable while the hit is computed at the call site.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 01:21:47 +02:00
e7ba367f9d docs(p18): stress round 2 -- five paid runs, and the mutation that found a hole
P18 parts D and E (order 20260914T105139Z), plus the two things measuring
them turned up.

DEL D -- five paid runs (gpt-4-1-mini, azure, PACE_SECONDS=2), same four
context sets, SAME parameters on all four (--max-rounds 3 --max-tokens
600000), plus one variance repeat of gate-nordvik. All four free
--live-dry-runs first: rc 0, and Grounding-offer numbers IDENTICAL to round 1
(435/272/982/3) -- the control that the Grounding structure did not change
what the gate measures.

MEASURED, round 1 -> round 2:
- runs that died on the token cap: 3 of 7 -> 0 of 5, and two sets now fit a
  LOWER cap than round 1 had to give them;
- guessed read paths: 12 -> 0;
- wall time, the two sets whose parameters are directly comparable: n100
  140.1 s -> 70 s, n500 73.5 s -> 69 s;
- 5 of 31 read_file calls opened documents BEYOND the default window, so the
  window WAS widened -- the trace does not record which knob (finding 1);
- (a) grounded in a fasit concept: 0 of 26, UNCHANGED. That is the mission
  gap, and DEL A did not close it.

The a4 falsification arm fails once in each round, on a different set. r761's
a4 is now rejected -- but NOT by B1: the model proposed "Kontraktsum" this
time, so the "appears nowhere" arm caught it, and B1's effect on that row is
proven offline, not live. NEW failure: tunnel-hauglia a4 VALIDATED on
"impulsventilator" (3/270 documents), and fv412 a1 on "bituminost barelag"
(4/1133). Both are ordinary Norwegian words from the standard's prose, not
cost codes. B1 cannot and should not fell them: this is an ANCHORING defect,
not a grounding one, and it is finding 2 with two named remedies and a
recommendation.

C2 isolated by re-judging round 1 with the new judge: kontrakt-sorasen goes
named=3 -> named=1, and the survivor is named_in_measure -- the model's own
words. Two of the three were the whole-base snippet artefact.

DEL E -- docs/2026-09-14-p18-stressrunde-2.md: round 1 against round 2, what
each fix bought (measured, never attributed), the B2 table, variance, and for
EACH remaining ugly finding a NAMED solution with an estimate.

THE MUTATION THAT FOUND A HOLE. B6 (revert run.py to compose ONE blob instead
of one document per concept file) left the WHOLE suite green: 1698 passed / 5
skipped. The composition arm drives _grounding_text with a Grounding it
builds ITSELF, so it cannot see what the RUN handed over -- and a blob has
exactly one boundary, so the floor can never be reached, the share can never
fire, and the measured defect is back intact. The rule is only as good as the
boundaries it is given.

Arm (h) is the gate that was missing: a crafted base with TWELVE concept
files all carrying the same token -- per document 12 of 17 and inert, as one
blob 1 of 1 and grounding -- with a control on a code only ONE file carries,
which must still validate. Measured RED against exactly that mutation. The
mutation was not dropped and the seam was not declared unwitnessed: it got a
witness.

Also: the debate's own bundle pointer (run.py _bundle_pointer) now explains
the window and the filter, alongside the tool description and the navigator
instruction updated in 9b47e5a -- a description that lies about the body IS
the model's instruction (the Fase 3 class). Golden transcript unaffected.

Mutations, all against the FULL suite in an isolated worktree, one at a time:
DEL A 7 of 7 red (control 1685/5), DEL B+C 9 of 10 red (control 1698/5), the
tenth being B6 above. Tables in the report s 9.

Verification: uv run pytest -q 1699 passed / 5 skipped (1670 on cfd9079; +29,
0 removed). ruff check + format clean, mypy clean (38 files). Golden
demo-transcript.stdout BYTE-UNCHANGED, shasum -a 1 of the CONTENT =
ea8c534773acdbe41ae68f2c55724d69aaf8be4f. No version bump, no push.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 00:03:26 +02:00
7a7c988253 feat(p18): an identifier that stands everywhere identifies nothing
P18 parts B and C (order 20260914T105139Z).

B1 -- stage 0b. P7 made it `item.code in grounding`: plain containment over
ONE concatenated string. P16 ran it against a delivered corpus and measured
what containment cannot tell apart: the falsification arm a4-indeksregulering
put 250 000 NOK on a single cost line coded R761 -- the knowledge base's OWN
NAME, carried by all 2 756 of its concept documents -- and the whole gate
said validated (stage 0 skipped, un-anchored run; checker approve).

The grounding is now carried as the DOCUMENTS it is made of (validator.
Grounding), not as a blob. A structure and not a second argument beside the
text: the boundaries and the text are one fact, and .text is derived, so the
gate and P8's report measure the same characters. run.py composes one
document per concept file where the base is already walked; generate.
_grounding_text folds each cost line in as a one-line document.

N and A are MEASURED, not chosen (14.09, four mounted vegnormal bases):
- every must_cite ref and mandate affected_code in the four context sets --
  shortest real identifier is FOUR characters (12.1, 52.1), so N = 3 sits one
  below the measurement and cannot refuse anything measured;
- document frequency of every code-shaped token per base -- 1 692 distinct
  and NOT ONE reaches 5 %. Highest anywhere 6/446 (1.35 %), highest a fasit
  names 3/446 (0.67 %), R761 2 756/2 756 (100 %). A = 0.05 therefore sits
  3.7x above the highest real token and 20x below the defect.
Length is NOT what makes the defect inert (R761 is four characters); the
share is. And a share is not a measurement without a denominator big enough
to take one (ansikt 4): one of three is 33 %, so an ABSOLUTE floor of 10
documents gates it. Highest absolute count any real identifier reaches is 6,
and every fixture in the repo is far below 10 -- which is why every pre-P18
gate is UNTOUCHED by this rule rather than exempted from it. Grounding.of
(one document) can never reach the floor by construction.

The refusal NAMES the denominator ("appears in 2756 of the 2756 documents
this run was given"), because Step 5 feeds that reason verbatim into the next
attempt's prompt: a proposer told only "ungrounded" answers with another
token of the same kind.

B2 SPIKE (measured, NOT built) FELLED the order's own alternative: option (b)
"ground in what the run OPENED" was run over P16's 16 code rows -- R761
stands in every OPENED document too, so (b) would NOT have caught the defect,
while B1 makes it inert and still grounds the real process line 65
ASFALTDEKKER (29/2756 = 1.05 %). (b) is not a substitute for B1.

C1 -- --docs-dir is optional once --bundle-dir is given (P16 FUNN 2). On the
bundle path docs_dir is never read: retrieval, the chunk tool and the "no
citable content" check all live in the road branch. Bound ONCE from
--bundle-dir, which is byte-identically what the README already tells an
operator to type by hand. NOT the "--docs-dir omvei": no such path is opened
and the road branch still refuses without a real --docs-dir (own arm).

C2 -- the judge's snippet arm counts only under citation_scope == "narrowed",
as (a) already does (PM decision, P16 s 6.2). P16's reason for (b') being
clean -- snippets are bodies while ref/title live in frontmatter, 0 of 446
n100 bodies -- holds for "Krav 4.1.2-1" but NOT for R761, where a process
number like 12.1 stands in the bodies. Under a whole-base citation list that
mark was "cited" before any model call.

tests: test_inert_identifier_loadbearing.py (7 arms; known positive is P16's
OWN artefact replayed against the base that run was given, known negative is
26 of 26 fasit references still grounding), test_docs_dir_optional_
loadbearing.py (5 arms). test_stress_judge_loadbearing.py's snippet arm split
into narrowed/whole-base -- the pair is the discriminator, same snippet, same
mark, only the scope differs. The grounding tests migrate from str to
Grounding.of (the honest reading of a caller that declared no boundaries).

Verification: uv run pytest -q 1698 passed / 5 skipped (1685 after part A,
strict superset, 0 removed). ruff check + format clean, mypy clean (38
files). Golden demo-transcript.stdout BYTE-UNCHANGED, shasum -a 1 of the
CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f. No version bump, no push.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 22:28:40 +02:00
9b47e5aa62 feat(p18): a listing is a WINDOW, and an invented path is refused by name
P18 part A (order 20260914T105139Z). P16 measured S7a-3's ladder against a
delivered corpus for the first time and found two things fixture bases
cannot show.

(1) One level is not bounded by being one level. Measured 14.09 on the four
mounted vegnormal bases: okf.directory_listing on krav/N200 is 169 974 chars
over 1 132 documents, krav/N100 69 250 over 445, krav/N500 39 853 over 269,
and R761's own root 110 912 over 2 728 SUBDIRECTORIES -- 27-113x the
1 500-char ceiling S7a-3 set, riding in every later prompt. That last number
is why the window covers BOTH kinds: a pagination over documents only would
have left the largest measured level unpaginated.

read_dir now answers with a window. offset/limit page directories first then
documents as ONE sequence (two independent windows make "the next ten" a
question with two answers); total is the denominator and is always carried;
limit is CLAMPED to 50, never refused. Default 10 chosen against the ceiling:
one entry is 121-209 chars (median 145) over the four bases. After: n100
1 493, n500 1 453, R761 479, n200 1 537 -- 2.5 % over, stated rather than
tuned away, because the ceiling is a character budget and the window is a
count. Largest single call any caller can make: ~7 600 chars.

filter narrows a level instead of paging it: case-insensitive SUBSTRING over
title + req_number/prosessnr and over a directory path, answering with
total_matches beside total. A substring and not a pattern for
_ground_against_input's reason one rung down -- a form the rule does not know
returns nothing, and an empty listing reads as "the base does not have this".
A filter that matches nothing is an ANSWER (total_matches: 0), never a
refusal. ORDER PREMISE FELLED before building on it: the order asks for a
separate top-level reader "like own_frontmatter" because parse_frontmatter
was last-write-wins -- P15 (f13dc64) already made a top-level key win, so
BundleFile.frontmatter IS the concept's own value and a second reader here
would be the second copy ko-(p) forbids.

(2) 0 of 26 fasit concepts were opened in 32 read_file calls (the order's
"24" is the four runs' DISTINCT paths, re-measured 14.09), and 10 of those
calls named a path the base does not hold. Each reached the model as MAF's
opaque "Error: Function failed." while counting toward the three consecutive
tool errors that end a request. read_file now refuses such a path by name
(BundlePathNotFound, funn-99 returned form) and names the nearest directory
that actually HOLDS documents -- chosen off context_files, never the
filesystem, because a directory can exist on disk and hold no navigated
concept (read_dir would then refuse the very path the refusal handed back)
and because context_files is what drops the type: verdict layer, so a refusal
can never advertise by name the one layer no listing mentions. Narrow by
construction: only an ABSENT path is translated; any other OSError propagates
untouched.

Two existing arms REWRITTEN, neither weakened:
- test_a_nonexistent_sibling_is_still_an_os_error was a tripwire whose own
  docstring said "when it goes red, someone has closed it, and that is a
  decision to be recorded". This is the record. Its narrowness half survives
  as a new arm driving a real PermissionError on a file that IS there.
- test_every_document_is_still_reachable_and_the_counts_add_up became
  STRONGER: the accounting must now page, so the same assertion also proves
  the window is complete and non-overlapping.

tests/test_navigation_window_loadbearing.py: 13 arms. Arms needing the
delivered bases SKIP with the root named (PORTFOLIO_VEGNORMAL_ROOT), as
MAJOR-3's ceiling arm does; the window algebra, the filter negative and the
refusal run over a synthetic base UNCONDITIONALLY, so the file can never be
silently absent in full.

Verification: uv run pytest -q 1672 passed / 5 skipped before the new file
(1670 on cfd9079). ruff check + format clean, mypy clean (38 files). Golden
demo-transcript.stdout BYTE-UNCHANGED, shasum -a 1 of the CONTENT =
ea8c534773acdbe41ae68f2c55724d69aaf8be4f. No version bump, no push.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 22:10:43 +02:00
a61a3ccda7 fix(p16): the pre-call announcement must state the cap the run will use, not the constants
announce() read _DEFAULT_MAX_ROUNDS/_DEFAULT_MAX_TOKENS directly, which was correct only while
main() could not do otherwise. MEASURED on the first free drill after --max-rounds landed: the same
stdout said "Stops at: 3 rounds / 100000 tokens" two lines above "max_rounds=8, max_tokens=120000".
The announcement is the ONE thing printed before the first paid call and its whole job is to say
what the run will do -- the Fase-3 class, introduced by the very flag being announced.

Load-bearing MEASURED: arm (f) red before the fix; M16 (read the constants again) -> 1 red, that
arm ALONE. Green control 1670/5, golden BYTE-UNCHANGED (ea8c534...).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 12:21:29 +02:00
5f94c92476 feat(p16): --max-rounds/--max-tokens -- the cap on a PAID run had no operator door
B2's free drill earned its keep on the first command. STATE.md, docs/2026-09-12-p14-kontekstsett.md
and the order all publish the same stress command ending "--max-rounds 8 --max-tokens 120000".
Measured: run.py accepts neither, all four --live-dry-run drills refused with "unrecognized
arguments", and main() never passed max_rounds/max_tokens to run_project at all -- so every CLI run
ever made was silently bound to _DEFAULT_MAX_ROUNDS=3 / _DEFAULT_MAX_TOKENS=100_000, with no way to
raise or lower the cap on a run being paid for. Three surfaces described a door that did not exist.

Widening, never breaking: both flags default to exactly those values, so every existing invocation
is byte-identical. Wired to BOTH dispatches -- run_portfolio takes the same two parameters and
main() dropped them there too -- and refused by name in report mode, which returns above every
dispatch (the F4 silent-drop gap).

Load-bearing MEASURED (6 arms, ALL RED before the fix), four mutations all red, green control
1669/5 (from 1663/5, superset, 0 removed), golden BYTE-UNCHANGED (ea8c534...).

MEASURED, REPORTED, NOT FIXED: single-project mode still requires --docs-dir even when
--bundle-dir is given and docs_dir is unused on the bundle path, so the documented command would
have refused for that reason too. The README's own form works; loosening the guard is its own call.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 12:10:59 +02:00
f21007c858 feat(p16): the stress judge -- and the order's own (a) was a gate that could only be green
Session 102's criterion ((a) built on the right fasit concept OR refused anchored, (b') names it,
(c) zero hallucinations) was adjudicated BY HAND. Measured 14.09: nothing in the tree read
contexts/<set>/fasit.json against an outbox at all, so "provable against the base" had no
repeatable form. portfolio_optimiser.stress reads ONLY artefacts that already exist -- the
per-approach proposal/outcome pair and {run_id}-debate.json -- so no run gains a field.

MEASURED BEFORE BUILDING: the order defines grounded as "OPENED or CITED", but on the S2c path
run_project stamps citations = bundle_citations(bundle), one per context file. On n100-2023 that
is 446 citations over 446 concepts, and 6 of 6 fasit paths are already "cited" before a single
model call. Honouring it literally would be the repo's own vacuous-gate class inside the gate
built to catch it, so a citation grounds an approach only under a NARROWED list (a declared
pre-pass cut); both halves are reported either way. (b') was checked for the same vacuity and is
clean -- snippets are bodies, ref/title live in frontmatter (0 of 446 n100 bodies carry
"Krav 4.1.2-1") -- so the order's definition stands.

A2: unanswerable questions had no runnable form (po is not a lookup tool), so they become a FOURTH
commissioned approach per set whose cost line the base carries no ground for, and fasit.json
carries must_refuse INSTEAD of unanswerable -- one form, never two copies of one fact. Rule U is
untouched and its known-positive is still red.

Load-bearing MEASURED (20 arms), eleven mutations all red on their own arm, green control 1663/5
(from 1643/5, superset, 0 removed), golden demo-transcript.stdout BYTE-UNCHANGED
(shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 11:58:48 +02:00
f13dc64a0a fix(okf): top-level frontmatter title survives a nested sources: title
P15 (order 20260912T220951Z). okf._frontmatter_from_text was linewise
last-write-wins over EVERY line regardless of indentation, so a curated
concept's own top-level `title:` got silently overwritten by the nested
`sources:\n  - title: ...` block's title. Fix: a top-level (unindented)
key always wins over an indented one of the same name; a nested line
with no top-level counterpart is still preserved (SPEC §4).

Red-before/green-after: new test
test_parse_frontmatter_top_level_title_survives_nested_sources_title
(tests/test_okf.py) failed on 45edbf5 (fm["title"] == "N500:2024",
expected the concept's own), green after the fix.

Re-measured on all four vegnormal-okf bases (concept files / distinct
titles): n100-2023 446/446 (was 1) - n200-2024 1133/1133 (was 1) -
n500-2024 270/270 (was 1) - r761-2025 2756/2407 (genuine repeated
process names, not a collapse). directory_listing on krav/N500:
269 documents / 269 distinct titles (was 1).

tests/test_context_sets_loadbearing.py:
- The P14 tripwire test (asserting parse_frontmatter DID collapse
  titles) is INVERTED, not deleted, per the order: it now asserts the
  fix holds, as a live regression guard.
- own_frontmatter() stays (not replaced by parse_frontmatter): measured
  29,500 field reads (type/title/req_number/prosessnr, all four bases)
  agree exactly except for quote-stripping (2,728/29,500, zero value
  mismatches) - own_frontmatter unquotes for fasit comparison,
  parse_frontmatter deliberately doesn't (D1/(a)/(i): unquote_scalar is
  the ONE unquoting rule).

docs/2026-09-12-p14-kontekstsett.md Part B correction: the "22 of 22
cost words absent from n100/n200/n500" claim was false - n500-2024
carries `kroner` as a false positive (substring match inside
"borkroner", drill bits, not money). The original 22-word list was
never persisted, so only ~9 of the 22 survive named. Replaced with a
newly named, persisted 22-word list and the actual re-measured count:
n100 22/22 absent - n200 22/22 - n500 21/22 (kroner via borkroner) -
r761 18/22 (4 genuine cost words). No gate touched (no fasit anchor is
`kroner`).

Verification: full suite 1643 passed / 5 skipped (was 1642/5 on
45edbf5, +1 new test, 0 removed) - `uv run pytest -q`. ruff check +
ruff format --check clean on the three changed source/test files.
Golden transcripts byte-unchanged: shasum -a 1
tests/golden/demo-transcript.stdout = ea8c534773acdbe41ae68f2c55724d69aaf8be4f,
demo-transcript.stderr = ede3e2f685ce6a14ad9888e9de421d1a66f6c611.
No version bump, no push (both forbidden by the order).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-13 07:40:57 +02:00
45edbf5957 test(p14): four stress-test context sets, gated -- and the fasit titles had to be read from the concept's OWN declaration
P14 valg (a): en bundle per kjoering, fire kjoeringer. Ingen modellkall, ingen
Azure, ingen produksjonskode roert -- leveransen er fire datasett, en gate og
maalingen bak dem.

FORMEN: contexts/<prosjekt>/ med mandate.json (Mandate ordrett), bundle.txt
(symbolsk basenavn + erklaert bundle_id -- aldri en absolutt sti, som ville
pinnet settet til en maskin og ridd ut i `git archive HEAD`), fasit.json og en
TOM docs/. Fasiten er EN maskinlesbar fil, ikke fasit.md + en tvilling: to
kopier av ett faktum er ko-(p), saa prosaen bor INNI JSON-en.

REGEL U (den maalbare formen for "basen kan ikke svare"): hvert ubesvarbart
spoersmaal erklaerer >=1 anchor, og admitteres iff HVER anchor er fravaerende
fra HELE teksten i HVERT konsept. Ikke "deler ingen noekkelord med noen tittel"
-- et tunnelspoersmaal deler "tunnel" med hundrevis av titler og det beviser
ingenting. Skanningen baerer alltid NEVNER; null konsepter er ROEDT.

MAALT, og verdt hele ordren: 22 av 22 proevde kostnadsord er FRAVAERENDE fra
n100/n200/n500 (r761 baerer 4). po sitt oppdrag er aa finne kostnadsbesparelser,
og tre av fire baser inneholder ikke ett pengeord.

FUNN, maalt og IKKE fikset (egen ordre): okf.parse_frontmatter er linjeorientert
last-write-wins, saa sources-blokkens innrykkede title overskriver konseptets
egen -- directory_listing paa krav/N500 returnerer 269 dokumenter, ALLE med
"title": "N500:2024". Navigasjonsstigens rung 2/3 skiller dem kun med et
UUID-filnavn og et tegnantall. En fasit-assert mot den tittelen ville vaert
VAKUOES, saa gaten leser toppnivaa-noekler og baerer en TRIPWIRE som asserterer
at kollapsen fortsatt finnes.

Load-bearing MAALT (36 armer), seks mutasjoner alle roede paa sin egen arm +
groenn kontroll 1642/5 (fra 1606/5, supersett, 0 fjernet) og golden BYTE-UENDRET
(shasum -a 1 av INNHOLDET = ea8c534773acdbe41ae68f2c55724d69aaf8be4f). Gaten var
ROED foer settene fantes.

AErlighets-grenser: de fire prosjektene er OPPDIKTET (hvert sett sier i sitt eget
honesty-felt hva jeg konstruerte); Regel U beviser at ORDET mangler, ikke at
spoersmaalet er ubesvarbart; de bundle-krevende armene SKIPPER uten basene; og
INGEN kjoering er gjort -- dette er maaleoppsettet, ikke maalingen. (c) er ikke
bygget: motoren finnes ferdig, det som mangler er ett nytt flagg, en run_id-
myntingsregel operatoeren maa ta, og tre partisjons-rader.

Ordre: 20260912T202210Z-7590723260-from-.claude
Maaling: docs/2026-09-12-p14-kontekstsett.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 23:48:20 +02:00
fed69790ac fix(okf): close the inert ingest-stamp guard, then land okf 0.8.5 -- and read the block sources form all four bases actually write
P13 measured this lift and REFUSED it, because okf >=0.8.5 emits the ownership
stamp as the V1 flow mapping `generated: { by: process:okf-ingest, at: ... }`
where 0.3.2 emitted `true`, and `_carries_complete_ingest_stamp` read the new
form as NOT a stamp -- write_concept_file's forgery refusal would have shipped
DISARMED with the whole fail-closed suite green. That blocker is closed first,
red-first, and then the pin moves.

ROW 1, THE SECURITY HALF. `_claims_ingest_ownership` widens the predicate from
"reads as boolean True" to "claims ingest ownership", of which the boolean is
the pre-V1 spelling. The recogniser for the new half is `decode_flow_value` --
the module's ONE flow decoder, the same argument write_concept_file already
makes for `verified`: the writer refuses exactly what the reader can read. A
value the decoder REFUSES is therefore not an ownership claim and writes
through, which is what keeps this from collapsing into "any non-empty
generated". Two arms red before the fix; no YAML library introduced.

THE PIN. okf v0.3.2 -> v0.8.5, guard v0.3.4 -> v1.4.0 spelled `tag =`, not
`rev =`, and not the declared floor 1.2.0 -- both P13 premises hold and the
reason now lives next to the pin in pyproject.toml. The ":40" comment is
corrected: okf has ONE runtime dependency, the guard, and that is what binds
the two lines together. 27/27 imported names resolve across five modules.

THE GOLDENS, REGENERATED AS A DECISION. Seven concept files across four
examples/ingest-golden-* bundles, one line each. Two were regenerated by the
REAL materializer; the other five are derived (http/sql/mcp cannot materialize
outside the tests' stubs) and then MEASURED -- all four golden suites compare
byte for byte against what the stubs produce, and all four are green. The four
`generated == "true"` asserts now read ONE source, conftest.
expected_generated_stamp: four literals for one emitter fact are four places a
later release can leave half-corrected, which is exactly how the pre-V1 form
survived until P13 measured it. tests/test_okf.py keeps its literal on purpose
-- that one round-trips a CURATED half-stamp through our own writer.

THE BLOCK READER. Measured with the full denominator: all four delivered
knowledge bases write `sources` as a BLOCK sequence and none in flow form
(n100 446/446, n200 1133/1133, n500 270/270, r761 2756/2756 = 4605/4605), and
`evidence_for` reported `unreadable` on 4605 of 4605 -- the falsification layer
had no address for any document in any base. `okf.decode_block_mappings` is the
second CARRIER of one grammar, never a second grammar: colon-SPACE separator,
unquote_scalar, duplicate keys refused, SPEC 5.2's actor rule applied. okf's
consume.read_sources was READ for the form and not called; po calls no okf
reader, which is measured and deliberate. After: 4605 present / 4605 entries.
Reading is not a licence to WRITE -- the emitter is untouched and both writers
still refuse what decode_flow_value refuses.

THREE FINDINGS. (1) The first block reader INVENTED data on `- { k: v }` items
-- SPEC-canonical, and the shape tests/golden/block-form-provenance writes for
`verified` -- decoding it as `{'{ id': '...'}`. No arm caught it: the 5.2 actor
rule shielded the fixture by accident. Closed with a flow-decoder branch and
four new arms. (2) One of my own arms was VACUOUS, found by my own mutation M5:
it claimed to prove the colon-SPACE rule and stayed green under first-colon,
because the two rules agree on every delivered value. Renamed, labelled, and
the claim moved to the arm that actually witnesses it. (3) OPEN, and it needs
the operator: the commons-owned worked example declares its second concept
`unreadable`/`block-sequence`, which is now false for po. `shared/` is
pull-only, so closing it needs a commons amendment; the test asserts the
divergence instead of skipping it, keeping the discriminating half (the example
says two entries were seen and the reader returns exactly two).

NINE EXISTING ARMS REWRITTEN, NONE WEAKENED. All nine pinned "the block form is
unreadable" -- the behaviour this order changes. Each keeps its claim on a
specimen that is still unreadable for a reason of its own (5.2: an entry naming
no actor), or pins the REVERSED direction where the old arm stood so the change
cannot be silent. Two got STRONGER: multi-verified.md was authored for "a reader
keeping the last entry reports machine-confirmed for a concept a human signed",
and that could not be tested while the form was unreadable. Three node ids were
renamed; nothing was removed in substance.

Suite 1582 -> 1606 passed / 5 skipped. Both demo goldens byte-unchanged
(ea8c534... / ede3e2f..., shasum -a 1 of the CONTENT, never the git blob id).
ruff check / ruff format / mypy green. shared/ untouched.

Six mutations, all red against the WHOLE suite, each with its own signature:
row 1 detached (2) / block reader detached (17) / flow-item branch detached (7)
/ a stray indented line folds into an INVENTED entry (4) / separator becomes the
first colon (1 -- and that is finding 2) / the stamp expectation reverts to
"true" (4).

Order: 20260912T195112Z-995611104-from-.claude
Record: docs/2026-09-12-p13b-okf-bump.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 23:02:38 +02:00
afdd9e0692 test(p13): the okf pin is MEASURED and deliberately NOT lifted -- the V1 stamp form disarms the forgery guard
The lift to llm-ingestion-okf v0.8.5 (which forces llm-ingestion-guard to
v1.4.0) was built and run in a worktree, never in the tracked tree. It is not
green, and the reason that decides it is not the red tests.

MEASURED. 27/27 imported names still resolve across five modules. Both demo
goldens stay byte-identical (ea8c534... / ede3e2f...), ruff check passes and
mypy clears 37 files. The suite goes 1579/3/5 (control) -> 1573/9/5. Eight of
the nine reds have ONE cause: the emitter moved from `generated: true` to the
V1 flow mapping `generated: { by: process:okf-ingest, at: ... }`, one line per
generated concept, seven files across four examples/ingest-golden-* bundles --
and 0 under shared/, so a future lift does not touch the pull-only subtree.

THE FINDING. okf._carries_complete_ingest_stamp reads the new form as NOT a
stamp (measured: True on the literal, False on the flow mapping), so
write_concept_file's IngestStampError refusal would land DISARMED -- and
test_ingest_stamp_fail_closed_loadbearing stayed GREEN through the whole bump
run. That is exactly the trap the _YAML_TRUE_LITERALS invariant row was written
for, arriving by a spelling it did not anticipate. A gate that stops gating is
not a row to name; it is a blocker.

TWO PREMISES FELLED before anything was built on them. Guard v1.2.0 -- the
lowest 1.x satisfying okf's declared >=1.2,<2.0 -- is NOT choosable: okf v0.8.5
pins the guard itself via [tool.uv.sources] tag = "v1.4.0" and uv refuses the
consumer's lower pin as conflicting URLs. And `rev = "v1.4.0"` is a DIFFERENT
url to uv than `tag = "v1.4.0"` even at the same value; only the tag= spelling
resolves.

DELIVERED. tests/test_okf_version_guard.py pins what is measured-green
(0.3.2 / 0.3.4) in two halves -- the installed distribution and pyproject --
with the refusal messages NAMING the eight-row cost of the lift, so the next
session cannot lift the pin without re-measuring. Iron Law: written red against
the v0.8.5/v1.2.0 target first (3 failed / 2 passed). Four mutations, each with
its own signature, all red: the okf assert never raises (1) / always raises (1)
/ the guard assert never raises (1) / the pin constant drifts to 0.3.3 (2 --
both halves, so the derivation is live and not two literals).

Also in the assessment: R761 navigated free for the first time (po had 0
references to it) -- 8.58/6.89/7.11 s, 131 MB max RSS, 5514 files -> 2756
concepts, 0 skipped links, against n100-2023's 0.21 s / 111 MB / 450 -> 446 /
0; and what the CLI can and cannot do with four bundles today.

Suite 1582 passed / 5 skipped (from 1577/5, strict superset, 0 removed).
Goldens byte-unchanged. Two coord messages closed; two STATE claims corrected
against measurement (upushed 3 -> 0, inbox "empty" -> 2).

Order: 20260912T190444Z-8080128610-from-.claude
Record: docs/2026-09-12-p13-okf-pin-r761.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 21:41:25 +02:00
c623498525 test(maf-guard): raise the floor to 1.18.0 first and prove it red on the installed 1.16.0
Iron Law, F15's procedure repeated: the guard moves BEFORE the pin, and the
move is measured (2 failed / 2 passed against the installed 1.16.0) rather
than asserted. The floor still lives in ONE constant and the pyproject assert
still DERIVES its expected string from it; a second literal would reintroduce
the drift F15 removed.

The stale arm gains 1.16.0 and 1.17.0, and its docstring is corrected rather
than left behind: 1.17.0 is rejected even though agent-framework-foundry
1.13.0 and agent-framework-openai 1.14.3 only require core>=1.17.0 -- a
dependency's own floor is not our floor.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 16:57:48 +02:00
ce22b0e0e6 docs(p11): okf 0.8.1 measured - source-quota carries the denominators, one prior constant ranked the price schedule out [skip-docs]
P11 (order 20260910T225652Z), NOK 0, no push. PATH okf is 0.8.1 (uv tool list,
__version__, --help flags). src/ untouched.

(ix) on 0.8.1 shipped: 629/623/6 and 629/620/9 - PM's predicted denominators,
now measured. --no-source-quota alone gives back 629/621/8 and 629/617/12 in
both arms; the exact numbers are an interaction with tie-shared-rank (both
arms) and stem-prefix (open arm). --title-covered never fires here (byte-
identical payload). 0.8.1 with the three new rules off reproduces the 0.7.0
payloads' delivered lists and budgets.

P10 section 6 diagnosed: a bisect over okf's own history (known-positive at
both ends) puts the price schedule's loss at okf 38104b7, whose only consume
code change is DOCUMENT_PRIOR_EXPONENT 1.0 -> 0.5; putting that one constant
back in a copy returns the old payload byte for byte. tie-shared-rank is ruled
out; known_positive 10349 -> 12563 is a version marker, not the mechanism. On
0.8.1 the schedule is over_budget_after_knapsack, not below_k.

okf check on 0.8.1 has 15 rules, not 16: rule 16 (bundle_mismatch) is on okf
main 7cca9e0, in no tag. Run from an export of 7cca9e0 the K3-15 pair is rc 1
and the right pair rc 0 - a real cross-corpus mismatch, not an okf defect.

Offer rows on four corpora: identifiers unchanged (65/50, 435, 982, 272).

The P10 doc gets dated additions in sections 4 and 6; the old sentences stand.
Test docstring: +2 lines naming the 0.8.1 rule count, no behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 01:38:58 +02:00
3da28c382e test(prepass): a delivered excerpt nobody named must be named, not silent [skip-docs]
RED. P9 measured `okf check` (0.7.0, 15 rules) rc 1 on both K2 payloads with 20
findings, every one of them `excerpt_unnamed`. On po's side the same absence is
SILENT: `_excerpt_header` appends `title:` only `if excerpt.title is not None`.

F1 (reporting) over F2 (refusing), and F2's price is a NUMBER: of the 20 payload
files in this repository 10 carry a `title` on every excerpt and 10 on none, and
among the ten F2 would refuse is the ONE git-tracked payload fixture
(tests/fixtures/prepass/bygg-energi-mikro-fixture.payload.json, 4 excerpts,
0 titled). F2 would need a tracked fixture rewritten to make its own rule green,
and would reverse PrepassExcerpt's written "a required field would refuse every
payload written before today".

6 failed, 1 passed. The one green is `test_no_existing_payload_is_refused` --
F1's promise, true before the rule and required to stay true through it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 07:23:27 +02:00
3082e704d2 docs(p9): okf v0.7.0 measured - the rename costs nothing, content adds 15 identifiers [skip-docs]
P8's grounding-offer measurement repeated on the okf v0.7.0 bundles, free, NOK 0.

N100/N200/N500 reproduce P8 EXACTLY (446/1133/270 concepts, 462041/1500962/408220
chars, 435/982/272 identifiers, 0 cost lines, derive REFUSED) - as predicted, because
vegnormal never went through okf's pandoc converter. K2 on the pinned v0.7.0 build
(453 concepts / 865 md) offers 65 identifiers against P8's 50: no row got worse
(0 lost, 15 new, derive REFUSED on both builds with a byte-identical message).

The cause is MEASURED, not guessed: CONTENT, not the concept rename. Concept NAMES
contribute 0 of 50 identifiers in the old build and 0 of 65 in the new one, because
the identifier forms require an uppercase head and a concept slug is lowercase - so
the rename cannot move the count at all. The 15 new tokens come from three documents
present under the SAME name in both builds with different bodies (snitt-e.md
28 -> 11029 chars, generell-orientering.md 1662 -> 47503); 102 of 414 shared concept
names differ in body length.

okf check (0.7.0, --skill/--payload): 15 rules; payload-n100 rc 0 with 0 findings;
the two K2 payloads rc 1 with 12 and 8 findings, all excerpt_unnamed - payload AGE
(recorded before the title field existed), not a v0.7.0 regression; known-negative
{} rc 1 with 9 findings, reproducing V3's figure, so the check can fail.

Two docstrings corrected to the new concept id, each naming the build its numbers
were measured on. src/ carries 0 occurrences of the old form. The recorder-test
string label and the two *-SYNTETISK fixtures are left untouched with the reason
stated: the label is arbitrary and looked up nowhere, and the fixtures simulate
converter output as it was.

No new seam, no new function, no contract change against okf. _ground_against_input
is untouched. No paid run.

Suite 1570 passed / 5 skipped, golden demo-transcript.stdout content sha1 unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 06:33:57 +02:00
2e04c00d6e style(format): ruff format on the three files P7/P8 left drifting [skip-docs]
Pure formatting, no behaviour change. ruff 0.15.18 reported
"3 files would be reformatted, 195 files already formatted" on HEAD 455d611;
after this commit "198 files already formatted". Suite 1570 passed / 5 skipped
before and after, golden demo-transcript.stdout content sha1 unchanged
(ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Two of the three came from P7 (277bb95), one from P8 (455d611).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 06:23:35 +02:00
455d611660 feat(run,generate): a run says what its delivered input can ground, before it spends an attempt [skip-docs]
P7 is right and landed, but re-measuring it exposed a consequence no row stated: with the gate
live, 29 of 29 cost codes in 13 of 13 delivered proposals fall across the three free recordings
(PM's denominator; 14 of 14 in 8 proposals on the PARSEABLE one -- the five blobs that separate
the numbers are refused by pydantic's `claimed <= total` and never reach stage 0b). All 29 were
invented, so the gate is right; but a gate that always refuses is as useless as one that never
does.

The cause is that the PROMPT asks for something the input cannot supply. `_build_messages`
requires each affected_item to "restate a cost line as the project's price schedule already
carries it", while K2's delivered input carries 9 occurrences / 2 distinct code-shaped tokens --
`SHA-01`/`SHA-10`, both document numbers off a page footer -- and `derive_cost_baseline` refuses
the base outright. There is no cost line in it to restate.

`GroundingOffer(chars, identifiers, cost_lines)` reports it. The PAIR is the diagnosis: "50
identifiers, 0 cost lines" says what neither number says alone. A REPORT, never a gate -- it
blocks nothing, because a blocking requirement IS `--require-cost-baseline` (F4/D-3, opt-in,
untouched), and `_ground_against_input` is untouched.

The callsite is MEASURED, not chosen: `generate.py` composes the grounding per attempt, after
`await _fetch_parsed`, so a report there could only speak once an attempt had been paid for;
`run.py` binds both halves above the `--live-dry-run` cut and before the first `debate.run`, so
the FREE trip says it. `delivered` is bound ONCE and the same variable feeds the report and
`_evaluate`; the report composes THROUGH `_grounding_text`, the gate's own composer.

A pattern is admissible here and not in the gate, and that is the difference between a report and
a falsifier: an unknown form is a token left uncounted -- an under-count, never a false rejection.
The forms are transcribed from the measurement; bare numbers are excluded with the number
(46 394 / 2 117 in K2). `grounding_offer_notice` is the ONE renderer and is silent when the run
CAN anchor -- omission, never an empty row.

Load-bearing MEASURED (tests/test_grounding_offer_loadbearing.py, 12 arms), nine mutations all
red against the WHOLE suite + green control 1570/5 (from 1558/5, strict superset, 0 removed) and
the golden byte-unchanged (shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Measurement: docs/2026-09-09-p8-forankringstilbudet.md

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 17:50:02 +02:00
277bb95777 feat(validator,generate,run): an identifier a proposal builds on must be in the input, or the verdict falls [skip-docs]
P6 (økt 108) ended in ValidatedProposal (verdict 5fd6272e3725fe68) on two cost
codes -- M-04-01 / M-04-03 -- that appear in NO prompt of that run. Measured
here first, verbatim: validate_proposal(p, baseline=None) validates it; the same
proposal against any non-empty CostBaseline is rejected naming both codes.

So the hole was never "fabrication goes uncaught" -- _reconcile_against_baseline
exists and is right -- but that the falsifier is reached only through
`if baseline is not None`. The input always exists; the baseline does not.

New stage 0b (_ground_against_input), OUTSIDE the baseline branch, after stage 0
so an anchored run's message is byte-identical to before. ONE Rejection, the
validator's own type, naming EVERY ungrounded identifier "; "-joined in the
proposal's own order (økt 94's completeness reason).

The rule has NO pattern -- `code in grounding`, exact substring -- and that is a
measurement: over the delivered corpora (K2 1108 files / 2 005 561 chars, the
three N payloads 8 excerpts each) the identifier forms are heterogeneous, and a
pattern chosen to cover them would be a rule about shapes. Bare numerals are the
one inert class (46 394 occurrences / 2 117 distinct in K2); the rule fails OPEN
there, never closed.

Evidence is three non-model-authored sources: what run_project DELIVERED (the
rendered cut/pointer/chunks plus the base's context_files -- never files, which
would make the type: verdict layer evidence), the project's own cost lines, and
the baseline's codes when anchored. The rendered PROMPT is deliberately NOT
evidence, on two measurements: gen_context IS the debate output on the S2c path,
and from attempt 2 the prompt carries the previous Rejection.reason verbatim --
which for this stage QUOTES the identifier it just refused. Grounding in the
prompt would let the gate's own refusal disarm it on its second round.

Prose scanning was chosen against WITH THE NUMBERS: a typed gate catches 2/2
(P6) and 2/2 (S7c) -- 100% of what reached a verdict. What stays uncaught, said
plainly: an ungrounded identifier that lives only in agent/debate prose and never
becomes an affected_item code (2 of 4 P6, 2 of 4 S7c, 1 of 2 P4).

Iron Law: 9 red / 2 green before the rule existed. Ten mutations all red against
the whole suite, green control 1558 passed / 5 skipped (from 1543/5, superset,
0 removed), golden demo-transcript.stdout BYTE-UNCHANGED (shasum -a 1 of the
CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Three existing fixtures changed, no gate weakened -- most of all
test_pre_amendment_bundle_runs_unchanged, which sent the SAME FABRICATED code and
asserted it validated: the økt-108 hole written down as an expectation.

No paid run. Order 20260909T113641Z-38938691-from-.claude.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 16:04:01 +02:00
eb4137415c feat(prepass,explore): padding dies at the prompt; a refusal the model can act on is a return value [skip-docs]
Order 20260908T195801Z. Findings 4 and 5 from the S7 acid test, then the two things
finding 99 measured and deliberately did not fix (D3, D2).

No user-facing surface changes: no new flag, no new command, no changed output
contract. Both seams are internal (the pre-pass rendering, and the shape a tool
answers a model with), so [skip-docs] rather than a README edit that would describe
nothing an operator can do differently.

FINDING 4 -- MEASURED, NOTHING BUILT. K2's price schedule IS readable without
guessing (8 column spans, 71 of 91 non-blank rows give >= 2 cells, the split stable
for K = 2..64). But 0 of 92 rows name all three of code/quantity/unit_cost -- also
under a looser substring match -- and 0 of 91 data rows carry code + quantity +
amount. The triple is not formatted away; it is not in the document. It is a price
SUMMARY plus nine rate cards whose unit-price columns are empty (pre-award). The
order's binding decision rule therefore falls against building:
--derive-cost-baseline keeps refusing, and MAJOR-4's own honesty limit holds.

FINDING 5 -- BUILT. Measured on the actual rendering path (concept_text, not the
raw file): the delivered excerpt is 104 lines / 67 245 chars, carrying 208 interior
whitespace runs, 117 of them >= 100 and the longest 887 -- 56 806 of 67 245
characters = 84.5 %, over 72 of 104 lines. collapse_padding, called from
_data_blocks (the one renderer both arms share, and therefore AFTER
verify_against_bundle -- collapsing in concept_text would break every payload's own
digest), gives -72.4 %: line count invariant, non-whitespace byte-identical, leading
indentation untouched, no number changed.

F99-D3 -- read_file / read_dir / read_bundle now RETURN their refusal. MAF turns a
tool raise into "Error: Function failed." (_tools.py:1410-1432, :1427) and counts it
against DEFAULT_MAX_CONSECUTIVE_ERRORS_PER_REQUEST = 3, so everything the refusing
arm knows is destroyed on the way out. The gates are unchanged; the property they
exist for -- the reason travels, the bytes never do -- is now asserted explicitly on
the returned value. The arm is keyed on named classes, never bare Exception, because
ExplorationError is itself a RuntimeError subclass.

F99-D2 -- the invariant row, plus one for finding 5 (a stated deviation from "one
row only": finding 5 is a separately built seam and the ledger's standing rule
requires its own row).

19 existing arms rewritten, never deleted and never weakened: where the class
carried a distinction, the refusal KIND carries it now.

13 mutations, all red against the whole suite (W1-W5, M1-M8), each restored from
scratchpad with shasum -c. Control 1543 passed / 5 skipped (from 1529/5, a strict
superset, 0 removed). Golden demo-transcript.stdout unchanged
(shasum -a 1 of the CONTENT = ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Measurement: docs/2026-09-08-funn-4-5-og-read-nekt.md

Co-Authored-By: Claude <Opus 5>
2026-09-08 23:36:54 +02:00
078a099898 fix(cli): a provider failure leaves the CLI as one line, and a guessed base id is correctable
Funn 99, measured offline against the artefacts the paid Q5=B run left behind — no paid
run here.

ROOT, verbatim from the records: the three failing quick_validate calls all sent
bundle_id="renholdstekniske_funksjonskrav" — a CONCEPT name guessed out of the seeded cut,
while the base's id is k2-trinn1-20260903. Both arguments parsed against the signature, so
it was _resolve_bundle's raise MAF counted, proven by quick_validations being EMPTY while
all three stand in tool_calls. Denominator: 12 tool calls, and those three came BEFORE
list_bundles.

The order's causal chain is FELLED: the quick_validate triple is records 4-6 and the run
continued for 13 more model calls; the triple immediately before the 400 is the navigator's
three read_file refusals on del-ii-bilag-7-prisskjema*. The limit fired TWICE.

(A) ChatClientException is caught on BOTH seams — the exploration dispatch and the full-run
dispatch — because the debate's own model calls go through the same provider. The line is
"run stopped:", not "run refused:" (a stated divergence from the order): the argv was fine
and tokens were already spent, which is the MAJOR-2 arm's own reason, verbatim. Caught
INSIDE the try/finally so the exploration artefact still lands.

(B) quick_validate answers an unknown base id with {"decision": "refused", ...} naming the
configured ids, and records it in the sink. MAF turns a tool raise into the opaque
"Error: Function failed." (_tools.py:1426), so the one thing the refusal knew and the model
did not never reached it — the replies show it guessing at the JSON format instead.
read_file/read_dir/read_bundle still raise: measured, reported, out of scope.

Seven mutations all red against the whole suite, green control 1529/5, golden ea8c534
unchanged. One existing gate REWRITTEN, not deleted; its second half is what keeps (B)
scoped. The test double raises from the reply_selector seam rather than a new
_inner_get_response body, so the S2.5 consolidation guard stays untouched.

Co-Authored-By: Claude <claude-opus-5>
2026-09-08 21:03:18 +02:00
2453246d4b feat(prepass): the excerpt's new fields reach the prompt, po stops dropping them twice
P2 measured that the producer's payload now carries title/req_number/sources/source_* on
every excerpt (14 members, was 9), but PrepassExcerpt ignored them (extra="ignore") and
_data_blocks rendered only concept_id/adjudication/trust_tier -- so (b') was a po verdict,
never a model verdict. title/req_number/sources are now named fields; source_* locators are
read via model_extra and a prefix scan (measured: the producer treats source_* as an
open-ended family, not a fixed allowlist), so a future producer's new source_foo key reaches
the prompt without a code change here. A P1-form payload renders byte-identical to before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 16:42:17 +02:00
9232f94041 feat(navigation,validator): read_dir names the rung that reads a document, and a run can require its anchoring
F3 and F4, the two findings the S7 acid test (session 98) reported and left. The order required
both descriptions to be treated as PREMISES. One held; the other was felled before anything was
built on it.

F3 -- premise FELLED, asymmetry real. The order read arm C's two refused calls as "the path names
a document that EXISTS". Measured against the base that ran: its root holds 27 directories named
del-ii-bilag-N-... and 12 documents named inbox-del-ii-bilag-N-....md, and the requested path
matches NEITHER -- it is the directory naming convention applied to a document whose real name
carries an inbox- prefix. So the two live rounds were the UNKNOWN-path class, and this delivery
does NOT recover them (gated). What IS real: read_file on a directory has named read_dir since
session 95, while read_dir on a document named neither the rung nor the path.
okf.DocumentPathRefused closes that one direction -- a ValueError, a SIBLING of BundlePathNotFound
rather than a subclass, built from context_files (never files) and through the same in_dimension
predicate the listing uses, quoting the document's REAL name so what it hands back resolves.

F4 -- premise HELD, option (c) felled by measurement. All four live artefacts stamped
cost_baseline_anchored: False and each arm invented its cost codes. derive_cost_baseline refuses
against the delivered base: K2's price schedule is a pandoc SIMPLE table with ONE column header,
so making --derive-cost-baseline reachable there would mean inventing a rule for an unmeasured
form -- MAJOR-4's own honesty limit. Chose (b) over (a): --require-cost-baseline /
run_project(require_cost_baseline=...), OPT-IN and never default, so every bundle without a
cost-baseline.json runs unchanged. The gate sits where both branches have bound baseline and ABOVE
the dry-run cut, so it fires on the free trip too and, on the paid one, before the first model
call. Three CLI refusals by name, each with an rc-0 control.

12 mutations, all red against the WHOLE suite. Green control 1493/5 -> 1511/5 (+18 node-ids, 0
removed); golden demo-transcript.stdout BYTE-UNCHANGED (shasum -a 1 of the CONTENT =
ea8c534773acdbe41ae68f2c55724d69aaf8be4f). No paid run: both findings measured offline.

Measurement: docs/2026-09-08-f3-f4-nekten-og-forankringen.md
Order: 20260908T020419Z-5837110336-from-portfolio-optimiser

Co-Authored-By: Claude <Opus 5>
2026-09-08 05:32:46 +02:00
76b939b3b8 feat(prepass): --prepass-seed makes the cut a starting point, and K2 says it costs
Q5 = B, bygget som MAALT OPSJON. --prepass-payload gir DEBATTEN et deklarert
kutt og trekker de fire navigatoerverktoeyene; --prepass-seed gir UTFORSKNINGEN
det samme kuttet som utgangspunkt og BEHOLDER verktoeyene.

Nekten M32/F4 staar ORDRETT. B er et nytt flagg, aldri en loesning av den, og
hjemmelen er konsumkontraktens SS 2.2: en skill maa ikke lese «outside what the
payload delivers or explicitly names as reachable». Andre ledd er hele arm B, og
PrepassDeclaration.rest_reachable er det som gjoer de to lesningene skillbare i
ettertid -- paakrevd uten default av cost_baseline_anchoreds grunn, fordi begge
defaults ville loeyet om hvilken arm som leste kuttet.

Soemmene:
  admit_payload er EN opptaks-gate (form -> montert base -> tom-leveranse-nekt)
    delt av begge doerer; to kopier ville latt en doer slippe inn det den andre
    nekter.
  render_seed deler header, regel->ANTALL-foldingen og DATA-blokkene med
    render_context. Det eneste som skiller dem er avsnittet som sier hva
    leseren kan gjoere videre.
  explore(seed_context=...) legger kuttet i TASK-MELDINGEN, aldri i prompt:
    _finish bygger Mandate.objective av prompt, og en kommisjon med 22 335
    tokens utdrag i objektivet er uleselig for den som skrev den. Tom streng gir
    en byte-identisk task-melding.
  trace_payload(prepass=...) skriver deklarasjonen fra en finally. MAALT baerende
    -- den seedede kjoeringen som doede paa en Azure-400 etterlot likevel kuttet
    deklarert.
  Fem nekter ved navn. --checkpoint-dir baerer en beslutning: en gjenopptatt
    etappe kjoerer i en prosess som aldri saa payloaden og ville overskrevet den
    parkerte etappens deklarasjon med prepass: null.
  --dimension-config er BEVISST ikke nektet (arm A nekter den): maalt bygger
    utforskningen navigator_tools(bundle_dirs) UTEN dimensjon, saa aa skope
    seedet ville nektet tekst den samme loekka kan aapne et oeyeblikk senere.

MAALT PAA K2 MED LEVENDE MODELL, og maalingen taler MOT aa gjoere B til default:
like-for-like gratis 4 317 -> 227 675 o200k (x 52,7), og betalt er manageren
x 21 paa samme antall prompter. Viktigere enn prisen: den USEEDETE kontrollen
hentet prisskjemaet i fire steg (del-ii-bilag-7-prisskjema/prissammenstilling-
sheet-1.md), mens BEGGE seedede armer lot vaere -- den ene med null verktoeykall
fordi manageren rutet til hypotesisereren i alle tre runder, den andre ved aa
gjette stier ut av kuttets egne konsept-navn og mynte en base-id som ikke finnes.
Erkjennelsen kom (manageren skrev i hver runde at utdragene ikke rakk),
handlingen ikke. Ingen av de 40 svarene brukte ett eneste av kontraktens fem
literaler. NOK 2,78 av taket 5, 0 x 429. Anbefaling skrevet, beslutning ikke
tatt -- den er operatoerens.

Load-bearing maalt: 26 armer, 16 mutasjoner alle roede mot HELE suiten, groenn
kontroll 1493 passed / 5 skipped (fra 1467/5; +26 node-ider, 0 fjernet), golden
demo-transcript.stdout byte-uendret (shasum -a 1 av innholdet =
ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

To armer var GROENNE AV FEIL GRUNN og ble rettet, ikke droppet: tool_calls alene
kan ikke skille «et verktoey ble kalt» fra «basen var aapen», fordi recorderen
appender FOER call_next; og id-enighets-armen maalte den stale digesten i stedet
for id-gaten. Sonden var dessuten feil foer koden var det -- foerste
diskriminator var norsk, og verktoeysvar serialiseres med \uXXXX-escapes.

Avvik, uttalt: implementasjonen ble skrevet FOER testfila. Roedmaalingen er gjort
etterpaa ved aa reversere src/ til HEAD (16 av 24 armer roede), deretter
restaurert med shasum-verifikasjon. Beviset er ekte, rekkefoelgen var ikke.

Rapportert, ikke fikset: ChatClientException (Azure 400, «No tool call found for
function call output») etter tre quick_validate-nekter paa rad -- den ligger
utenfor main()s nekt-tuppel og forlater CLI-en som traceback.

Azure-konfigurasjonen er uendret; endepunktet utledes inline fra az og er aldri
lagret i fil. Ruff + mypy rene.

Maaling: docs/2026-09-07-prepass-mater-q5b-k2.md
Ordre: 20260907T234344Z-9062321009-from-.claude

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 04:02:09 +02:00
183191e51c docs(prepass): before/after on K2, the mutation battery and the invariant row
Steg 8. Ingen produksjonskode i denne commiten utover de TRE testarmene tre GROENNE
mutasjoner tvang fram.

MAALING (`docs/2026-09-07-okf-prepass-i-debatten.md`), instrumentet validert mot en
kjent-positiv FOER foerste tall (K2 rotlisting = 3 954 tegn / 1 495 o200k-tokens,
S7a-3s publiserte tall eksakt; skriptet asserterer paa det og nekter aa rapportere
ellers):
- KUTTET paa levert K2: 629 = 621 + 8, tre regler navngitt. Payload paa disk 40 425
  tok; renderingen debatten ser 5 162 tok -- 87 % av kostnaden er withheld-lista, som
  aldri naar prompten.
- FOER/ETTER, tre armer fordi "FOER" ikke er ett tall. Like-for-like (en debatt som
  GAAR stigen mot det deklarerte kuttet): 10 prompter / 7 031 tok -> 3 prompter /
  10 641 tok. **+51 %, og dokumentet paastaar ikke at dette er en besparelse.** Det som
  kjoepes er nevnerne, et re-maalbart `ref`, og at stempelet slutter aa overdrive:
  siteringer 629 -> 8.
- SPOERSMAALET BASEN IKKE SVARER PAA: aatte arkitekttegninger, og **5,5x dyrere** enn
  det gode spoersmaalet (28 312 mot 5 162 tok). En TOM leveranse er bevis for fravaer;
  en FULL er IKKE bevis for tilstedevaerelse -- og med verktoeyene trukket har debatten
  ingen vei til aa oppdage det selv. Uttalt som den reelle handelen, ikke oppdaget i
  drift.

LOAD-BEARING: 32 mutasjoner, ALLE ROEDE mot HELE suiten, maks en per Bash-kall,
restaurert fra `scratchpad/` med `shasum -c` (aldri `git checkout`). Groenn kontroll
**1467 passed / 5 skipped** (fra 1387/5; **+80 node-ider, 0 fjernet**, maalt med `comm`
mot en baseline tatt foer foerste commit). Golden `shasum -a 1` av INNHOLDET =
ea8c534773acdbe41ae68f2c55724d69aaf8be4f, BYTE-UENDRET. `pyproject.toml` uroert.

TRE MUTASJONER VAR GROENNE FOERST, og alle tre var TESTFEIL -- ikke soemfeil. Ingen ble
droppet:
- M9: injeksjonsarmen satte `text_sha256` fra den INJISERTE teksten, saa
  `text_sha256`-grenen fyrte i stedet, og armen matchet paa "text" -- en DELSTRENG av
  `text_sha256`. En angriper kontrollerer begge avledede medlemmer, saa fixturen setter
  naa `text_sha256` til digesten av den EKTE teksten; da er likhetssjekken det eneste
  som staar i veien. Ny parallell arm beviser at de to sjekkene ikke er en sjekk.
- M18: fixturbasen navigerer til fire ikke-verdict-konsepter og payloadet leverer ALLE
  fire, saa "siter de leverte" og "siter alt navigerbart" ga SAMME sett. Armen flytter
  naa ett utdrag til `withheld` og asserterer `delivered < navigable` FOER den maaler.
- M32: uten raden falt kjoeringen gjennom til "--explore requires --explore-config",
  som ogsaa navngir `--explore`. Oekt 57s regel ordrett -- assert aldri paa en
  delstreng to nekter deler. Armen bruker naa en argv `--explore` ellers ville blitt
  AKSEPTERT paa.

Co-Authored-By: Claude <claude-opus-5>
2026-09-07 14:39:50 +02:00
ca98888358 feat(prepass): --prepass-payload with six refusals by name, README and the hosted surface untouched
Steg 6 og 7 av planen. Flagget lastes fail-fast ved siden av `--mandate` (samme
try/except, saa manglende/ugyldig fil lander paa `run refused:` uten traceback) og
traades inn i BEGGE `run_project`-dispatchene -- dry-run og full kjoering. En egen arm
SPIONERER paa argumentet, ikke paa exit-koden: et flagg som parses, valideres og
droppes er F4-klassen, og de to utfallene er samme rc.

SEKS NEKTER, hver ved NAVN, hver med en rc-0-kontroll paa en argv som ellers ville
blitt AKSEPTERT:
- `report_forbidden` -- report-modus returnerer OVER hver dispatch, saa en utelatelse
  er et stille DROPP. Kontrollen bruker en JSON-ARRAY-ledger; et objekt ville gjort
  armen roed av feil grunn (maalt i oekt 89).
- `single_only` -- navngir `--portfolio`, ALDRI det delte `--prepass-payload`-tokenet:
  en droppet rad faller gjennom til `--bundle-dir`-kravet, hvis melding ogsaa navngir
  flagget, saa en arm paa det delte tokenet ville staatt groenn mot sin egen mutasjon.
- krever `--bundle-dir`; nektet med `--proposals-from-mandate` (returnerer over
  debatten, saa flagget ville vaert stille inert), med `--dimension-config` (pre-passet
  kuttet uten aa kjenne dimensjoner, saa aa aere skopet ville droppe utdrag
  deklarasjonen teller som LEVERT -- da er nevnerne feil for kjoeringen som publiserte
  dem) og med `--explore` (utforskningen leser HELE basen med de fire verktoeyene
  payloadet trekker, saa kjoeringen som helhet ville lest langt utenfor kuttet den
  erklaerer).

Blokka ligger paa FUNKSJONS-nivaa etter mode-dispatchen, aldri nestet under en annen
grens -- under en av dem ville en bar kombinasjon falt rett gjennom.

`hosting.py` er BEVISST URØRT (briefens non-goal, MAJOR-4/S7b-presedensen): feltet
kommer inn i ingen av de tre settene, saa den generiske `unknown field(s)`-400-en
svarer alt, og Fase 4es to halvdeler staar. Gatet av en testarm i stedet for en
redigering -- inkludert den negative halvdelen (hvert videresendt felt ER en
`run_project`-parameter, hvert konsumert er det ikke).

README-blokka navngir alle seks partnerne, uttrykker seg i kundevendt terminologi
(aldri "OKF bundle") og sier BEGGE aerlighets-grensene hoeyt: dette kjoeper et
DEKLARERT kutt, ikke en billigere kjoering; og en TOM leveranse er bevis for fravaer
mens en FULL ikke er bevis for tilstedevaerelse. Uttrekkeren tar BLOKKA (ikke en
delstreng over hele fila -- `--portfolio` og `--report` staar overalt), med
`--plan-review`-blokka som kjent-positiv kontroll.

1466 passed / 5 skipped (fra 1446/5, +20, 0 fjernet). ruff + mypy rene. Golden
`shasum -a 1` av INNHOLDET = ea8c534773acdbe41ae68f2c55724d69aaf8be4f, BYTE-UENDRET.

Co-Authored-By: Claude <claude-opus-5>
2026-09-07 11:27:35 +02:00
d3476d55ee feat(prepass): one renderer and one artefact carry the declaration out of the run [skip-docs]
[skip-docs]: CLI-flagget og README-blokka kommer i neste commit.

`prepass_notice` er ENESTE renderer, tar den ALT OPPLOESTE deklarasjonen (aldri en
payload-sti -- en renderer som leste fila paa nytt ville vaert en andre oppløsning fri
til aa vaere uenig med kjoeringen den beskriver), og returnerer `None` uten payload.
Omisjonen er entydig her paa en maate `proposal_review_notice`s bevisst ikke er: en
kjoering sier ingenting om et kutt fordi operatoeren ikke ga noe, og det finnes
noeyaktig EN maate aa gi et paa. Golden-transkriptet er det uavhengige, eksisterende
vitnet for den halvdelen. To kallsteder: dry-run og full kjoering.

`outbox.write_prepass` skriver `{run_id}-prepass.json` IFF et payload ble gitt --
`write_proposal_reviews`-regelen, og her gjoer den en andre jobb: siden et payload
TREKKER navigatoerverktoeyene er `{run_id}-debate.json`s `tool_calls` tom ved
konstruksjon paa denne stien, og DEN fila skrives ubetinget nettopp fordi et tomt spor
ER S2c-regresjonen. Uten dette artefaktet ved siden av ville "trukket med vilje" og
"regrert" lest likt. TILSTEDEVAERELSEN er det som skiller dem, og en egen arm
observerer BEGGE filene paa SAMME kjoering.

Skrevet fra `finally` (`write_parse_failures`-presedensen) ved et DIREKTE kall, ikke
via `_write_or_report`: den helperens paakrevde `in_flight` bindes foerst inne i
genererings-blokka, og `None` der ville latt en `OSError` fortrenge en `BudgetExceeded`
i luften -- noeyaktig defekten helperen finnes for. En egen arm driver et budsjettstopp
midt i debatten og krever at deklarasjonen likevel ligger der.

Regel -> ANTALL i baade rendereren og artefaktet; en arm beviser at 20 tilbakeholdte
konsept-ider naar HVERKEN stdout eller artefaktet mens antallet gjoer det.

1446 passed / 5 skipped (fra 1440/5, +6, 0 fjernet). ruff + mypy rene. Golden
`shasum -a 1` av INNHOLDET = ea8c534773acdbe41ae68f2c55724d69aaf8be4f, BYTE-UENDRET.

Co-Authored-By: Claude <claude-opus-5>
2026-09-07 11:17:38 +02:00
7aa06d581f feat(prepass): the debate is handed the declared cut and the navigator tools are withdrawn [skip-docs]
[skip-docs]: CLI-flagget og README-blokka kommer i steg 6/7; `prepass_payload` er
foreloepig bare naabar for en bibliotek-kaller.

`run_project(prepass_payload=...)` forgrener bundle-armen. UTEN payload er hver linje
uendret -- pekeren, de fire verktoeyene, siteringer over hele den navigerte basen. MED
et payload faar debatten et DEKLARERT KUTT og verktoeyene trekkes (SS 2.2: kontekst
pre-passet holdt tilbake ble holdt tilbake med vilje; en debatt som holder BEGGE er fri
til aa gaa rundt kuttet den nettopp erklaerte).

Nekten PROPAGERER, aldri en stille degradering tilbake til pekeren -- `load_mandate`s
regel. Maalt paa NULL modellkall, ikke paa exit-koden.

`delivered == 0` nektes ved NAVN foer debatten, med nevnerne, spoersmaalet og ref-en
sitert: maalt er den tilstanden bare naabar naar hvert konsept feilet leksikalsk (den
andre tomme saken nekter produsenten selv), altsaa bevis for FRAVAER. Uten den falt
kjoeringen gjennom til `run.py`s siteringsvakt, hvis melding navngir `docs_dir` -- som
er `None` paa denne stien.

`bundle_excerpt_citations` (datasource) siterer de LEVERTE konseptene alene: et stempel
som siterer hele korpuset for et forslag som saa fire, gjenoppfinner den uerklaerte
paastanden sømmen finnes for. Kroppen tas fra den NAVIGERTE `BundleFile`, ikke fra
payloadets `text` -- den leverte teksten er NFC-normalisert med strippet hale, saa en
locator over den ville ikke indeksert fila den navngir. Deler dermed ogsaa
`bundle_citations`' verdict-eksklusjon i stedet for aa gjenta den.

MCP-appenden ligger BEVISST under forgreningen: dette trekker navigatoerverktoeyene,
ikke verktoeylista. Maalt: en tom liste naar traaden som `tools: None`, saa ingen
uproevd tom-array-form innfoeres.

`RunResult.prepass` og `DryRunReport.prepass` DEFAULTER (`skipped_links`-halvdelen:
`None` er det sanne utsagnet "ingen payload ble gitt"), bundet i BEGGE grener saa ingen
`NameError` venter paa veg-stien. `ProvenanceStamp` er BEVISST urørt -- stempelet
beskriver gaten som doemte EN kandidat, dette er et RUN-nivaa-faktum om hva kjoeringen
i det hele tatt fikk lese.

Tilbaketrekkingen asserteres paa `fresh_workflow(tools=...)`, ALDRI paa
`debate_tool_calls`: maalt er det sporet allerede tomt MED alle fire verktoeyene, fordi
en `ScriptedChatClient` aldri emitterer et verktoeykall. En arm skrevet paa det kan
ikke skille de to implementasjonene. GATE, IKKE VEGG: en egen arm beviser at den gatede
ExpeL-folden fortsatt naar hypotese-prompten (0.82) under et payload.

1440 passed / 5 skipped (fra 1425/5, +15, 0 fjernet). ruff + mypy rene. Golden
`shasum -a 1` av INNHOLDET = ea8c534773acdbe41ae68f2c55724d69aaf8be4f, BYTE-UENDRET.

Co-Authored-By: Claude <claude-opus-5>
2026-09-07 11:10:54 +02:00
651c83f9df feat(prepass): render the cut for the prompt and declare its denominators [skip-docs]
[skip-docs]: fortsatt ingen brukervendt flate -- CLI-flagget kommer i steg 6/7.

`render_context` er hva debatten faar I STEDET FOR pekeren: de leverte utdragene, de
tre nevnerne, spoersmaalet kuttet ble regnet for, og de tilbakeholdte konseptene som
REGEL -> ANTALL. Ids-ene baeres aldri -- maalt paa et 629-konsepts korpus er
withheld-lista alene 34 451 o200k-tokens, og et regelnavn er faktumet en leser kan
handle paa mens en liste over ids de ikke kan aapne er kostnad uten informasjon.

Avslutningslinja er ikke pynt: maalt LIVE gir et spoersmaal basen ikke svarer paa
fortsatt aatte utdrag, og med verktoeyene trukket har debatten ingen vei til aa
oppdage det selv. Renderingen sier derfor at et levert utdrag er et LEKSIKALSK treff
og ikke et svar, og navngir `[sourced-not-sufficient]` fra SS 4s markoer-sett.

`adjudication`/`trust_tier` merkes som PRODUSENTENS erklaering, ikke vaar: B4 slo fast
at en tillitsgrad utledes lokalt fra dokumentets eget `verified`, og repoets egne baser
baerer ingen -- saa en lokal utledning ville rapportert `unverified` for alt og sagt
ingenting. Aa adoptere produsentens verdi i stillhet ville vaert det gale svaret.

TAKET BINDER OVERHEADET, IKKE TOTALEN, og det er hele poenget: den leverte teksten er
alt bundet av produsentens eget budsjett (SS 7.3), saa et absolutt tak ville enten
nektet et legitimt stort kutt eller vaert saa loest at det fanget ingenting. Det po maa
garantere er KOSTNADSFORMEN -- O(levert tekst) + O(1), aldri O(korpus) -- som er
noeyaktig det som regrerer hvis withheld-lista sniker seg tilbake. Taket bor i TESTEN
(`_CATALOGUE_EXCERPT_CHARS`-regelen), med en korpus-skala kontroll paa 620
tilbakeholdte og en flat kontroll paa at den lista alene er >5x taket.

TO ARMER BLE SKREVET OM UNDER MAALINGEN, begge repoets vakuositets-klasse:
- Et absolutt tak paa 8 000 tegn var roedt paa fixturen (11 183) og ville uansett vaert
  meningsloest paa K2 -- det maalte ikke egenskapen, det maalte fixturstoerrelsen.
- `entry.concept_id not in rendering` var roed FORDI et levert konsept SITERER det
  tilbakeholdte. Asserten gjelder naa HEADEREN po selv forfatter; en delstreng-assert
  over hele renderingen kan ikke skille "vi lekket lista" fra "basen er kryss-lenket".

1425 passed / 5 skipped (fra 1413/5, +12, 0 fjernet). ruff + mypy rene. Golden
`shasum -a 1` av INNHOLDET = ea8c534773acdbe41ae68f2c55724d69aaf8be4f, BYTE-UENDRET.

Co-Authored-By: Claude <claude-opus-5>
2026-09-07 10:59:57 +02:00
ad9686517a feat(prepass): payload model, loader, shape gate and the binding to the mounted base [skip-docs]
[skip-docs]: ingen brukervendt flate ennaa -- CLI-flagget kommer i steg 6/7 med
README-blokka, og CLAUDE.md-raden i steg 8. Modulen er ikke naabar utenfra i denne
commiten.

Steg 1 og 2 av planen, committet sammen fordi begge armene bor i samme testfil og
ble skrevet mot samme ekte produsent-payload.

`prepass.py` konsumerer et OKF-konsumkontrakt-SS-8-payload som INPUT-fil (SS 2.4:
transporten er ikke del av kontrakten). po produserer ingen payload og vendrer ingen
produsent -- en subprosess mot produsentens checkout ville bundet pakka til en sti paa
en maskin og doedd i `git archive HEAD`, og en kopi ville vaert kø-(p)-driften.

`check_payload_shape` -- seks nekter, hver ved navn: ukjent revisjon (`==`, aldri
prefiks), nevnere som ikke lukker (SS 5.2, med DEKLARERT og OBSERVERT som to tall),
`len(excerpts)` og `len(withheld)` mot sine tellere (SS 8.1s to halvdeler), `spent >
limit` (SS 7.3) og en kjent-positiv som ikke ble reprodusert (SS 7.4). De to siste er
DEFENSIVE og uvitnet: produsenten nekter dem selv foer den emitterer.

`verify_against_bundle` -- fem sjekker mot den MONTERTE basen: erklaert id (aldri
mountet, S7a-3), joinen gjennom `safe_resolve` med `PathSecurityError` re-reist som
nekt, filas `sha256`, og -- baerende -- at `text` er RE-UTLEDBAR fra det monterte
dokumentet. Den siste lukker injeksjonsflaten: et payload kan ikke levere bytes basen
ikke holder. De to gatene som bodde i det tilbaketrukne `read_file` er reist paa nytt
her paa DOKUMENTET (verdict-laget og SS 4.1a), aldri delegert til produsentens egne
regler.

MAALT foer bygging, ikke antatt:
- Pre-passet nekter ALLE tre av repoets leverte baser ("declares no bundle_id") --
  S7a-3s egen maaling sett fra produsentsiden. `shared/` er pull-only, saa fixturen er
  bygget paa en KOPI med erklaert id, som ogsaa gir S7a-3s slakk-tilfelle gratis.
- `concept_id + ".md"`, `sha256` = HELE fila, og tekst-utledningen (frontmatter delt
  paa `splitlines()`, NFC, per-linje rstrip) holder for alle fire utdrag i et EKTE
  produsent-payload. En naiv `split("\n")` er UENIG med produsenten -- derfor er
  regelen transkribert fra kilden og gatet av fixturen, aldri gjettet.

Fixturen `tests/fixtures/prepass/` er produsentens output ORDRETT (okf `54a0bc2`), saa
ingen arm er groenn mot en form ingen pre-pass emitterer.

TO TEST-FIXTURER VAR FEIL, funnet ved aa kjoere roedt: excerpt/withheld-armene brakk
ogsaa nevnerne, saa lukke-sjekken fyrte foerst; og dimensjons-armen kjoerte mot et
umerket dokument, der `in_dimension` aldri dropper uskopet kunnskap. Begge armene ville
staatt groenne mot en implementasjon uten sjekken de er skrevet for.

1413 passed / 5 skipped (fra 1387/5, +26, 0 fjernet). ruff + mypy rene.

Co-Authored-By: Claude <claude-opus-5>
2026-09-07 10:52:20 +02:00
78e8e39147 fix(validator): stage 0 navngir ALLE baseline-overtredelser, ikke bare den foerste
K2-funn (b), maalt live i oekt 94 (docs/2026-09-06-major2-levende-k2.md § 4,
kjoering 3, max_attempts=3):

  1000/500 -> "quantity 1000 ... baseline 1250" -> 1188/350
           -> "unit_cost 350 ... baseline 850"  -> 1000/850 -> forsoekene brukt opp

Steg 5 mater avvisningsgrunnen ORDRETT inn i neste forsoeks prompt, saa en
melding som navngir ETT felt leses som en instruks om aa rette det feltet.
Modellen fant hver riktig verdi og aldri begge samtidig: den rettet feltet
avvisningen navnga og brakk det andre. Loekka oscillerte i stedet for aa
konvergere.

Endringen gjelder KUN meldingens fullstendighet. D6 er uendret: en hvilken som
helst overtredelse avviser fortsatt, i samme stage, foer loeseren.
max_attempts heves IKKE og eksponeres IKKE.

Hver overtredelse beholder dagens setning ORDRETT, sammenfoeyd med "; ", saa
NOEYAKTIG EN overtredelse rendres byte-identisk med foer - det er dette som
holder de eksisterende delstreng-assertene i S4.0-, reserve- og
levert-bundle-gatene staaende. Skilletegnet er valgt framfor linjeskift fordi
Rejection.reason ogsaa lander i outbox-JSON, de hostede payloadene og
terminal-notisene.

Rekkefoelgen er FORSLAGETS egen (linjer i oppgitt rekkefoelge, quantity foer
unit_cost i en linje), saa to identiske forsoek gir to identiske prompter. En
ukjent kostkode bidrar med sin ENE setning og ingen magnitude-setninger: det
finnes ingen baseline-linje aa avvike fra, og en sammenligning mot ingenting er
nettopp den fabrikasjonen dette steget finnes for.

Load-bearing MAALT (tests/test_stage0_all_violations_loadbearing.py, 6 armer),
seks mutasjoner ALLE ROEDE mot HELE suiten + groenn kontroll 1387/5 (fra
1381/5; +6, 0 fjernet - strengt supersett) og golden demo-transcript.stdout
BYTE-UENDRET (shasum -a 1 av INNHOLDET = ea8c534773acdbe41ae68f2c55724d69aaf8be4f):
returner ved foerste overtredelse (4 roede) - rapporter kun den siste (4) -
ustabil rekkefoelge i linja (1) - over-rapporter et felt som er INNENFOR
toleransen (28) - drift enkelt-overtredelsens form (1) - la en ukjent kode
ogsaa emittere magnitude-setninger (19).

Co-Authored-By: Claude <claude-opus-5>
2026-09-07 00:38:50 +02:00
ab747bc7e8 fix(tools): read_file paa en KATALOG nektes ved navn og peker paa read_dir
Funn (c) fra oekt 94s levende K2-maaling (docs/2026-09-06-major2-levende-k2.md
§ 3/§ 4): en levende modell gikk stigen list_bundles -> read_bundle -> read_dir
-> read_file, naadde et nivaa med underkataloger (30-1, 30-7, 521-001 ...) og
kalte read_file paa en av dem. Tre slike kall i EN kjoering.

SJEKKET FOERST, som ordren ber om: ordren aapner for to aarsaker. Det er den
ANDRE. En listing SKILLER allerede de to slagene strukturelt - hver
directory_listing-payload svarer med "directories" (oppfoeringer noeklet "path",
med subtre-telling) og "documents" (noeklet "name", med type/title/chars) som TO
distinkte noekler, og en katalog opptrer aldri blant dokumentene. Arm (4) pinner
det, fordi det er en egenskap denne fila naa HVILER paa. Derfor ingen
trailing-/-markoer: formen sier allerede hva som er hva, og aa endre payloaden
ville flyttet en listing tre eldre gates maaler byte for byte.

Det som MANGLET var den andre halvdelen. MAALT foer arbeidet, paa den shippede
nestede eksempelbasen:

    read_file(id, "a") -> IsADirectoryError: [Errno 21] Is a directory: <abs sti>

En OSError, altsaa krasj-kanalen i stedet for CLI-ens nekt-tuppel og hostings
400-arm, og den navnga verken hva stien VAR eller hvilken sprosse kalleren
skulle brukt. DirectoryPathRefused (ValueError, BundlePathNotFound- og
DimensionScopeRefused-presedensen) navngir begge. Regelen bor i VERKTOEYET, saa
den gjelder begge kallere (utforskningen, og siden S2c debatten) - samme
plassering som verdict-gaten, og FOER den: declares_verdict_type leser stiens
frontmatter, saa paa en katalog ville den reist noeyaktig den OSError-en denne
grenen finnes for aa erstatte.

MAALT, RAPPORTERT, IKKE FIKSET (utenfor ordren): det levende kallet var
read_file(".../30-7.md") - modellen la .md paa et katalogNAVN, som resolverer
til en sti som ikke finnes i det hele tatt, ikke til katalogen. Maalt paa samme
base gir den FileNotFoundError, altsaa samme form (OSError paa krasj-kanalen der
en navngitt nekt hoerer hjemme). Denne ordren fikser katalog-tilfellet; arm (5)
pinner maalingen av naboen, saa gapet er et faktum i suiten og ikke en setning i
en rapport.

RoedT foerst: import-feil paa DirectoryPathRefused (klassen fantes ikke).
Ingen endring i prompter.

Suite 1381 passed / 5 skipped (fra 1368/5; +13, 0 fjernet - strengt supersett).
ruff + mypy rene. Golden demo-transcript.stdout BYTE-UENDRET,
shasum -a 1 av INNHOLDET = ea8c534773acdbe41ae68f2c55724d69aaf8be4f.

Ordre 20260906T212735Z-2448754-from-.claude, funn (c).

Co-Authored-By: Claude <claude-opus-5>
2026-09-06 23:38:44 +02:00
c6886ed605 fix(proposer): affected_items maa baere BASELINE-linja, ikke den foreslaatte reduksjonen
Funn (a) fra oekt 94s levende K2-maaling (docs/2026-09-06-major2-levende-k2.md
§ 4, rad 3-4). Genererings-prompten beskrev feltet som

    affected_items (list of {code, quantity, unit_cost})

og sa ingenting om HVEM sine tall det er. Stage 0 (S4.0) avstemmer hver linje
mot prosjektets egen kostbaseline, saa de eneste tallene som kan passere er
prisskjemaets EGNE - men en modell som blir bedt om aa halvere et volum leser
quantity som feltet tiltaket sitt hoerer i, fyller inn den REDUSERTE verdien
(1000 der baselinen sier 1250) og avvises. Den levende kjoeringen hadde lest det
prisede skjemaet TO ganger gjennom navigasjonsstigen og gjettet likevel: tallene
var tilgjengelige, KONTRAKTEN for feltet var ikke uttalt.

Besparelsen har sitt eget felt, og prompten sier det i samme aandedrag - "dette
er baseline-tallene" uten "og reduksjonen din hoerer der" etterlater modellen med
en verdi den er fortalt aa ikke putte noe sted.

Blokka rir paa BASE-prompten, ikke paa en gren, og det er hva arm (3) finnes for:
defekten ble maalt paa et REVIDERT forsoek, saa en instruksjon som bare naadde
forsoek 1 ville vaert fravaerende fra noeyaktig den prompten den ble maalt i.
prior_rejection / prior_feedback / approach appendes ETTER basen, saa hver
komposisjon baerer den fortsatt.

RoedT foerst: 6 av 7 armer roede (kontrollen groenn - den asserterer at
referanseprosjektet faktisk HAR kostlinjer, saa gaten ikke beskriver noe
imaginaert). Ingen endring i skjema, validator eller stage 0.

Suite 1381 passed / 5 skipped (fra 1368/5; +13, 0 fjernet - strengt supersett).
ruff + mypy rene. Golden demo-transcript.stdout BYTE-UENDRET,
shasum -a 1 av INNHOLDET = ea8c534773acdbe41ae68f2c55724d69aaf8be4f.

Ordre 20260906T212735Z-2448754-from-.claude, funn (a).

Co-Authored-By: Claude <claude-opus-5>
2026-09-06 23:38:16 +02:00
a790ce3ae1 fix(hosting): _run_kwargs reads _REFUSED_BY_NAME instead of a hardcoded literal
MAJOR-2 review, fokuspunkt 4: `_REFUSED_BY_NAME = ("proposal_review",)` was
documented as the tuple the door reads, but the `if "proposal_review" in
payload` refusal never consulted it — two copies of one fact, free to drift.

Måling FØR bygging: `grep -rn 'no terminal to answer' tests/` gave 0 treff,
men `grep -rn proposal_review tests/ | grep -i hosting` fant
`test_proposal_review_loop_loadbearing.py:1677` (T21) — so half of the fact
(refusing `proposal_review` itself, with the exact CLI-pointing message) WAS
already pinned. The unpinned half was the second name: any OTHER entry added
to `_REFUSED_BY_NAME` fell through to the generic `unknown field(s)` check
instead of being refused by name before it.

Fix: `_run_kwargs` now iterates `_REFUSED_BY_NAME`; `proposal_review` keeps
its verbatim message, any other name gets a message naming the field.

New test `test_the_named_refusal_reads_the_constant_not_a_literal`
(tests/test_hosting_loadbearing.py) is unit-level against `_run_kwargs`
directly: confirms an unlisted name is refused generically (control), then
monkeypatches `_REFUSED_BY_NAME` wider and confirms the SAME payload is now
refused by name, before the generic check.

Mutation check (done and reverted): setting `_REFUSED_BY_NAME = ()`
temporarily turns the EXISTING T21 test red — proving the constant is now
load-bearing. Verified against pre-fix code that the same mutation left T21
green (the old literal-based `if` never consulted the constant at all), so
the fix closes the exact drift the review flagged.

Suite: 1368 passed / 5 skipped (was 1367/5 on 17998af), 0 regressions.
ruff + mypy clean. Golden demo-transcript.stdout byte-unchanged
(shasum -a 1 = ea8c534773acdbe41ae68f2c55724d69aaf8be4f).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-06 07:21:04 +02:00
257f44f0cb fix(major2): en skrivefeil i finally fortrenger ikke lenger stoppet i flukt
Funn d71a5d72 (MINOR, MISSING_ERROR_HANDLING, run.py:1202). Begge
finally-skriverne ligger i propageringsstien til nettopp det unntaket de er
bevis for: en OSError fra mkdir/write_text mens BudgetExceeded eller
ProposalReviewInputError er i flukt ERSTATTER den - og hverken
`except ProposalReviewInputError` (:3404) eller nekt-tuppelen (:3412) fanger
OSError, saa operatoeren fikk traceback og grunnen til at kjoeringen stoppet var
borte. Review-skriveren gaar paa HVER kjoering med reviewer; parse-skriveren
har samme form.

`_write_or_report` er EN kopi for begge kallstedene (koe-(p)): en regel om hva
en skriver faar gjoere med et unntak i flukt, kopiert, blir en regel anvendt paa
bare det ene. Vakten er BETINGET, aldri en blanket except - uten noe i flukt
finnes ingen stoppgrunn aa beskytte, og en kjoering som ikke fikk skrevet
utboksen maa si fra ved aa feile. Feilen SIES uansett, fordi et fravaerende
artefakt ellers leses som en kjoering uten noe aa registrere (T10/T11).

`in_flight` fanges eksplisitt (`except BaseException as stop: ... raise`), ikke
via `sys.exc_info()`, som ville lest et ytre except-lag hos en bibliotekkaller
som en flukt her.

MAALT mot HELE suiten, to mutasjoner, hver med sin egen signatur, kontroll
1367 passed / 5 skipped og golden `demo-transcript.stdout` BYTE-UENDRET
(`shasum -a 1` av INNHOLDET = ea8c534773acdbe41ae68f2c55724d69aaf8be4f):
MC vakten detached (2 roede - de to in-flight-armene) - MD svelg ubetinget
(1 roed - KONTROLL-armen alene, altsaa er betingelsen selv gatet).
Iron Law: begge in-flight-armene skrevet FOERST og maalt roede mot uendret
run.py; kontroll-armen var groenn foer fiksen, som er nettopp
diskrimineringen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 22:29:56 +02:00
d4a0220b68 test(major2): hver genereringstelling paret med debattens telling (SC1/SC4)
Funn 5935d942 (MISSING_TEST, tests/...:820). Briefens SC1 gjorde paringen til
vakten mot aa telle en debatt-tur som en genereringsprompt - og ingen arm i
fila asserterte debattens turantall mellom treated og control (grep «debate»
traff bare en kommentar og et filnavn).

`_debate_entries` er KOMPLEMENTET av `_GENERATION_MARK`, ikke en positiv
debatt-markoer: paastanden som gates er nettopp «ingenting som IKKE er
generering flyttet seg», og en positiv markoer ville latt en tur klassifisereren
ikke kjenner drive usett.

- T5-run: de tre genereringstellingene paret med likhet paa debatt-oppfoeringene.
- T8: fikk sink + en control-kjoering (alltid-godkjenn reviewer) - dens
  attempt-indekser [0,1] per kandidat ER en genereringstelling.
- Begge har en VAKUITETSVAKT (debatt-lista maa vaere ikke-tom): to tomme lister
  er like gratis.
- T13: record-indeksene er DOKUMENTERT som SC4s telle-proxy ved den doera -
  barnet kjoerer i egen interpreter og `--scripted-replies` har ingen
  sink-dump, saa prompt-nivaaet maales in-process (T1 for verbatim, T5-run/T8
  for paringen). Reviewens andre alternativ.

MAALT mot HELE suiten, to mutasjoner, begge roede paa NOEYAKTIG de to parede
armene og paa ingen andre: MA klassifisereren returnerer konstant tom liste
(2 roede - vakuitetsvakten) og MB filteret droppet, saa generering telles som
debatt (2 roede). Kontroll 1364 passed / 5 skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 22:29:27 +02:00
c81a90c88a fix(major2): --proposal-review komponerer med --resume, og armen som beviser det
To funn, ETT commit, fordi de deler vitnet ved konstruksjon: 02159d21
(run.py:2738, PLAN_EXECUTE_DRIFT) er vakten, da485928
(tests/...:1344, MISSING_TEST) er armen som skulle sett den. AA dele dem ville
krevd en kastbar duplikattest.

VAKTEN: `if args.checkpoint_dir is not None` nektet --proposal-review for HVER
argv med en checkpoint-katalog - men --resume KREVER --checkpoint-dir
(run.py:2776), saa komposisjonen nekten selv anbefaler («pass --proposal-review
at --resume instead») var unaabar, og run.py:2218-helpen, README.md:546 og
MAJOR-2-raden beskrev en sti ingen argv kunne ta. Vakten nekter naa en PARK
(en etappe som returnerer foer noen kandidat finnes), aldri et LOEFT.

ARMEN: `test_the_door_composes_with_resume` sendte hverken --resume,
--checkpoint-dir eller --review-inbox - den var en vanlig enkeltkjoering T13
allerede dekket, altsaa groenn mot nettopp den defekten den var navngitt for.
Den driver naa en EKTE resume: dag 1 parkerer en ekte plan-review gjennom
run.main, ekspertens svar legges i en ekte innboks, og dag N sender
--resume ... --checkpoint-dir ... --review-inbox ... --proposal-review med
run_project innspilt og _refuse_model som kontroll paa null modellkall.

MAALT, i denne rekkefoelgen (Iron Law): armen skrevet FOERST og kjoert mot
uendret vakt -> ROED med nettopp nektlinja i stderr («pass --proposal-review at
--resume instead»), calls == []. Etter vakt-fiksen: 49 passed i fila, og
park-nekten (test_the_door_and_a_parked_exploration_contradict, M39) staar
groenn - den sender --explore uten --resume.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 22:01:49 +02:00
4a3655fc07 docs(major2): invariantraden - eksperten svarer paa forslaget paa bordet, og svaret brukes
CLAUDE.md: en invariantrad etter verdict-gate-raden, i husets skjelett - nevneren
(grep "prior_" finner EN soem, maskinens egen), D6 og hvorfor alternativet ble
forkastet, den ledger-bevisste attempts_remaining, honoured = "hentet faktisk",
kanalvalget begrunnet av tre maalinger, den kaller-eide sinken, skriv-iff-reviewer,
de fem nektene ved navn og den hostede pre-whitelist-sjekken. Load-bearing-blokka
lister M1-M40 med roedtall, M29 staar som et FUNN (uvitnet baerer), og de to
armene som var groenne av feil grunn i steg 1-8 er skrevet ned som repoets
vakuoes-gate-klasse, sekstende og syttende gang. Aerlighets-grensene til slutt.

README.md: --proposal-review i enkeltprosjekt-flagglista og en prosablokk etter
--plan-review/--checkpoint-dir-paret, i samme form - hva operatoeren ser og
skriver, at en revise KJOEPER ett forsoek til under de eksisterende takene, at
approve ikke er en ekspertdom, og alle fem nektene navngitt (--checkpoint-dir-en
peker paa --resume, doera som VIRKER). Kundevendt vokabular.

Ny arm: test_the_readme_block_names_every_flag_the_cli_refuses_the_door_with,
Fase-3-formen (raa tekst, uttrukket blokk, kontroll paa at uttrekkeren finner noe
som finnes). Intet M-nummer - lista lukket ved M40 - saa den ble drevet ROED TO
ganger: mot README-en foer blokka fantes, og med blokka paa plass men
--checkpoint-dir omskrevet til aa beskrive nekten uten aa navngi flagget.

1364 passed / 5 skipped. Golden demo-transcript.stdout BYTE-UENDRET
(shasum -a 1 av INNHOLDET = ea8c534773acdbe41ae68f2c55724d69aaf8be4f). Node-ID-ene
er et strengt supersett: 1319 -> 1369, 0 fjernet. mypy og ruff rene.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 20:40:16 +02:00
ee7da3bb94 feat(major2): run_mandate_across_bundles threads one reviewer; run_portfolio takes none [skip-docs]
Ordre 20260904T173146Z-8102814273-from-portfolio-optimiser, steg 8 av 10.

Dispatcheren faar proposal_reviewer og videresender den til hver per-base run_project.
EN gjenstand, aldri en kopi per base: dispatchen er SEKVENSIELL, saa en terminal-reviewer
komponerer. Assertert med `is`, ikke `==` - en fersk reviewer per base ville vaert en annen
gjenstand med identisk oppfoersel, som `==` paa en vanlig callable ikke kan skille (samme
identitets-leksjon test_multibase_loadbearings store-arm ble rettet til).

Doera faar sitt EGET vitne (S7a-3-regelen: hver doer som aapner en base faar sin egen
mutasjon, fordi en uvitnet kopi kan regrere alene).

run_portfolio faar INGENTING - samtidige boelger deler en terminal, som er --portfolio-
partisjonens egen grunn - og fravaeret er ASSERTERT, saa en senere "symmetri"-endring er en
roed test og ikke en stille utvidelse.

RODT foer impl: T22 (TypeError paa ukjent keyword).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 07:10:31 +02:00
bfc634d806 feat(major2): the hosted surface refuses the proposal review by name and points at the CLI [skip-docs]
Ordre 20260904T173146Z-8102814273-from-portfolio-optimiser, steg 7 av 10.

En PRE-whitelist-sjekk paa raa payload, ETTER isinstance-vakten (en ikke-objekt-body skal
beholde sin 400, ikke bli en 500) og FOER den generiske unknown-field-sjekken, som ellers
ville svart foerst.

Plasseringen er MAALT, ikke valgt: _CONSUMED_FIELDS maa vaere disjunkt fra run_projects
parametre (Fase 4es negative halvdel) mens proposal_reviewer ER en av dem, saa navnet kan
ikke bo i noen av de tre listene. F4-presedensen er IKKE analog - enable_plan_review er en
NOESTET noekkel inne i det whitelistede explore_contract, som er derfor den kan navngis der.

DISKRIMINATOREN ER TEKSTEN, IKKE STATUSEN: whitelisten svarer alt enhver ukjent nokkel med
400 "unknown field(s)", saa en detachet navngitt nekt ville fortsatt gitt 400 med feltnavnet.
Den navngitte meldingen peker paa CLI-doera og paa /readiness, og kontrollen (et ordinaert
ukjent felt) asserterer at den generiske meldingen deler ingenting av det.

Null run_project-kall, ikke bare en 400 (oekt 57).

_response_payload er IKKE utvidet: flaten nekter revieweren, saa feltet kunne kun vaert tomt.

RODT foer impl: tre armer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 07:09:57 +02:00
fedf9897c4 feat(major2): --proposal-review answers the review at the terminal; four refusals by name, EOF stops the run [skip-docs]
Ordre 20260904T173146Z-8102814273-from-portfolio-optimiser, steg 6 av 10.

(a) argparse --proposal-review. (b) rader i BEGGE partisjonene (report_forbidden og
single_only) - report-modus og portefoelje returnerer over dispatchen, saa en utelatelse
er et STILLE DROPP, ikke en nekt (F4-gapet). (c) TRE navngitte nekter i EN topp-nivaa-blokk
if args.proposal_review: - plasseringen er MAALT, ikke plassert paa oeyemaal: naboen
--scripted-replies/--live-dry-run er nostet under if args.scripted_replies, og
--explore/--live-dry-run under if args.explore, saa under noen av dem ville et bart
--live-dry-run --proposal-review falt rett gjennom til dry-run-dispatchen og droppet flagget.
(d) reviewer bygget paa KALLSTEDET + except ProposalReviewInputError -> "run stopped:" rc 1,
en DISTINKT kanal fra "run refused:". (e) proposal_review_notice printes fra kjoeringens EGEN
post. (f) _load_scripted_replies' aerlighetsgrense navngir review-stien.

--resume KOMPONERER (A3 verifisert av en arm, ikke utsatt): resume-blokka gir mandatet og
faller gjennom til SAMME full-run-dispatch.

RODT foer impl: 9 armer. T13 og T16 kjoerer i et BARN (P4). Nekt-armene kjoerer in-process
med _default_factory som REISER - ved exit-koden ser en nekt etter forbruket identisk ut med
en foer (oekt 57).

TO ARMER BLE FALSIFISERT AV MAALINGEN FOER de kunne gate noe:
(1) T18s rc-0-kontroll avslorte at F4-testens ledger-fixtur ({"entries": []}) faar rc 1 av
SavingsLedger.load ("must be a JSON array"), ikke av partisjonsraden - armen ville vaert
groenn mot en fjernet rad. Fixturen er naa en JSON-array, og kontrollen beviser at argv-en
ellers ville blitt AKSEPTERT.
(2) notice-null-armen ga BudgetExceeded i stedet for en avvist kjoering: et to-stegs
proposer-manus mot max_attempts=3 faller til default-svaret, som aldri parser, og rundeboka
fyrer - noeyaktig aerlighetsgrensen _load_scripted_replies uttaler, reprodusert ved uhell.
Manuset har naa like mange steg som forsoek.

Planens T19 er foldet inn i T13 og uttalt: "et bart, uskriptet flaggparse" ville kalt en
levende modell, saa argparse-vitnet er barnets egen unrecognized-arguments-assert.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 07:06:22 +02:00
c391fb5d67 feat(major2): terminal proposal reviewer - text never repr, closed vocabulary, EOF is never a sign-off [skip-docs]
Ordre 20260904T173146Z-8102814273-from-portfolio-optimiser, steg 5 av 10.

terminal_proposal_reviewer er et SOESKEN av explore.terminal_plan_reviewer ved FORM -
kopiert, ikke delt: en felles "terminal reviewer"-abstraksjon over to doerer er den
enkeltbruks-generaliseringen repoet nekter til en tredje doer finnes.

Kandidaten rendres som TEKST fra den typede IR-en (BLOCKER-1): maal, kostlinjer, krevd
besparelse, validatorens persentiler, checkerens dom (D3) og attempts remaining. Begge
halvdeler er gatet - en POSITIV sentinel bare tekst-stien kan sende, og den NEGATIVE
formen paa selve defekten (" object at 0x"), fordi den positive alene ville vaert
tilfreds med en renderer som printer ingenting.

Stroemmene resolveres ved KALL-tid, ikke i fabrikken.

Fail-closed paa ekspertens EGEN input: skrivefeil, blank linje og bar "revise" spoerres
paa nytt; D1(a) nekter en revise som ikke kan kjoepes AT THE DOOR med et faktum, aldri
med et botemiddel CLI-en ikke kan utfoere (det finnes ingen --max-attempts, og D1 legger
ingen til) - en tredje doer-TILSTAND, ikke et tredje ord. EOF reiser
ProposalReviewInputError: aa lese stillhet som godkjenning ville latt en kjoering baere
en kandidat ingen signerte, usynlig.

RODT foer impl: fem armer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-05 06:59:54 +02:00