Commit graph

3 commits

Author SHA1 Message Date
332eb5965b
test(b-gate): 31 arms against the gate's own denominators, 30 red on an assert about behaviour
The PM checkpoint on 207337c judged the gate DELVIS: row 1's M=13 is a curated list in the
gate's OWN b_gate.json (the run path has 41 po-calls, 7 of 10 outbox writers), rows 4, 5 and 6
have denominators with no source at all, 6 of 10 cheat-attacks got through, and the
never-Claude guard sees 433 of 512 published files.

This commit is the red half. Every arm fails on an ASSERT about behaviour, never at collection:
the four names that do not exist yet (ENTRY_KINDS, run_path_calls, registered_entry,
published_files) are stubbed here with DELIBERATELY wrong values — everything is a door, the
run path calls nothing, the surface is empty — so each arm measures the defect rather than the
absence of a symbol.

30 av 31 red on assert. The one that is green is the rc-0 control
(test_a_valid_attestation_is_the_only_thing_that_turns_row6_green): a valid attestation must
turn row 6 green both before and after, or the row refuses everything, which proves as little
as refusing nothing. All 34 pre-existing arms stay green — measured, not assumed.

The planted claude-invocations are base64 in the test file for the same reason the contract's
patterns are: tests/ is itself part of the surface row 3 scans, and a cleartext variant here
would register as its own finding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 19:14:17 +02:00
f7ade7aa8b
feat(b-gate): the gate that measures po as a toolbox, red on six measured rows [skip-docs]
python -m portfolio_optimiser.evals.b_gate — one command, offline, no model call, exit 1 today:

  1 steg i kjørestien kallbare utenfra        3 av 13   RØD
  2 roller som kan leveres utenfra            0 av 2    RØD
  3 vakter mot en vei fra po til Claude       6 av 6    GRØNN  (435 published files)
  4 løpet drevet uten et eneste modellkall    0 av 3    RØD
  5 Foundry-veien urørt og samme artefaktfamilie 1 av 2 RØD
  6 kjøreboka finnes og er kjørt              0 av 2    IKKE MÅLT

Every denominator is read off the source, never off a list in the gate. Row 1 counts the steps
of the run path that resolve to a symbol AND have a call site; a step is externally callable only
when a CLI (or MCP-registered) entry reaches it without any chat-client construct on the way —
which is why the ten run.py steps are red and round_builder's two plus the v1 gate are green. Row
2 reads the roles off workflow._MAKER_CHECKER_ROLES. Row 3's patterns each carry a known-positive
AND a known-negative fixture, so a guard that cannot hit is not counted as a zero.

Three decisions the operator cannot answer without reading code, made here and stated in the
gate's own output:

* the external door is a CLI subcommand, not MCP — po already has five main() and two console
  commands, and MCP would need a server the run path does not have. The gate still counts an
  MCP-registered door, so the choice does not bind the next order.
* the budget guard in B is NOT po's: BudgetMiddleware is fail-closed on missing usage and is
  never constructed without a chat client, so keeping it here would turn fail-closed into
  fail-open. The ceiling in B is the Claude Code session's own spend, which po neither sees nor
  steers. The Foundry path keeps its ceiling unchanged.
* row 6 is IKKE MÅLT, never green, until the operator attests that the runbook actually drove an
  analysis — a file the gate never writes, the same rule as the v1 gate's attestation.

Row 3's pattern text is base64 in the config so the contract cannot register as its own finding;
that is what lets the row run without an exclusion list, and a row without exclusions is a row
nobody can switch off by adding a filename.

Suite after: 2106 passed / 5 skipped / 5 xfailed (was 2072/5/5; +34 new, none changed).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 08:29:10 +02:00
0a784065d0
test(b-gate): 34 red tests for the gate that measures po as a toolbox Claude Code drives
Operator decision 19.09.2026: in development and test Claude Code LEADS and portfolio-optimiser
is the toolbox. po never calls Claude; production stays on Foundry. This commit writes the
measurement RED — the contract first, the capability later.

Six rows, each with a denominator taken from the SOURCE and counted independently here:

1. toolbox complete — every deterministic step of the run path, anchored to (module, symbol) and
   to the scope that calls it; k = steps reachable from a CLI entry WITHOUT a chat client.
2. what the model delivered can be delivered from outside — denominator read off
   workflow._MAKER_CHECKER_ROLES, k = roles with passing named probes.
3. po has no path to Claude — six patterns, each carrying its own known-positive AND
   known-negative fixture. Pattern text is base64 in the config so the gate's own contract cannot
   register as a hit against the surface it scans.
4. no model calls in toolbox mode — three named probes.
5. the Foundry path untouched — profiles, factory and injection seam read from source, plus a
   schema-comparison probe.
6. the runbook exists — NOT MEASURED until the operator attests, and the gate never writes that
   attestation itself.

Red now for the reason that matters: the tests fail to import a module that does not exist yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 08:01:29 +02:00