feat(persona): build the shared expert-reviewer persona as a framework-neutral Agent Skill
The expert reviewer was only a hardcoded verdict_input dict inside the offline simulation. Build it as the real, shared artifact target picture §8 calls for: shared/skills/expert-reviewer/ — a SKILL.md persona prompt (energy-advisor / M&V role + the realization-gap methodology the validator cannot compute) plus a canonical references/example-verdict.json. shared/ stays pure data; the MAF side reads it via portfolio_optimiser.persona.load_persona_example (call-time, fail-fast) and the Claude-SDK sibling reads the same JSON with its own loader. This de-stubs the simulation: its persona judgement (decision + rationale + traced marker) is now sourced from the artifact at call time, not an inline literal — so the shared persona is genuinely consumed and cannot rot silently. decision is binary (approved/rejected, the FeedbackContract the run path accepts); approved_with_adjustment is rejected there and lives only in the bundle seed frontmatter + the promotion gate, so the realization correction is carried in the rationale prose. Load-bearing trio (tests/test_persona_skill_loadbearing.py), each proven RED on its own detach: structure + framework-neutrality, the example is valid pipeline input (incl. FeedbackContract, on a throwaway copy), and the simulation's marker follows the artifact file. Suite 149->152. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MHR8iKxJRxDiDfNw8HZmWE
This commit is contained in:
parent
a9144cb9bb
commit
6f861a0078
8 changed files with 281 additions and 13 deletions
14
CLAUDE.md
14
CLAUDE.md
|
|
@ -100,6 +100,20 @@ Python ≥3.10. MAF (`agent-framework-core` 1.9.0). Pakkehåndtering: `uv`. To b
|
||||||
lever på Claude-SDK-siden (minimal API-kjøring). Skriptet klient = MAF-side stillas, IKKE delt
|
lever på Claude-SDK-siden (minimal API-kjøring). Skriptet klient = MAF-side stillas, IKKE delt
|
||||||
(`shared/` forblir framework-nøytralt). Kjøres `uv run python -m portfolio_optimiser.simulation`.
|
(`shared/` forblir framework-nøytralt). Kjøres `uv run python -m portfolio_optimiser.simulation`.
|
||||||
Load-bearing: `tests/test_simulation_loadbearing.py` blir RØD når promoteringen detaches.
|
Load-bearing: `tests/test_simulation_loadbearing.py` blir RØD når promoteringen detaches.
|
||||||
|
- **Delt ekspert-persona som Agent Skill (§8, framework-nøytral):** ekspert-reviewer-personaen bor i
|
||||||
|
`shared/skills/expert-reviewer/` (`SKILL.md` + `references/example-verdict.json`) og er den ENE
|
||||||
|
delte artefakten begge stacker instansierer reviewer-en fra. `shared/` forblir REN DATA — MAF-siden
|
||||||
|
leser den via `portfolio_optimiser.persona.load_persona_example` (call-time, fail-fast), Claude-SDK-
|
||||||
|
søskenet med sin egen loader mot samme JSON. Dette AV-STUBBER simuleringen: persona-dommen (decision
|
||||||
|
+ rationale + sporet markør) hentes nå fra artefaktet ved call-time, ikke en inline-literal — så
|
||||||
|
personaen er genuint konsumert og kan ikke råtne stille. **Decision er binær** (`approved`/`rejected`
|
||||||
|
— `FeedbackContract` run-stien tar; `approved_with_adjustment` avvises der, bor kun i bundle-seedens
|
||||||
|
frontmatter + promoterings-gaten); realiseringskorreksjonen lever i rationale-prosaen, ikke et tredje
|
||||||
|
enum. SKILL.md-prosaen nevner ALDRI en konkret framework (maf-guarden er import-formet). Load-bearing-
|
||||||
|
trio (`tests/test_persona_skill_loadbearing.py`): struktur+framework-nøytralitet (RØD på framework-
|
||||||
|
import), eksempelet er gyldig pipeline-input inkl. `FeedbackContract` (RØD på skjema-/kontrakt-drift,
|
||||||
|
på en throwaway-kopi — aldri den git-tracked fixturen), og sim-ens markør følger artefakt-fila (RØD i
|
||||||
|
det øyeblikk personaen re-inlines).
|
||||||
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
- **STATE.md er local-only** (gitignored). Voyage session-state er efemert; STATE.md er kanonisk kontinuitet.
|
||||||
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
- Prosess: Voyage-plugin (`/trekbrief → /trekplan → /trekexecute → /trekreview`) per større fase.
|
||||||
|
|
||||||
|
|
|
||||||
18
README.md
18
README.md
|
|
@ -43,12 +43,28 @@ for the swarm and the expert persona (honesty per §1). The scripted client is M
|
||||||
not part of the framework-neutral `shared/` core. Load-bearing: `tests/test_simulation_loadbearing.py`
|
not part of the framework-neutral `shared/` core. Load-bearing: `tests/test_simulation_loadbearing.py`
|
||||||
goes red the moment promotion is detached (the marker never crosses into Run B).
|
goes red the moment promotion is detached (the marker never crosses into Run B).
|
||||||
|
|
||||||
|
## Shared expert-reviewer persona (§8)
|
||||||
|
|
||||||
|
The expert reviewer is a **framework-neutral Agent Skill** in [`shared/skills/expert-reviewer/`](shared/skills/expert-reviewer/)
|
||||||
|
— a `SKILL.md` persona prompt (the energy-advisor / M&V role, the realization-gap methodology the
|
||||||
|
validator cannot compute) plus a canonical `references/example-verdict.json`. It is the shared
|
||||||
|
artifact both reference implementations consume to instantiate the reviewer; `shared/` stays pure
|
||||||
|
data, so the MAF side reads it via `portfolio_optimiser.persona.load_persona_example` and the
|
||||||
|
Claude-SDK sibling reads the same JSON with its own loader. This **de-stubs the simulation**: its
|
||||||
|
persona judgement (decision + rationale + traced marker) is now sourced from the artifact at call
|
||||||
|
time, not a hardcoded literal — so the shared persona is genuinely consumed and cannot rot silently.
|
||||||
|
The persona's `decision` is binary (`approved` / `rejected`, the feedback contract the run path
|
||||||
|
accepts); the realization correction lives in the rationale prose. Load-bearing trio
|
||||||
|
(`tests/test_persona_skill_loadbearing.py`): structure + framework-neutrality (red on a framework
|
||||||
|
import), the example is valid pipeline input (red on schema/contract drift), and the simulation's
|
||||||
|
marker follows the artifact file (red the moment the persona is re-inlined).
|
||||||
|
|
||||||
## Docs
|
## Docs
|
||||||
|
|
||||||
- [`docs/plan/2026-06-26-maalbilde-agentic-loop.md`](docs/plan/2026-06-26-maalbilde-agentic-loop.md) — target picture: the agentic cost-saving loop + OKF knowledge architecture (north star).
|
- [`docs/plan/2026-06-26-maalbilde-agentic-loop.md`](docs/plan/2026-06-26-maalbilde-agentic-loop.md) — target picture: the agentic cost-saving loop + OKF knowledge architecture (north star).
|
||||||
- [`docs/research/2026-06-23-prior-art-platform.md`](docs/research/2026-06-23-prior-art-platform.md) — prior-art & platform research (incl. implementation register §15).
|
- [`docs/research/2026-06-23-prior-art-platform.md`](docs/research/2026-06-23-prior-art-platform.md) — prior-art & platform research (incl. implementation register §15).
|
||||||
- [`docs/plan/2026-06-23-incremental-plan.md`](docs/plan/2026-06-23-incremental-plan.md) — incremental delivery plan (deterministic backbone).
|
- [`docs/plan/2026-06-23-incremental-plan.md`](docs/plan/2026-06-23-incremental-plan.md) — incremental delivery plan (deterministic backbone).
|
||||||
- [`shared/`](shared/) — framework-neutral shared core (concept + example OKF knowledge bundles), reused unchanged by both reference implementations.
|
- [`shared/`](shared/) — framework-neutral shared core (concept + example OKF knowledge bundles + the expert-reviewer persona skill), reused unchanged by both reference implementations.
|
||||||
|
|
||||||
## Stack
|
## Stack
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -20,7 +20,12 @@ so the only thing that differs is the agent framework itself.
|
||||||
bundle (OKF / LLM-wiki): one office building, one LED-retrofit measure, with a seed expert
|
bundle (OKF / LLM-wiki): one office building, one LED-retrofit measure, with a seed expert
|
||||||
verdict encoding the realization gap and a golden-suite of expected validator outcomes. A
|
verdict encoding the realization gap and a golden-suite of expected validator outcomes. A
|
||||||
small **dev fixture** for exercising the agentic loop; a realistic full-scale example comes later.
|
small **dev fixture** for exercising the agentic loop; a realistic full-scale example comes later.
|
||||||
- *(planned)* the method specification and the expert-reviewer persona.
|
- [`skills/expert-reviewer/`](skills/expert-reviewer/) — the **expert-reviewer persona** as a
|
||||||
|
framework-neutral Agent Skill: a `SKILL.md` persona prompt (energy-advisor / M&V role + the
|
||||||
|
realization-gap methodology the validator cannot compute) and a canonical
|
||||||
|
`references/example-verdict.json`. Both reference implementations instantiate the reviewer from
|
||||||
|
this one artifact; `shared/` stays pure data (each stack reads the JSON with its own loader).
|
||||||
|
- *(planned)* the method specification.
|
||||||
|
|
||||||
## Rules
|
## Rules
|
||||||
|
|
||||||
|
|
|
||||||
68
shared/skills/expert-reviewer/SKILL.md
Normal file
68
shared/skills/expert-reviewer/SKILL.md
Normal file
|
|
@ -0,0 +1,68 @@
|
||||||
|
---
|
||||||
|
name: expert-reviewer
|
||||||
|
description: Adopt the expert energy-advisor persona to judge a deterministically-validated cost-saving proposal — render a verdict (approve / approve-with-adjustment / reject) that encodes the realization gap the validator cannot compute. Use after the deterministic validator has accepted a proposal's numbers and a human-grade domain judgement is needed.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Expert reviewer — energy advisor (M&V)
|
||||||
|
|
||||||
|
You are an experienced energy advisor and measurement-and-verification (M&V) professional. Your
|
||||||
|
role in the loop is the **human-grade judgement** that comes *after* the deterministic validator
|
||||||
|
has already confirmed a proposal's numbers are arithmetically sound and within a feasible range.
|
||||||
|
You are not a calculator and you are not a second validator — you supply the experiential knowledge
|
||||||
|
the math cannot reach.
|
||||||
|
|
||||||
|
This persona is **framework-neutral**: it is consumed unchanged by every implementation of the
|
||||||
|
method. It depends on no specific agent toolkit, transport, or vendor.
|
||||||
|
|
||||||
|
## What you receive
|
||||||
|
|
||||||
|
1. A **validated savings proposal** for one project measure: the measure, the affected cost items,
|
||||||
|
the claimed saving, and the validator's confirmation that the claim sits within the feasible
|
||||||
|
(e.g. P90) range.
|
||||||
|
2. The project's **curated knowledge bundle** — project documents, the assessment methodology, the
|
||||||
|
verified literature on realization gaps, and the hard constraints (budget, what cannot change).
|
||||||
|
|
||||||
|
## What you produce
|
||||||
|
|
||||||
|
A single verdict, two fields:
|
||||||
|
|
||||||
|
- `decision` — `approved` or `rejected`. The feedback the loop consumes is **binary**. The
|
||||||
|
*approve-with-correction* case — the signature case in energy work, where the measure is worth
|
||||||
|
doing but the modelled saving overstates the expected actual — is an `approved` decision whose
|
||||||
|
rationale records the correction. Reserve `rejected` for measures that should not proceed
|
||||||
|
(infeasible in practice, unsafe, mandated spec, or a realization gap that erases the benefit).
|
||||||
|
- `rationale` — prose that carries the knowledge the validator cannot compute. For an approval that
|
||||||
|
corrects, the rationale MUST state the **realization rate** you expect and the **expected actual**
|
||||||
|
saving, and *why* — the specific operational mechanism, not a generic hedge. This is where the
|
||||||
|
learning signal lives; it is folded back into the next run's hypothesis.
|
||||||
|
|
||||||
|
The canonical machine-readable shape is in [references/example-verdict.json](references/example-verdict.json).
|
||||||
|
|
||||||
|
## The judgement — the realization gap
|
||||||
|
|
||||||
|
The deterministic validator proves the *modelled* saving is correct from the parameters. Your job
|
||||||
|
is to judge the **realization gap**: the systematic bias between that modelled saving and what the
|
||||||
|
building will *actually* realize in operation. This gap is **not** parameter spread (the validator's
|
||||||
|
risk simulation already covers that) — it is a directional skew the parameters do not carry, visible
|
||||||
|
only in accumulated operating experience:
|
||||||
|
|
||||||
|
- **Hours-of-use overestimation (usually dominant):** the assumed schedule typically exceeds metered
|
||||||
|
burn time — daylight, empty rooms, occupancy controls. A timetable-stipulated 3000 h often meters
|
||||||
|
materially lower.
|
||||||
|
- **In-service rate < 1:** not every installed unit is necessarily mounted and operating at the time
|
||||||
|
of evaluation.
|
||||||
|
- **Behaviour and persistence:** rebound (more light because it is "now free") and overridden controls
|
||||||
|
erode the saving over time.
|
||||||
|
|
||||||
|
You cannot derive the realization rate from the proposal's parameters — that is exactly why a human
|
||||||
|
judgement is required here and a deterministic rule is not. Ground every correction in the bundle's
|
||||||
|
verified literature; never invent a number.
|
||||||
|
|
||||||
|
## Discipline
|
||||||
|
|
||||||
|
- **Provenance:** your verdict is stamped with who judged it, on which experiment, and when. Only an
|
||||||
|
approved (or approved-with-adjustment) verdict is eligible to be promoted back into the project's
|
||||||
|
knowledge base; a rejection never contaminates it.
|
||||||
|
- **Context-bound learning:** state the context your correction holds for (building type, the source
|
||||||
|
of the hours-of-use assumption). The next similar proposal in the same context should inherit it.
|
||||||
|
- **Honesty:** if you lack the experience to judge a measure, say so and do not fabricate a rate.
|
||||||
|
|
@ -0,0 +1,5 @@
|
||||||
|
{
|
||||||
|
"decision": "approved",
|
||||||
|
"marker": "realiseringsgrad=0.79",
|
||||||
|
"rationale": "Godkjent med realiseringskorreksjon. Den modellerte besparelsen er teknisk korrekt fra parameterne og validatoren bekrefter at den er innenfor feasibelt omraade. Men i drift realiseres erfaringsvis ~79% av en timeplan-stipulert LED-besparelse i kontorbygg (realiseringsgrad=0.79) pga. overestimerte driftstimer og in-service rate < 1; forventet faktisk besparelse ca 23700 NOK/aar."
|
||||||
|
}
|
||||||
52
src/portfolio_optimiser/persona.py
Normal file
52
src/portfolio_optimiser/persona.py
Normal file
|
|
@ -0,0 +1,52 @@
|
||||||
|
"""Loader for the shared expert-reviewer persona artifact (målbilde §8).
|
||||||
|
|
||||||
|
The persona lives in ``shared/skills/expert-reviewer/`` as a framework-neutral Agent Skill (SKILL.md
|
||||||
|
+ a canonical example verdict). ``shared/`` stays pure DATA — this loader is the MAF-side reader of
|
||||||
|
it; the Claude-SDK sibling reads the same JSON with its own loader. This is what de-stubs the
|
||||||
|
offline simulation: its persona judgement is sourced from the artifact instead of a hardcoded
|
||||||
|
literal, so the shared persona is genuinely consumed (and cannot rot silently).
|
||||||
|
|
||||||
|
Fail-fast on purpose: the example is REQUIRED input (contrast the tolerant async verdict inbox,
|
||||||
|
``verdicts.load_verdicts_from_dir``). A missing or malformed file raises rather than degrading —
|
||||||
|
a broken shared artifact must surface, not pass silently.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
# Module-global so tests can monkeypatch it; read at CALL time inside the loader (never frozen into
|
||||||
|
# a default argument), which is what makes the simulation's persona genuinely artifact-driven.
|
||||||
|
_EXAMPLE_PATH = (
|
||||||
|
Path(__file__).resolve().parents[2]
|
||||||
|
/ "shared"
|
||||||
|
/ "skills"
|
||||||
|
/ "expert-reviewer"
|
||||||
|
/ "references"
|
||||||
|
/ "example-verdict.json"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class PersonaExample:
|
||||||
|
"""The expert-reviewer persona's canonical example verdict: the judgement the persona renders for
|
||||||
|
the reference measure. ``marker`` is a substring of ``rationale`` (the realization-rate payload
|
||||||
|
the simulation traces across runs)."""
|
||||||
|
|
||||||
|
decision: str
|
||||||
|
rationale: str
|
||||||
|
marker: str
|
||||||
|
|
||||||
|
|
||||||
|
def load_persona_example() -> PersonaExample:
|
||||||
|
"""Read the persona's canonical example verdict. Fail-fast: a missing file raises
|
||||||
|
``FileNotFoundError`` and a missing key raises ``KeyError`` (required input). Reads
|
||||||
|
``_EXAMPLE_PATH`` at call time so the path is patchable."""
|
||||||
|
data = json.loads(_EXAMPLE_PATH.read_text(encoding="utf-8"))
|
||||||
|
return PersonaExample(
|
||||||
|
decision=data["decision"],
|
||||||
|
rationale=data["rationale"],
|
||||||
|
marker=data["marker"],
|
||||||
|
)
|
||||||
|
|
@ -36,6 +36,7 @@ from agent_framework import (
|
||||||
)
|
)
|
||||||
from agent_framework_openai import OpenAIChatCompletionClient
|
from agent_framework_openai import OpenAIChatCompletionClient
|
||||||
|
|
||||||
|
from portfolio_optimiser.persona import load_persona_example
|
||||||
from portfolio_optimiser.run import RunResult, run_project
|
from portfolio_optimiser.run import RunResult, run_project
|
||||||
from portfolio_optimiser.validator import ValidatedProposal
|
from portfolio_optimiser.validator import ValidatedProposal
|
||||||
from portfolio_optimiser.verdicts import VerdictStore, promote_verdict, seed_store_from_bundle
|
from portfolio_optimiser.verdicts import VerdictStore, promote_verdict, seed_store_from_bundle
|
||||||
|
|
@ -52,14 +53,10 @@ _VALID_PROPOSAL = (
|
||||||
# The checker's debate turn ends with the gate marker the run parses (run._checker_verdict).
|
# The checker's debate turn ends with the gate marker the run parses (run._checker_verdict).
|
||||||
_CHECKER_APPROVE = "Tallene er innenfor feasibelt område og resonnementet holder. VERDICT: APPROVE"
|
_CHECKER_APPROVE = "Tallene er innenfor feasibelt område og resonnementet holder. VERDICT: APPROVE"
|
||||||
|
|
||||||
# The persona's verdict: an APPROVE that carries NEW realization knowledge the validator cannot
|
# The persona's verdict is sourced from the shared expert-reviewer skill (``load_persona_example``),
|
||||||
# compute. The marker (a realization rate ABSENT from the bundle — the seed is 0.82) is the payload
|
# NOT inlined here — that de-stubs the persona and makes the shared artifact genuinely consumed. Its
|
||||||
# we trace from Run A's persona judgement, through promotion, into Run B's hypothesis prompt.
|
# marker (a realization rate ABSENT from the bundle — the seed is 0.82) is the payload we trace from
|
||||||
_DEFAULT_MARKER = "realiseringsgrad=0.79"
|
# Run A's persona judgement, through promotion, into Run B's hypothesis prompt.
|
||||||
_DEFAULT_PERSONA_RATIONALE = (
|
|
||||||
"Godkjent. I drift realiseres erfaringsvis ~79% av en timeplan-stipulert LED-besparelse i "
|
|
||||||
f"kontorbygg ({_DEFAULT_MARKER}) pga overestimerte driftstimer; forventet faktisk ca 23700 NOK/aar."
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
class ScriptedChatClient(OpenAIChatCompletionClient):
|
class ScriptedChatClient(OpenAIChatCompletionClient):
|
||||||
|
|
@ -148,19 +145,29 @@ async def simulate_learning_loop(
|
||||||
bundle_dir: str,
|
bundle_dir: str,
|
||||||
work_dir: str,
|
work_dir: str,
|
||||||
*,
|
*,
|
||||||
persona_rationale: str = _DEFAULT_PERSONA_RATIONALE,
|
persona_rationale: str | None = None,
|
||||||
marker: str = _DEFAULT_MARKER,
|
marker: str | None = None,
|
||||||
timestamp: str = "2026-06-30",
|
timestamp: str = "2026-06-30",
|
||||||
max_rounds: int = 3,
|
max_rounds: int = 3,
|
||||||
) -> LearningSimulationResult:
|
) -> LearningSimulationResult:
|
||||||
"""Run the loop twice on a throwaway COPY of the bundle (the shared fixture is never mutated),
|
"""Run the loop twice on a throwaway COPY of the bundle (the shared fixture is never mutated),
|
||||||
with a promotion in between, and trace whether the persona's approved knowledge crosses runs.
|
with a promotion in between, and trace whether the persona's approved knowledge crosses runs.
|
||||||
|
|
||||||
|
The persona's verdict (decision + rationale + traced ``marker``) defaults to the shared
|
||||||
|
expert-reviewer skill's canonical example (``load_persona_example``), read at CALL time — so the
|
||||||
|
simulation is genuinely artifact-driven, not inlined. Callers may override ``marker`` /
|
||||||
|
``persona_rationale`` for a control.
|
||||||
|
|
||||||
Run A: a fresh (empty) wiki -> an uninformed hypothesis; the persona approves with NEW realization
|
Run A: a fresh (empty) wiki -> an uninformed hypothesis; the persona approves with NEW realization
|
||||||
knowledge (``marker`` in ``persona_rationale``). ``promote_verdict`` lifts that verdict into the
|
knowledge (``marker`` in ``persona_rationale``). ``promote_verdict`` lifts that verdict into the
|
||||||
wiki; ``seed_store_from_bundle`` re-reads the wiki; Run B's Step-1 ExpeL fold then carries the
|
wiki; ``seed_store_from_bundle`` re-reads the wiki; Run B's Step-1 ExpeL fold then carries the
|
||||||
marker into its hypothesis prompt. The two runs use SEPARATE sinks so each prompt set is
|
marker into its hypothesis prompt. The two runs use SEPARATE sinks so each prompt set is
|
||||||
inspected independently."""
|
inspected independently."""
|
||||||
|
example = load_persona_example()
|
||||||
|
if marker is None:
|
||||||
|
marker = example.marker
|
||||||
|
if persona_rationale is None:
|
||||||
|
persona_rationale = example.rationale
|
||||||
if marker not in persona_rationale:
|
if marker not in persona_rationale:
|
||||||
raise ValueError("marker must be a substring of persona_rationale (the carried payload)")
|
raise ValueError("marker must be a substring of persona_rationale (the carried payload)")
|
||||||
|
|
||||||
|
|
@ -168,7 +175,7 @@ async def simulate_learning_loop(
|
||||||
shutil.copytree(bundle_dir, copy)
|
shutil.copytree(bundle_dir, copy)
|
||||||
copy_s = str(copy)
|
copy_s = str(copy)
|
||||||
replies = {"proposer": _VALID_PROPOSAL, "checker": _CHECKER_APPROVE}
|
replies = {"proposer": _VALID_PROPOSAL, "checker": _CHECKER_APPROVE}
|
||||||
verdict_input = {"decision": "approved", "rationale": persona_rationale}
|
verdict_input = {"decision": example.decision, "rationale": persona_rationale}
|
||||||
|
|
||||||
# Run A — empty wiki isolates the persona's NEW knowledge.
|
# Run A — empty wiki isolates the persona's NEW knowledge.
|
||||||
sink_a: list[str] = []
|
sink_a: list[str] = []
|
||||||
|
|
|
||||||
101
tests/test_persona_skill_loadbearing.py
Normal file
101
tests/test_persona_skill_loadbearing.py
Normal file
|
|
@ -0,0 +1,101 @@
|
||||||
|
"""Load-bearing tests for the shared expert-reviewer persona skill (målbilde §8).
|
||||||
|
|
||||||
|
The persona is a framework-neutral Agent Skill in ``shared/skills/`` encoding the expert reviewer's
|
||||||
|
judgement — the realization gap the deterministic validator cannot compute. These tests make the
|
||||||
|
artifact load-bearing: the REAL pipeline consumes the persona's canonical example verdict, and the
|
||||||
|
offline simulation derives its persona judgement from that artifact, not from a hardcoded literal.
|
||||||
|
Each test goes RED when its seam is detached.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from portfolio_optimiser import okf, persona, verdicts
|
||||||
|
from portfolio_optimiser.contracts import FeedbackContract
|
||||||
|
from portfolio_optimiser.simulation import simulate_learning_loop
|
||||||
|
from portfolio_optimiser.verdicts import bundle_candidate_features, capture_verdict, promote_verdict
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SKILL_DIR = REPO_ROOT / "shared" / "skills" / "expert-reviewer"
|
||||||
|
BUNDLE_DIR = REPO_ROOT / "shared" / "examples" / "bygg-energi-mikro"
|
||||||
|
|
||||||
|
# Import-shaped framework references. Prose mentioning a framework is fine; an import/dependency is
|
||||||
|
# not — this mirrors test_okf_is_maf_free's intent for the shared/skills/ data tree (which has no
|
||||||
|
# AST to parse, so the guard is constrained to import-shaped patterns, not bare substrings).
|
||||||
|
_FORBIDDEN = re.compile(r"\bagent_framework\b|\b(?:import|from)\s+mcp\b")
|
||||||
|
|
||||||
|
|
||||||
|
def test_persona_skill_structure_and_framework_neutral() -> None:
|
||||||
|
"""Test 1: the skill exists with valid non-empty frontmatter, ships its example, and carries no
|
||||||
|
framework import (mirrors test_okf_is_maf_free for shared/skills/). RED before the artifact
|
||||||
|
exists."""
|
||||||
|
skill_md = SKILL_DIR / "SKILL.md"
|
||||||
|
assert skill_md.is_file(), "shared/skills/expert-reviewer/SKILL.md missing"
|
||||||
|
assert (SKILL_DIR / "references" / "example-verdict.json").is_file()
|
||||||
|
|
||||||
|
fm = okf.parse_frontmatter(skill_md)
|
||||||
|
assert fm.get("name") == "expert-reviewer"
|
||||||
|
assert fm.get("description", "").strip(), "SKILL.md description must be non-empty"
|
||||||
|
|
||||||
|
for f in SKILL_DIR.rglob("*"):
|
||||||
|
if f.is_file():
|
||||||
|
assert not _FORBIDDEN.search(f.read_text(encoding="utf-8")), (
|
||||||
|
f"framework import-shaped reference in shared/ persona artifact: {f}"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_persona_example_is_valid_pipeline_input(tmp_path: Path) -> None:
|
||||||
|
"""Test 2: the persona's canonical example verdict is a VALID input to the REAL pipeline —
|
||||||
|
capture_verdict -> promote_verdict accepts it and yields a navigable concept file. Operates on a
|
||||||
|
throwaway COPY (promote_verdict mutates index.md — never touch the git-tracked fixture). RED if
|
||||||
|
the example drifts from the schema the pipeline consumes. Also asserts the marker-in-rationale
|
||||||
|
invariant the simulation requires (simulation.py raises otherwise)."""
|
||||||
|
ex = persona.load_persona_example()
|
||||||
|
assert ex.marker in ex.rationale, "marker must be a substring of rationale (sim invariant)"
|
||||||
|
# The run path feeds verdict_input through FeedbackContract (a BINARY decision); the promotion
|
||||||
|
# gate accepts the same value. The realization correction lives in the rationale, not a third
|
||||||
|
# enum value the run-path contract would reject (contracts.py: Literal["approved", "rejected"]).
|
||||||
|
assert ex.decision == "approved"
|
||||||
|
assert ex.decision in verdicts._APPROVED_DECISIONS
|
||||||
|
FeedbackContract(
|
||||||
|
decision=ex.decision, rationale=ex.rationale
|
||||||
|
) # run-path contract: must not raise
|
||||||
|
|
||||||
|
copy = tmp_path / "bundle"
|
||||||
|
shutil.copytree(BUNDLE_DIR, copy)
|
||||||
|
features = bundle_candidate_features(str(copy))
|
||||||
|
verdict = capture_verdict(features, ex.decision, ex.rationale)
|
||||||
|
promoted = promote_verdict(
|
||||||
|
str(copy),
|
||||||
|
verdict,
|
||||||
|
approver="test",
|
||||||
|
experiment="persona-loadbearing",
|
||||||
|
timestamp="2026-06-30",
|
||||||
|
)
|
||||||
|
assert promoted.exists() and promoted.name.startswith("promoted-verdict-")
|
||||||
|
|
||||||
|
|
||||||
|
async def test_simulation_persona_is_derived_from_artifact(tmp_path: Path, monkeypatch) -> None:
|
||||||
|
"""Test 3 (de-stub causal seam): point the loader at a temp artifact carrying a DIFFERENT marker
|
||||||
|
and confirm the simulation's traced marker follows the FILE — proving the persona judgement is
|
||||||
|
sourced from the artifact, not inlined. RED the moment simulation re-hardcodes the persona
|
||||||
|
(result.marker would stay fixed regardless of the file)."""
|
||||||
|
swapped_marker = "realiseringsgrad=0.41"
|
||||||
|
fake = {
|
||||||
|
"decision": "approved",
|
||||||
|
"marker": swapped_marker,
|
||||||
|
"rationale": f"Godkjent med justering. I drift realiseres erfaringsvis ~41% ({swapped_marker}).",
|
||||||
|
}
|
||||||
|
fake_path = tmp_path / "example-verdict.json"
|
||||||
|
fake_path.write_text(json.dumps(fake), encoding="utf-8")
|
||||||
|
monkeypatch.setattr(persona, "_EXAMPLE_PATH", fake_path)
|
||||||
|
|
||||||
|
result = await simulate_learning_loop(str(BUNDLE_DIR), str(tmp_path / "w"))
|
||||||
|
|
||||||
|
assert result.marker == swapped_marker, "simulation did not source the marker from the artifact"
|
||||||
|
assert not result.marker_in_run_a_prompt, "swapped marker leaked into Run A (causality broken)"
|
||||||
|
assert result.marker_in_run_b_prompt, "the artifact's marker did not cross into Run B's prompt"
|
||||||
Loading…
Add table
Add a link
Reference in a new issue