New § Stability section: the PM's 8/8 and 7/8 onea66a98, the cause (correct lower-case verdicts, 3 of 34 case-8 and 2 of 22 case-6 recorded replies), the grader changes and why, and the measured result on7f7161b: case 8 12 of 12 (10 alone + 2 full), case 6 8 of 8 (6 + 2), two full runs 8 of 8 each. The status block and "every full run gave 8 of 8" now name whose sample that was. Correction to 687710a's message: it said a grader file carries only `pattern`. The docs list `flags` for the regex grader. The letter classes stay, for the reason now written in the doc: a flag applies to the whole pattern, and the opposite-verdict refusal must stay case-sensitive. Suite after git add: 1236 tests, 1234/0/2. No version bump: plugin code is untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
20 KiB
Prøvesett — Voyage's fixed test set as a plugin eval suite
Status: v2, eight cases. v1 (seven cases) first run 6 of 7; v2 first run 8 of 8 (both 2026-09-23, Claude Code 2.1.280). Each version's cases and expectations were committed before its first run; the run did not change them. Red is data here, not a defect in the suite. What v2 changed and why: § v2 below.
v2's first 8 of 8 was one sample, not a stable state. Measured the same day, before the stability fix: case 4 passed 5 of 6 here, 6 of 12 and 7 of 10 in the PM re-measurement. After the fix, the builder's sample gave every case 5 of 5 over three full runs. The PM re-measurement of that commit gave 8 of 8 and 7 of 8: case 8 failed on correct replies written in lower case. After the verdict-grader fix, case 8 12 of 12, case 6 8 of 8, two full runs 8 of 8 each. Numbers per case, with the denominator: § Stability.
The unit suite (npm test) proves Voyage's modules in isolation. This suite proves
something the unit suite cannot: that the commands, as a model reads and follows
them in a real headless session, keep the promises they make. Each case is a real
claude -p child with only Voyage loaded.
What "pass" and "fail" mean (tool definition, verbatim)
Source: Claude Code docs, Test plugins with evals (code.claude.com/docs/en/plugin-evals),
§ How a case is scored and § Grader types.
A run's score is the fraction of its graders that passed, weighted if you set weights, and the case's score is the mean across its runs. A case passes when its score meets the
--threshold,1.0by default.
The four graders this suite uses, and when each passes:
| Grader | Passes when (verbatim) |
|---|---|
regex |
"The JavaScript regex pattern is found in the target. Set match: not_contains to require absence" |
tool_used |
"The number of calls to tool whose JSON-encoded input matches the optional input_match regex is between min, default 1, and max, default unlimited. To assert a tool was never called, set both min: 0 and max: 0" |
file_exists |
"A file Claude created matches the path glob, or none does with exists: false. Only files created during the run count" |
regex looks at last_message ("Claude's final response text. This is the default")
unless the case sets target: trace ("The session as JSON, one message per line").
With runs: 1 and the default threshold, a case passes only when every grader
passes, and fails when any single one does not. n of N below counts cases, not graders.
The cases
Every case has runs: 1, deterministic graders only (no llm, no baseline — they cost
judge calls), and runs without the no-plugin baseline arm. Each tests one thing Voyage
promises and ends in seconds. Nothing starts the /trekplan exploration swarm or the
/trekreview reviewer agents; the no-agent grader in every case asserts it.
| Case | Voyage promise under test | Expected (committed before the run) | Run 1 |
|---|---|---|---|
plan-requires-brief |
/trekplan without a brief prints usage and stops |
"A brief is required"; no Agent, no Write | pass |
plan-project-not-initialized |
/trekplan --project on a missing folder stops before exploring |
"project directory not initialized"; no Agent, no plan.md |
pass |
plan-rejects-unknown-export |
a removed export format is named and refused | "format 'pr' is not supported"; no Agent, no Write | pass |
plan-halts-without-phase-signals |
a brief_version 2.1 brief without phase_signals halts at the sequencing gate |
brief-validator runs, reply names phase_signals; no Agent, no plan.md |
pass |
review-requires-project |
/trekreview without --project refuses |
"--project is required"; no Agent, no Write | fail (2 of 3 graders) |
review-validate-flags-bad-finding-id |
known-positive: --validate fails a review whose finding ID is not 40-char hex, and names the rule |
review-validator runs; REVIEW_BAD_FINDING_ID in the trace; reply says FAIL; no Agent, no Write |
pass |
review-validate-passes-clean-review |
known-negative: --validate passes a valid review and raises nothing |
review-validator runs; reply says PASS; no REVIEW_* error code in the trace; no Agent, no Write |
pass |
The v1 table above is kept as it was run. v2 rows and grader changes: § v2.
The one red case (v1; fixed in v2, see § v2)
review-requires-project failed on its text grader only: the child replied in one turn
without reproducing the exact error line Error: --project <dir> is required. The
Agent and Write guards held. A diagnostic re-run of the same case alone passed, with the
exact line. So the case is unstable, not broken. It is still red on purpose. This
guard lives only in the command's prose, and a model does not always repeat prose word
for word. Loosening the pattern after seeing the result would fit the grader to the run.
If this should hold every time, the fix belongs in Voyage: let code emit the message, as
the validators already emit their rule codes (REVIEW_BAD_FINDING_ID is printed by
review-validator.mjs, not composed by the model — that is why names-rule greps the
trace, not the reply).
Known-positive, demonstrated
Proof that the known-positive case can fail: in a scratch copy of Voyage (never this
tree), the finding-ID check in lib/validators/review-validator.mjs was disabled. The
validator then accepted the planted F-001 ID, and
review-validate-flags-bad-finding-id dropped from 1.0 to 0.6. names-rule and
says-fail failed; validator-ran, no-agent and no-write still passed. The case
therefore falls when this part of Voyage breaks.
First run, recorded
| Result | 6 of 7 cases passed (graders 25 of 26; overall score 0.952 = mean of the case scores) |
| Children | 7 claude -p runs (7 cases × 1 run × 1 arm) |
| Wall time | 54 s (the tool reports durationSeconds: 53) |
| Cost | $1.67, the tool's list-price estimate; on a subscription it counts against usage limits instead |
| Output | evals/results/<timestamp>/aggregate-result.json + report.html (gitignored) |
v2 (2026-09-23): the intent gate gets a case, case 5's error moves into code
Expectations committed in 3892a83 before the v2 run. Offline structure tests in
tests/lib/proevesett.test.mjs pin what each grader can see.
| Change | Why |
|---|---|
New case plan-halts-without-intent-approval: a brief_version 2.1 brief with phase_signals and no intent_approved_hash. Graders: gate-ran (intent-approval.mjs --check in a Bash call), names-code (the gate's own JSON code for BRIEF_INTENT_NOT_APPROVED in the trace), no-agent, no-plan-file |
Veikart steg 1's gate had no case. No v1 trace reached it: case 4 halts at the sequencing gate first |
review-requires-project: the arg parser got a CLI that checks --project itself and prints Error: --project <dir> is required. + usage (exit 1, code ARG_REQUIRED_MISSING). trekreview.md runs it with $ARGUMENTS and relays its stderr. The reply-regex grader project-required is replaced by parser-ran + names-missing-project (the parser's code in the trace, the same form as names-rule) |
The fix this doc recommended for the red case. Also: the prose passed "$@", which the Bash tool never has. The parser never ran (0 Bash calls in the case-5 traces), so the model composed the line from prose. That failed 2 of 10 on backticks |
no-error-code → code\W{1,12}REVIEW_[A-Z_]+, bound to the validator's JSON output |
The listed codes missed REVIEW_WRONG_TYPE and REVIEW_VERSION_FORMAT |
says-pass / says-fail reject a preceding NOT and a longer word |
Raw substrings: "did not PASS" passed says-pass |
no-bash-write beside every no-write: redirect, tee, touch, cp, mv, rm, sed -i in the Bash command |
no-write sees only the Write tool |
All new patterns were run against the ten traces of the PM re-measurement. There were no false positives, and the known-positive and known-negative still split.
v2 first run: 8 of 8 (graders 36 of 36, 65 s, $1.97; the tool's list-price estimate). That was one run. It did not hold: see § Stability.
Known-positive for the new case, demonstrated in a scratch copy (never this tree):
- With
--checkmade to accept every brief (intent-approval.mjs), the case dropped to 0.75.names-codefailed; the other graders passed. - Removing only the gate's command line from
trekplan.mddid not fell the case (1.00). The child rebuilt the--checkcall from the prose around it. So the case measures whether the gate works, not whether that exact line is there. The line is pinned bytests/validators/intent-approval.test.mjs, which executes it.
Stability (2026-09-23): case 4 measures its own gate
What was unstable. The PM re-measurement of v2 got 6 of 8 and 7 of 8 in two full
runs. Case 4 plan-halts-without-phase-signals passed 6 of 12 on the v2 commit
and 7 of 10 on the commit before it. The failure was the same every time. The
case-4 brief had no intent marker, so the child sometimes ran the intent gate first
and halted on BRIEF_INTENT_NOT_APPROVED. That halt was correct, but the brief-validator
never ran, so validator-ran and halt-names-signals failed. The case measured the
wrong gate. Re-measured here before the fix: 5 of 6, and the one failure had the
same trace.
| Change | Why |
|---|---|
The case-4 scaffold writes a current intent_approved_hash into the brief |
The intent gate now passes, so only the sequencing gate can stop the case. The marker is written literally, not by intent-approval.mjs --stamp: --stamp emits a brief-approved record into the live stats that the yardstick reads |
New case-4 grader no-intent-halt (trace, not_contains, any BRIEF_INTENT_* code) |
If the marker ever goes stale, the case fails and names the reason. A halt on the intent gate no longer counts as a pass |
New case-4 grader names-code (the validator's BRIEF_V51_MISSING_SIGNALS code in the trace) |
The known-positive below: without it, the case passed with the gate's code switched off |
says-fail / says-pass take FAILED / PASSED as the verdict word |
v2's word boundary rejected them. v1's raw substring accepted them. A correct reply must not fail the case |
Known-positive for case 4, demonstrated in a scratch copy (never this tree). The
fault planted is the one the case exists to catch: the BRIEF_V51_MISSING_SIGNALS
check in brief-validator.mjs was made if (false).
- Without
names-code, the case still passed 3 of 3. The child read the rule in the command prose, halted, namedphase_signals, and even pointed at the plantedif (false). A correct halt does not show that the gate's code fired. - With
names-code, it fell 0 of 3.names-codefailed each time, and the other five graders passed. - Before committing, the pattern was checked on this session's recorded traces. It matched 10 of 10 green runs and 0 of 3 fault runs.
says-fail (case 6), cause not measured. The PM saw 1 failure in 7 runs on
says-fail. That run's trace was deleted, so the reply cannot be read. The
hypothesis was that the reply said FAILED. It is not confirmed: 0 of 13 kept
replies (8 here, 5 from the PM) used any spelling but FAIL. What is certain is that the
v2 pattern rejects FAILED, and a reply saying FAILED is correct. So the pattern is
widened, and the failure stays unexplained. If it comes back, it is not this.
Measured stability after the fix (a384ce6, Claude Code 2.1.280). Full run =
all eight cases, the command under § How to run it. There were three: one with
--runs 3 and two with --runs 1. That gives 5 children per case.
| Case | Full runs | Alone | Total |
|---|---|---|---|
plan-halts-without-intent-approval |
5 of 5 | – | 5 of 5 |
plan-halts-without-phase-signals |
5 of 5 | 10 of 10 | 15 of 15 |
plan-project-not-initialized |
5 of 5 | – | 5 of 5 |
plan-rejects-unknown-export |
5 of 5 | – | 5 of 5 |
plan-requires-brief |
5 of 5 | – | 5 of 5 |
review-requires-project |
5 of 5 | – | 5 of 5 |
review-validate-flags-bad-finding-id |
5 of 5 | 8 of 8 (v2 grader, before widening) | 13 of 13 |
review-validate-passes-clean-review |
5 of 5 | – | 5 of 5 |
In 4 of the 15 case-4 runs, the child ran the intent gate as its own call before the brief-validator. In 1 more, it ran both gates in one call. All 5 passed the intent gate, went on to the sequencing gate and passed the case. Before the fix, a run that ran the intent gate alone first halted there. Every full run in this sample gave 8 of 8; the PM's two did not (next section). A pass rate of 15 of 15 does not prove the rate is 1: it bounds it. Five runs per case is a small denominator. The claim is "no failure in N", not "cannot fail".
Stability (2026-09-23): the verdict graders read any letter case
What failed. The PM re-measurement of ea66a98 (the commit above) gave 8 of 8
and 7 of 8 in two full runs. Case 8 review-validate-passes-clean-review passed
10 of 12 (9 of 10 alone, 1 of 2 in full runs). Both failures were correct replies:
the validator ran and exited 0, and the reply said "proj/review.md passed the strict
schema check". Only says-pass failed. Its pattern \bPASS(?:ED)?\b is
case-sensitive, and the reply had no PASS. The fault is older than § Stability's
fix: v2's \bPASS\b was stricter still.
Measured, not assumed. The children write the verdict in lower case, in both cases. Over all recorded replies (the PM's and this change's own runs), the old patterns would have failed 3 of 34 correct case-8 replies ("passed", "passes") and 2 of 22 correct case-6 replies ("fails", "failed"). That also explains the class behind the unexplained case-6 failure above. That one reply is still gone, so it is not confirmed.
| Change | Why |
|---|---|
The verdict word in any letter case: [Pp][Aa][Ss][Ss], with the forms pass/passes/passed and fail/fails/failed |
Correct replies used each of these |
A preceding n't refuses, like not already did |
"doesn't pass" is the wrong verdict |
The whole reply is refused if the opposite verdict stands in capitals (FAIL/FAILED for says-pass, PASS/PASSED for says-fail), unless negated |
A correct case-6 reply said "FAIL: … did not pass … The other checks passed." Accepting lower case alone would let a wrong case-8 reply of that shape pass |
| The opposite word in lower case does not refuse | Correct replies carry it in side clauses ("did not pass", "other checks passed"). Refusing it would bring back the instability this change removes |
Letter classes, not flags: i |
The grader accepts flags (docs, § Grader types), but a flag applies to the whole pattern, and the refusal must stay case-sensitive. The pattern starts with (?<![\s\S]), so it anchors at the start of the reply with or without m |
On the recorded replies, the new patterns take 34 of 34 (case 8) and 22 of 22
(case 6), and each grader matches 0 of the other case's replies. Pinned offline in
tests/lib/proevesett.test.mjs, red first on the commit before each change.
Measured stability after the fix (Claude Code 2.1.280, same command as § How to run it):
| Case | Full runs | Alone | Total |
|---|---|---|---|
review-validate-passes-clean-review |
2 of 2 | 10 of 10 | 12 of 12 |
review-validate-flags-bad-finding-id |
2 of 2 | 6 of 6 | 8 of 8 |
| the six other cases | 2 of 2 each | – | 2 of 2 each |
Both full runs gave 8 of 8. In these 20 runs of cases 6 and 8, 1 reply gave its verdict only in lower case ("failed", case 6, full run 2). Before the final fix, an intermediate pattern (lower case only, no refusal) had case 8 10 of 10 and case 6 6 of 6 alone, with 2 lower-case-only replies. Those are not counted above. Eval cost of this change: $10.85 (list-price estimate).
What the numbers allow. Case 8 has 12 of 12 on this commit. Across all its
recorded replies, 34 of 34 would pass the new grader. That bounds the grader fault.
It does not show that the case cannot fail. With runs: 1, step 4 still reads one
sample per case per run.
Measured facts about the harness this suite relies on
These were measured by probe cases in a scratch copy, not taken from the docs.
- Commands are invocable in the child as
/voyage:<command>. The command body is expanded into the prompt; noSkilltool call is involved. ${CLAUDE_PLUGIN_ROOT}is substituted in the expanded command text, so thenode ${CLAUDE_PLUGIN_ROOT}/lib/...lines in the commands reach the real validators. In the child's Bash environment the variable is empty. A case prompt must never rely on it directly.- The plugin directory is readable from the child's OS sandbox, so Voyage's
validators run for real under
--allow-tools Bash. - Fixtures come from an inline
scaffold.shthat writes them with heredocs. A run starts in an empty workspace and "can't read the eval directory". The docs do not say how anadd_dirspath maps into the run, so the suite does not depend on it.
What the suite does NOT cover
- The agent swarms. It never tests the
/trekplanexploration, synthesis or review phases, the/trekreviewreviewers and coordinator, or/trekresearch. Every case asserts zero Agent calls. Whether the swarm plans or reviews well is outside this suite. The offline gold-scored eval (tests/lib/gold-eval.test.mjs, seedocs/eval-corpus/README.md) covers the coordinator contract on recorded payloads. /trekbrief, because its interview needsAskUserQuestion, which a headless child does not have./trekexecute,/trekcontinueand/trekendsessionhave no case yet.- Plugin contribution. The suite runs without the no-plugin baseline (
--ablation none), so it does not measure how much Voyage adds over plain Claude. It measures only whether Voyage keeps its own promises. - Variance. With
runs: 1, a single run can pass or fail by chance; the red case above shows exactly that. One result is a sample, not a rate. - Semantic quality of any output. Every grader is a pattern, a tool count or a file glob.
How to run it
From the plugin root:
claude plugin eval . --trust-plugin --ablation none --runs 1 --scaffold --no-publish --allow-tools Bash Write
--scaffoldruns each case'sscaffold.shas you, outside the sandbox. The four scaffolds in this suite only write fixture files into the run's workspace.--allow-tools Bash Writegrants the tools the commands need (Bash runs under the OS sandbox). Write is granted so that theno-writeandno-plan-filegraders measure a refusal, not a missing permission.Agentis listed in every case'sallowed_toolsfor the same reason.- The exit code is 1 whenever any case scores below 1.0.
How veikart step 4 ("Trygg å endre") uses it
The suite is the fixed yardstick that a rule change is measured against:
- Run the suite on the base commit and record
n of N(aggregates.casesPassed/aggregates.casesTotalinaggregate-result.json). - Run it again with the change applied.
- If
nis smaller, the change stops. A lowernis not argued away. It is cleared only by showing that the same case is also red on the base commit, measured the same day. Otherwise the change broke a promise. - A new promise gets a new case, and its expectation is committed before its first run,
the same as here.
Ngrows; it never shrinks to makenlook better.
Because runs: 1 is a sample, an unstable case such as review-requires-project in v1
will sometimes stop a change that is innocent. That is the intended price. Moving the guard
into code removes the instability; editing the grader does not. v2 did the move: the
parser now writes the line, and the grader reads the parser's code. One green run does
not show the instability is gone. That takes more runs. § Stability has them:
review-requires-project 5 of 5, and case 4 15 of 15 after its own fix.