Offline structure tests for evals/ (the suite itself needs headless runs):
- no case reaches the steg 1 intent gate (PM: 0 traces touched it). The
new case's scaffold is in: its brief passes brief-validator --soft
(2.1 WITH phase_signals) and --check gives BRIEF_INTENT_NOT_APPROVED —
that test is green, so the fixture reaches the gate.
- review-requires-project: trekreview.md composes 'Error: --project <dir>
is required.' in prose (8/10 in the PM run, backticks broke the regex),
and its arg-parser line passes "$@", which the Bash tool never has.
- no-error-code misses REVIEW_WRONG_TYPE; PASS/FAIL graders are raw
substrings; no-write cannot see a write through Bash.
7 tests, 6 red.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Seven `claude plugin eval` cases under evals/, runs: 1, deterministic graders
only (regex, tool_used, file_exists), no ablation. Each case tests one thing
Voyage promises and stops within minutes, headless:
- plan-requires-brief, plan-project-not-initialized, plan-rejects-unknown-export,
review-requires-project: argument guards stop before any Agent or Write.
- plan-halts-without-phase-signals: a brief_version 2.1 brief without
phase_signals halts /trekplan at the sequencing gate, before the swarm.
- review-validate-flags-bad-finding-id (known-positive) and
review-validate-passes-clean-review (known-negative): /trekreview --validate
names REVIEW_BAD_FINDING_ID on a planted bad ID and stays clean on a valid file.
Each case's expected_outcome is committed here, before the suite has run.
Fixtures are written by an inline scaffold.sh (needs --scaffold): chose inline
heredocs because a run cannot read the eval directory, and the add_dirs path
mapping is not documented. evals/results/ is ignored.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>