test(evals): Voyage's fixed test set as a plugin eval suite, expectations written before the first run
Seven `claude plugin eval` cases under evals/, runs: 1, deterministic graders only (regex, tool_used, file_exists), no ablation. Each case tests one thing Voyage promises and stops within minutes, headless: - plan-requires-brief, plan-project-not-initialized, plan-rejects-unknown-export, review-requires-project: argument guards stop before any Agent or Write. - plan-halts-without-phase-signals: a brief_version 2.1 brief without phase_signals halts /trekplan at the sequencing gate, before the swarm. - review-validate-flags-bad-finding-id (known-positive) and review-validate-passes-clean-review (known-negative): /trekreview --validate names REVIEW_BAD_FINDING_ID on a planted bad ID and stays clean on a valid file. Each case's expected_outcome is committed here, before the suite has run. Fixtures are written by an inline scaffold.sh (needs --scaffold): chose inline heredocs because a run cannot read the eval directory, and the add_dirs path mapping is not documented. evals/results/ is ignored. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
parent
1779411b49
commit
d311f3b49a
40 changed files with 422 additions and 0 deletions
4
.gitignore
vendored
4
.gitignore
vendored
|
|
@ -9,6 +9,10 @@ test-results/
|
|||
playwright-report/
|
||||
blob-report/
|
||||
|
||||
# `claude plugin eval` run output (aggregate-result.json, report.html). Results are
|
||||
# data about one run on one machine; the suite (evals/<case>/) is what is tracked.
|
||||
/evals/results/
|
||||
|
||||
# Editor files
|
||||
*.swp
|
||||
*.swo
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue