Locks the ground-truth control BEFORE running /trekplan, so Q4 (does the adversarial review catch real bugs?) is scored as recall against a fixed target, not post-hoc. Feature under test: voyage-doctor (project-coherence validator). Expected plan + 7 pre-registered real risks (R1-R7) + scoring rubric. Input brief lives in gitignored .claude/projects/; its SC/NG are embedded verbatim in this doc so the control is reproducible. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
7.4 KiB
S22 — Happy-path dogfood (Blind spot #1 + #4)
Run: S22, 2026-06-19. Method: dogfood Voyage's own pipeline (/trekplan → /trekexecute) on a real, small Voyage feature, scored against a pre-registered ground-truth scorecard committed before the pipeline runs.
Answers the two open questions the S14 audit named but never measured (devils-advocate-results.md §"What this audit might have missed" #1 and #4):
- Q1 — plan quality: does
/trekplanproduce a correct, useful plan on a real feature? - Q4 — review efficacy: does the adversarial review (
plan-critic10-dim +scope-guardian) catch real defects?
Operator-chosen shape (S22): real new Voyage feature (not a planted defect, not a known backlog bug); execute runs in an isolated git worktree and the code is discarded — the deliverable is the measurement, not the feature.
Methodology — why pre-registration
The audit's deepest critique of itself (#4) was that nobody ever measured whether the review catches real bugs — both attack and defense assumed it. The failure mode for a dogfood is post-hoc rationalization: run the pipeline, then declare whatever it found "the important stuff." To avoid that, the expected plan and the real-risk list below are written and git-committed before /trekplan is launched (provable by commit order). Q4 is then scored as recall against a fixed target, not "did it find something."
The feature (voyage-doctor) was chosen because it has three properties that make Q4 measurable on an honest feature:
- Strong reuse anchors (
discoverProject,validateBrief/Research/Plan) → a naive plan reimplements parsing; a good plan composes. Tests whetherscope-guardian/plan-criticflag reinvention. - A subtle state-conditional correctness rule (
research_status↔research/contents) → easy to get wrong. Tests whetherplan-criticcatches a wrong conditional. - Explicit Non-Goals (read-only, no auto-fix) with a Context that mildly invites scope-creep → tests whether the review verifies Non-Goal coverage.
Input brief: .claude/projects/2026-06-19-voyage-doctor/brief.md (gitignored per repo convention; Success Criteria + Non-Goals reproduced verbatim below so the control is locked). Brief validates clean: brief-validator.mjs --json → {valid:true, errors:[], warnings:[]}.
Brief — Success Criteria (verbatim, locked)
- SC1 — New
lib/validators/project-doctor.mjsexports a function returning{valid, errors, warnings}, building ondiscoverProject()+ the per-artifact content validators rather than re-parsing files itself. - SC2 — Detects ≥3 problem classes: (i) present-but-invalid artifact (surface underlying validator findings, tagged by artifact), (ii)
research_status: completewith emptyresearch/, (iii) project-dir slug ≠ briefslugfrontmatter. - SC3 — CLI (
node lib/validators/project-doctor.mjs <projectDir> [--json]): human report default,--json, non-zero exit on invalid — consistent withbrief-validator.mjs. - SC4 —
node:testcoverage for the three SC2 coherence branches + a test asserting delegation to existing validators (no re-implementation).
Brief — Non-Goals (verbatim, locked)
- NG1 — No auto-fix, no mutation; strictly read-only.
- NG2 — No new artifact types, no schema changes to brief/research/plan.
- NG3 — No network / external calls.
- NG4 — Not a replacement for the per-artifact validators or
checkPhaseRequirements(); composes, does not supersede.
Pre-registered ground truth (LOCKED before /trekplan)
A. Expected plan (Q1 oracle)
A competent plan should, at minimum:
- Reuse, not reimplement — call
discoverProject(dir)for the artifact set, thenvalidateBrief/validateResearch(Dir)/validatePlanon the present artifacts. No new frontmatter/markdown parsing. - Aggregate findings into one
{valid, errors[], warnings[]}using the existingissue()/result.mjshelpers, tagging each finding with the artifact it came from. - Implement the 3 coherence checks of SC2 with correct conditionals (see real-risks R2 below).
- Add a CLI shim copying
brief-validator.mjs'simport.meta.urlpattern (human +--json+ exit 0/1, usage → exit 2). - TDD — node:test cases for each coherence branch + a delegation test; fixtures = throwaway project dirs.
- Stay read-only — no writes, honoring NG1.
- Decompose into a small number of steps (module → checks → CLI → tests), each independently testable.
A plan scores well on Q1 if it hits 1–6 without inventing scope beyond the brief.
B. Real risks / defects the review SHOULD catch (Q4 oracle — fixed target)
Each risk is something a flawed plan could plausibly contain. Scoring records, for each: did the plan avoid it, and if the plan tripped it, did plan-critic or scope-guardian flag it.
| ID | Risk | Which reviewer should catch it if the plan trips it |
|---|---|---|
| R1 | Plan reimplements brief/research/plan parsing instead of reusing the validators (duplication, drift). | scope-guardian (reuse gap) / plan-critic (maintainability) |
| R2 | The research_status ↔ research/ conditional is wrong: e.g. flags empty research/ even when research_topics: 0 / skipped (false positive), or misses complete+empty (false negative). |
plan-critic (correctness / edge cases) |
| R3 | No error isolation: a malformed/unreadable artifact throws and aborts the whole doctor instead of becoming a finding. | plan-critic (error handling / robustness) |
| R4 | Missing-artifact vs present-but-invalid not distinguished (both collapse to one code), so the report is ambiguous. | plan-critic (correctness) |
| R5 | Slug-from-dirname parse is naive: breaks on slugs containing hyphens or dirs lacking the YYYY-MM-DD- prefix. |
plan-critic (edge cases) |
| R6 | Scope-creep against NG1: plan adds auto-fix / "repair" / writing a report file. | scope-guardian (creep vs Non-Goal) |
| R7 | Tests assert only the happy path; the SC2 conditional branches (esp. R2's false-positive case) are untested. | plan-critic (test coverage) / test-strategist |
Q4 score = (real risks correctly handled by the plan) + (risks the plan tripped that the review flagged) / total applicable. A risk the plan handles correctly is not counted against the review (nothing to catch) but is recorded as "plan avoided it." The review's job is the residual: of the risks the plan got wrong, how many did it surface?
C. Scoring rubric
- Q1 (plan quality): PASS / PARTIAL / FAIL against expected-plan items 1–6, plus a one-line qualitative verdict. Independently judged by main context reading the produced
plan.md. - Q4 (review efficacy): for each Ri —
plan: avoided | tripped, and if trippedreview: caught | missed. Headline = review recall on tripped risks (caught / tripped). Also note any real problems the review raised that are NOT in R1–R7 (true positives outside the pre-registered set → credit) vs. noise/false-positives (debit). - Execute (worktree): did
/trekexecuteproduce code that (a) matches the plan and (b) passes its own tests +node --test? Run in isolated worktree, then discard. Records a yes/partial/no, not a quality grade.
RESULTS
(appended after the run — empty until /trekplan + /trekexecute complete)