The S22 dogfood noted defect #4: the installed plugin cache is at v5.1.1 while the repo under development is v5.5.0, so operators dogfooding the installed skill are not exercising the dev tree. S27 assessed whether this is a code task or an operator note. Verified the skew is still live — ~/.claude/plugins/cache/ktg-plugin-marketplace/voyage/ 5.1.1/.claude-plugin/plugin.json = 5.1.1 vs repo .claude-plugin/plugin.json = 5.5.0 — but it is an operator/cache state, not a Voyage code defect: the 5.5.0 source tree is correct, and /trekplan resolving to the cached install rather than the dev tree is expected Claude Code plugin-cache behavior. The only remediation is an environment action (refresh the cache via /plugin), out of repo-code scope. A runtime version warning was rejected as scope creep — the installed context cannot know a newer dev tree exists without querying the marketplace. Operator-chosen: close no-op with a doc marking. Appended an ASSESSED (S27) note to defect #4 and a verification-log row in docs/S22-happy-path-dogfood.md. No source code or tests changed; suite unchanged at 728 (726/2/0, bare node --test). No doc-consistency pin reads the S22 doc, so nothing to break.
19 KiB
S22 — Happy-path dogfood (Blind spot #1 + #4)
Run: S22, 2026-06-19. Method: dogfood Voyage's own pipeline (/trekplan → /trekexecute) on a real, small Voyage feature, scored against a pre-registered ground-truth scorecard committed before the pipeline runs.
Answers the two open questions the S14 audit named but never measured (devils-advocate-results.md §"What this audit might have missed" #1 and #4):
- Q1 — plan quality: does
/trekplanproduce a correct, useful plan on a real feature? - Q4 — review efficacy: does the adversarial review (
plan-critic10-dim +scope-guardian) catch real defects?
Operator-chosen shape (S22): real new Voyage feature (not a planted defect, not a known backlog bug); execute runs in an isolated git worktree and the code is discarded — the deliverable is the measurement, not the feature.
Methodology — why pre-registration
The audit's deepest critique of itself (#4) was that nobody ever measured whether the review catches real bugs — both attack and defense assumed it. The failure mode for a dogfood is post-hoc rationalization: run the pipeline, then declare whatever it found "the important stuff." To avoid that, the expected plan and the real-risk list below are written and git-committed before /trekplan is launched (provable by commit order). Q4 is then scored as recall against a fixed target, not "did it find something."
The feature (voyage-doctor) was chosen because it has three properties that make Q4 measurable on an honest feature:
- Strong reuse anchors (
discoverProject,validateBrief/Research/Plan) → a naive plan reimplements parsing; a good plan composes. Tests whetherscope-guardian/plan-criticflag reinvention. - A subtle state-conditional correctness rule (
research_status↔research/contents) → easy to get wrong. Tests whetherplan-criticcatches a wrong conditional. - Explicit Non-Goals (read-only, no auto-fix) with a Context that mildly invites scope-creep → tests whether the review verifies Non-Goal coverage.
Input brief: .claude/projects/2026-06-19-voyage-doctor/brief.md (gitignored per repo convention; Success Criteria + Non-Goals reproduced verbatim below so the control is locked). Brief validates clean: brief-validator.mjs --json → {valid:true, errors:[], warnings:[]}.
Brief — Success Criteria (verbatim, locked)
- SC1 — New
lib/validators/project-doctor.mjsexports a function returning{valid, errors, warnings}, building ondiscoverProject()+ the per-artifact content validators rather than re-parsing files itself. - SC2 — Detects ≥3 problem classes: (i) present-but-invalid artifact (surface underlying validator findings, tagged by artifact), (ii)
research_status: completewith emptyresearch/, (iii) project-dir slug ≠ briefslugfrontmatter. - SC3 — CLI (
node lib/validators/project-doctor.mjs <projectDir> [--json]): human report default,--json, non-zero exit on invalid — consistent withbrief-validator.mjs. - SC4 —
node:testcoverage for the three SC2 coherence branches + a test asserting delegation to existing validators (no re-implementation).
Brief — Non-Goals (verbatim, locked)
- NG1 — No auto-fix, no mutation; strictly read-only.
- NG2 — No new artifact types, no schema changes to brief/research/plan.
- NG3 — No network / external calls.
- NG4 — Not a replacement for the per-artifact validators or
checkPhaseRequirements(); composes, does not supersede.
Pre-registered ground truth (LOCKED before /trekplan)
A. Expected plan (Q1 oracle)
A competent plan should, at minimum:
- Reuse, not reimplement — call
discoverProject(dir)for the artifact set, thenvalidateBrief/validateResearch(Dir)/validatePlanon the present artifacts. No new frontmatter/markdown parsing. - Aggregate findings into one
{valid, errors[], warnings[]}using the existingissue()/result.mjshelpers, tagging each finding with the artifact it came from. - Implement the 3 coherence checks of SC2 with correct conditionals (see real-risks R2 below).
- Add a CLI shim copying
brief-validator.mjs'simport.meta.urlpattern (human +--json+ exit 0/1, usage → exit 2). - TDD — node:test cases for each coherence branch + a delegation test; fixtures = throwaway project dirs.
- Stay read-only — no writes, honoring NG1.
- Decompose into a small number of steps (module → checks → CLI → tests), each independently testable.
A plan scores well on Q1 if it hits 1–6 without inventing scope beyond the brief.
B. Real risks / defects the review SHOULD catch (Q4 oracle — fixed target)
Each risk is something a flawed plan could plausibly contain. Scoring records, for each: did the plan avoid it, and if the plan tripped it, did plan-critic or scope-guardian flag it.
| ID | Risk | Which reviewer should catch it if the plan trips it |
|---|---|---|
| R1 | Plan reimplements brief/research/plan parsing instead of reusing the validators (duplication, drift). | scope-guardian (reuse gap) / plan-critic (maintainability) |
| R2 | The research_status ↔ research/ conditional is wrong: e.g. flags empty research/ even when research_topics: 0 / skipped (false positive), or misses complete+empty (false negative). |
plan-critic (correctness / edge cases) |
| R3 | No error isolation: a malformed/unreadable artifact throws and aborts the whole doctor instead of becoming a finding. | plan-critic (error handling / robustness) |
| R4 | Missing-artifact vs present-but-invalid not distinguished (both collapse to one code), so the report is ambiguous. | plan-critic (correctness) |
| R5 | Slug-from-dirname parse is naive: breaks on slugs containing hyphens or dirs lacking the YYYY-MM-DD- prefix. |
plan-critic (edge cases) |
| R6 | Scope-creep against NG1: plan adds auto-fix / "repair" / writing a report file. | scope-guardian (creep vs Non-Goal) |
| R7 | Tests assert only the happy path; the SC2 conditional branches (esp. R2's false-positive case) are untested. | plan-critic (test coverage) / test-strategist |
Q4 score = (real risks correctly handled by the plan) + (risks the plan tripped that the review flagged) / total applicable. A risk the plan handles correctly is not counted against the review (nothing to catch) but is recorded as "plan avoided it." The review's job is the residual: of the risks the plan got wrong, how many did it surface?
C. Scoring rubric
- Q1 (plan quality): PASS / PARTIAL / FAIL against expected-plan items 1–6, plus a one-line qualitative verdict. Independently judged by main context reading the produced
plan.md. - Q4 (review efficacy): for each Ri —
plan: avoided | tripped, and if trippedreview: caught | missed. Headline = review recall on tripped risks (caught / tripped). Also note any real problems the review raised that are NOT in R1–R7 (true positives outside the pre-registered set → credit) vs. noise/false-positives (debit). - Execute (worktree): did
/trekexecuteproduce code that (a) matches the plan and (b) passes its own tests +node --test? Run in isolated worktree, then discard. Records a yes/partial/no, not a quality grade.
RESULTS
Run: 2026-06-19, interactive main-context dogfood of /trekplan --brief … → /trekexecute on the voyage-doctor feature. Brief validated clean; plan written to .claude/plans/trekplan-2026-06-19-voyage-doctor.md (gitignored); execute ran in an isolated git worktree (/private/tmp/claude-voyage-doctor-exec) and was discarded — main HEAD unchanged at aeee4c6, working tree clean.
Pipeline actually exercised: Phase 1 parse → 4b brief-reviewer (PROCEED) → Phase 5 swarm (7 agents: architecture-mapper, dependency-tracer, risk-assessor, task-finder, test-strategist, git-historian, convention-scanner; research-scout skipped — no external tech; ~345k subagent tokens) → Phase 7 synthesis → Phase 8 plan (passes plan-validator --strict) → Phase 9 plan-critic + scope-guardian → revise → execute (TDD) in worktree. Not run (bounded scope, recorded honestly, not silently skipped): the effort: high gemini-bridge 2nd-opinion pass; the full /trekexecute disciplined-executor harness (the plan was executed directly via TDD — the question is plan quality, not the executor's manifest ceremony).
⚠️ Contamination caveat (load-bearing — read first)
The pre-registration (docs/S22-happy-path-dogfood.md) was committed into the repo before exploration, so the Phase-5 swarm read the answer key: architecture-mapper, task-finder, and risk-assessor explicitly cite "the locked Success Criteria and real-risk oracle" and echo R1–R7. The oracle leaked into the swarm. This inflates any "plan handled the R-risks" claim. It does not weaken the Q4 finding — see below, the defects review caught were outside R1–R7. Lesson: when dogfooding a planning tool on its own repo, the scorecard must live outside the explored tree (or be committed after exploration). This is itself a finding about how to run this experiment.
Q1 — does the happy path produce a good plan? Yes, but not a self-sufficient one.
- The plan was executable and correct: all 4 steps implementable; execute produced working code; 15/15 new tests pass; full suite 720 (718 pass / 2 skip / 0 fail) = 705 baseline + 15, zero regression; CLI works (
project-doctor: PASS, exit 0, valid--json). - The plan's code analysis was unusually accurate — plan-critic independently verified all 8 "load-bearing gotchas" against the source and rated every one TRUE (
combine()first-parsed-only,validateResearchDirvalid-on-empty,discoverProjectun-try/catch'd, the trekreview exemption, etc.). The swarm surfaced real, code-grounded defects beyond my oracle (thecombine()clobber, the crash vector, FM_MISSING guard) — a strong positive signal even net of contamination. - But the plan was not executable as first written. plan-critic scored it C (71/100) vs the planner's self-score B+ (88) — a ~17-pt self-inflation. The 3 majors (below) had to be fixed before a clean execute. Verdict: the happy path produces a high-quality draft plan; the adversarial review is load-bearing, not decorative — without it the plan ships a real underspecification and a contradiction.
Q4 — does the adversarial review catch real bugs? Yes — decisively, and beyond the pre-registered target.
This is the audit's "never measured by anyone" question. Measured here: plan-critic caught 3 real majors the planner (Opus 4.8 + a 7-agent swarm) genuinely missed, none planted, none in the R1–R7 oracle:
| Finding | Real? | In R1–R7 oracle? | Consequence if shipped |
|---|---|---|---|
PC-1 discoverProject exposes research as a file-array, not a dir path; Step 1 said "validateResearchDir on the research/ dir" without deriving join(dir,'research') |
✅ verified | ❌ no | Executor blocked / invents the path — confirmed at execute (the path derivation was genuinely absent) |
PC-2 the plan's own #1 Critical risk (PROJECT_DIR_UNREADABLE crash-isolation) was untested — a missing dir returns empty via discoverProject's guard, never hitting the try/catch |
✅ verified | ❌ no (R7 covered "conditional branches", not this meta-gap) | The top risk ships unverified; a meta-catch the planner missed |
PC-3 export-name contradiction: brief SC1 doctorProject vs plan/manifest diagnoseProject |
✅ verified | ❌ no | Executor following SC1 literally fails the manifest |
plan-critic also did not trust the plan — it re-verified the code claims itself. scope-guardian returned ALIGNED: every SC covered, every Non-Goal (NG1–NG4) respected, every cited file:line confirmed exact; one borderline minor (progress/review beyond SC2). The two converged on that single minor.
Recall vs the pre-registered R1–R7 target: the plan handled R1, R2, R4, R5, R6 correctly (composed not reimplemented; correct research truth-table incl. trekreview + topics>0 guard; presence-gated missing-vs-invalid; anchored slug strip; read-only). R3/R7 were partially tripped — the crash-isolation risk was handled in code but under-tested, and plan-critic caught exactly that (PC-2). So review recall on the one tripped pre-registered sub-risk = 1/1. The more important result: the oracle under-predicted where the defects would be. The real majors lived in plan→execute handoff fidelity (an unspecified path, a name contradiction, an untested top-risk), not in the algorithmic risks I anticipated. Adversarial review earned its keep precisely on the class my foresight (and the swarm's) missed.
Execute outcome + an execute-phase finding
Execute succeeded (15/15, suite green, CLI works). One finding only execute could surface: the revised plan's PC-2 fix prescribed t.mock.method to force discoverProject to throw — but it's an ESM named import (read-only namespace binding), so t.mock.method can't redefine it. The executor had to substitute dependency injection (opts.discover). So even the post-review plan carried a residual gap that only contact with the runtime exposed — a reminder that plan review is not a substitute for execution.
Pipeline defects the dogfood surfaced (the bonus the S14 audit could not get — it never ran the pipeline)
/trekplanPhase 9 is broken as documented. It instructs plan-critic + scope-guardian to "Write structured JSON output to/tmp/…out.json", then runsplan-review-dedup.mjson those files. Both agents' frontmatter grants onlyRead/Glob/Grep— noWrite/Bash— so the files are never created and the dedup step cannot run. Both agents fell back to returning JSON inline. Severity: MAJOR (a documented, wired step that cannot execute). Fix options: grant the reviewersWrite, or have the orchestrator persist the returned JSON before calling the dedup helper.- Oracle-into-swarm contamination (see caveat) — a real trap for dogfooding planning tools on their own repo.
plan_versionnot parsed.The plan template emitsRESOLVED (S26).plan_version: 1.7as prose in the "Generated by" line;plan-validatorthen warnsPLAN_NO_VERSION. Minor template/validator mismatch — the validator looks for a frontmatter/parseable field the template doesn't emit.PLAN_VERSION_REGEX(lib/parsers/plan-schema.mjs) was^-anchored and only matched frontmatter; relaxed to/(?:^|)plan_version:.../mso it honors the parser's documented "frontmatter or doc body" contract and parses the backtick-wrapped prose form the template emits. Regression-pinned by runningextractPlanVersionagainst the actualtemplates/plan-template.md`.- Version skew. The installed plugin (skill the operator invokes) is cached at v5.1.1; the repo under development is v5.5.0. The dogfood ran the v5.1.1 command text against v5.5.0
lib/. Harmless here (the phases are stable across the bump) but worth noting: operators dogfooding the installed plugin are not testing the dev tree. ASSESSED (S27) → closed no-op. The skew is still live (~/.claude/plugins/cache/ktg-plugin-marketplace/voyage/5.1.1/vs repo.claude-plugin/plugin.json=5.5.0), but it is an operator/cache state, not a Voyage code defect: the source tree at 5.5.0 is correct, and/trekplanresolving to the cached install rather than the dev tree is expected Claude Code plugin-cache behavior (already noted in STATE's worktree-dogfood gotcha). The only remediation is an environment action (refresh the installed cache via/plugin), out of repo-code scope. A runtime version warning was rejected as scope creep — the installed context cannot know a newer dev tree exists without querying the marketplace.
Honest limitations
- Contamination caps confidence in the "plan handled R1–R7" half of Q1 (the swarm saw R1–R7). The Q4 half is robust because the caught defects were outside the oracle.
- n = 1, one small feature, one domain (an internal validator the planner knew well). Generalization to larger/unfamiliar features is unproven. A feature in an unfamiliar codebase would stress the swarm more and likely lower plan quality.
- No token/$ measurement (audit Blind spot #3 remains open): ~345k Phase-5 + ~110k Phase-9 + ~26k brief-review subagent tokens observed, but not normalized to a per-run cost. Recorded, not analyzed.
- The
voyage-doctorimplementation worked (15 passing tests, full suite green) and is genuinely useful, but was discarded per operator scope (deliverable = measurement). The plan persists locally at.claude/plans/trekplan-2026-06-19-voyage-doctor.md; productizing it is a future-session candidate, not part of S22.
Bottom line
The happy path works and produces high-quality, executable plans — and the adversarial review is load-bearing: it caught 3 real majors the full planning swarm missed, the highest-value being defects in plan→execute handoff fidelity that neither the planner nor the pre-registered oracle anticipated. The flagship "context-engineering via specialized agents + adversarial review" claim is, for this one case, demonstrated rather than asserted — with the honest caveats that it is n=1, the oracle leaked into the swarm, no cost was measured, and the pipeline itself shipped a broken Phase-9 dedup step (defect #1) that this very run exposed.
Verification log (Verifiseringsplikt)
| Claim | How verified |
|---|---|
| Brief validates clean | node lib/validators/brief-validator.mjs --json → {valid:true,errors:[],warnings:[]} |
| Plan passes schema | node lib/validators/plan-validator.mjs --strict --json → valid:true, 4 steps, 0 errors (1 soft PLAN_NO_VERSION warning = defect #3, resolved in S26) |
| 15 new tests pass; suite green | node --test in worktree → new file 15/15; full suite 720 (718/2/0) = 705 baseline +15 |
| CLI works on live dir | node lib/validators/project-doctor.mjs .claude/projects/2026-06-19-voyage-doctor → PASS, exit 0; --json → valid JSON |
| Execute isolated; main untouched | git worktree remove → git status clean, HEAD aeee4c6, lib/validators/project-doctor.mjs absent from main |
| plan-critic's 3 majors are real | Each re-checked against source: PC-1 (project-discovery.mjs research:string[], no dir field), PC-2 (project-discovery.mjs:39 empty-guard precedes any throw), PC-3 (brief SC1 vs plan/manifest names) — all confirmed |
| Reviewers lack Write (defect #1) | Agent registry: voyage:plan-critic + voyage:scope-guardian Tools = Read, Glob, Grep; both reported the write failure at runtime |
| Contamination | architecture-mapper/task-finder/risk-assessor outputs explicitly cite docs/S22-happy-path-dogfood.md and R1–R7 |
| Version skew is cache, not code (defect #4) | cache/.../voyage/5.1.1/.claude-plugin/plugin.json = 5.1.1 vs repo .claude-plugin/plugin.json = 5.5.0; skew confirmed live but is install-cache state, source tree correct. Assessed S27 → no-op (operator/cache, not a code defect) |