# S22 — Happy-path dogfood (Blind spot #1 + #4) **Run:** S22, 2026-06-19. **Method:** dogfood Voyage's own pipeline (`/trekplan → /trekexecute`) on a real, small Voyage feature, scored against a **pre-registered** ground-truth scorecard committed *before* the pipeline runs. **Answers the two open questions the S14 audit named but never measured** (`devils-advocate-results.md` §"What this audit might have missed" #1 and #4): - **Q1 — plan quality:** does `/trekplan` produce a *correct, useful* plan on a real feature? - **Q4 — review efficacy:** does the adversarial review (`plan-critic` 10-dim + `scope-guardian`) catch *real* defects? **Operator-chosen shape (S22):** real new Voyage feature (not a planted defect, not a known backlog bug); execute runs in an **isolated git worktree** and the code is **discarded** — the deliverable is the measurement, not the feature. --- ## Methodology — why pre-registration The audit's deepest critique of itself (#4) was that *nobody ever measured whether the review catches real bugs — both attack and defense assumed it.* The failure mode for a dogfood is **post-hoc rationalization**: run the pipeline, then declare whatever it found "the important stuff." To avoid that, the expected plan and the real-risk list below are **written and git-committed before `/trekplan` is launched** (provable by commit order). Q4 is then scored as *recall against a fixed target*, not "did it find something." **The feature** (`voyage-doctor`) was chosen because it has three properties that make Q4 measurable on an honest feature: 1. **Strong reuse anchors** (`discoverProject`, `validateBrief/Research/Plan`) → a naive plan reimplements parsing; a good plan composes. Tests whether `scope-guardian`/`plan-critic` flag reinvention. 2. **A subtle state-conditional correctness rule** (`research_status` ↔ `research/` contents) → easy to get wrong. Tests whether `plan-critic` catches a wrong conditional. 3. **Explicit Non-Goals** (read-only, no auto-fix) with a Context that mildly invites scope-creep → tests whether the review verifies Non-Goal coverage. **Input brief:** `.claude/projects/2026-06-19-voyage-doctor/brief.md` (gitignored per repo convention; Success Criteria + Non-Goals reproduced verbatim below so the control is locked). Brief validates clean: `brief-validator.mjs --json` → `{valid:true, errors:[], warnings:[]}`. ### Brief — Success Criteria (verbatim, locked) - **SC1** — New `lib/validators/project-doctor.mjs` exports a function returning `{valid, errors, warnings}`, building on `discoverProject()` + the per-artifact content validators rather than re-parsing files itself. - **SC2** — Detects ≥3 problem classes: (i) present-but-invalid artifact (surface underlying validator findings, tagged by artifact), (ii) `research_status: complete` with empty `research/`, (iii) project-dir slug ≠ brief `slug` frontmatter. - **SC3** — CLI (`node lib/validators/project-doctor.mjs [--json]`): human report default, `--json`, non-zero exit on invalid — consistent with `brief-validator.mjs`. - **SC4** — `node:test` coverage for the three SC2 coherence branches + a test asserting delegation to existing validators (no re-implementation). ### Brief — Non-Goals (verbatim, locked) - **NG1** — No auto-fix, no mutation; strictly read-only. - **NG2** — No new artifact types, no schema changes to brief/research/plan. - **NG3** — No network / external calls. - **NG4** — Not a replacement for the per-artifact validators or `checkPhaseRequirements()`; composes, does not supersede. --- ## Pre-registered ground truth (LOCKED before `/trekplan`) ### A. Expected plan (Q1 oracle) A competent plan should, at minimum: 1. **Reuse, not reimplement** — call `discoverProject(dir)` for the artifact set, then `validateBrief`/`validateResearch(Dir)`/`validatePlan` on the present artifacts. No new frontmatter/markdown parsing. 2. **Aggregate findings** into one `{valid, errors[], warnings[]}` using the existing `issue()`/`result.mjs` helpers, tagging each finding with the artifact it came from. 3. **Implement the 3 coherence checks** of SC2 with correct conditionals (see real-risks R2 below). 4. **Add a CLI shim** copying `brief-validator.mjs`'s `import.meta.url` pattern (human + `--json` + exit 0/1, usage → exit 2). 5. **TDD** — node:test cases for each coherence branch + a delegation test; fixtures = throwaway project dirs. 6. **Stay read-only** — no writes, honoring NG1. 7. **Decompose into a small number of steps** (module → checks → CLI → tests), each independently testable. A plan scores well on Q1 if it hits 1–6 without inventing scope beyond the brief. ### B. Real risks / defects the review SHOULD catch (Q4 oracle — fixed target) Each risk is something a *flawed* plan could plausibly contain. Scoring records, for each: did the **plan** avoid it, and if the plan tripped it, did **plan-critic or scope-guardian flag it**. | ID | Risk | Which reviewer should catch it if the plan trips it | |----|------|------------------------------------------------------| | **R1** | Plan reimplements brief/research/plan parsing instead of reusing the validators (duplication, drift). | scope-guardian (reuse gap) / plan-critic (maintainability) | | **R2** | The `research_status` ↔ `research/` conditional is wrong: e.g. flags empty `research/` even when `research_topics: 0` / `skipped` (false positive), or misses `complete`+empty (false negative). | plan-critic (correctness / edge cases) | | **R3** | No error isolation: a malformed/unreadable artifact throws and aborts the whole doctor instead of becoming a finding. | plan-critic (error handling / robustness) | | **R4** | Missing-artifact vs present-but-invalid not distinguished (both collapse to one code), so the report is ambiguous. | plan-critic (correctness) | | **R5** | Slug-from-dirname parse is naive: breaks on slugs containing hyphens or dirs lacking the `YYYY-MM-DD-` prefix. | plan-critic (edge cases) | | **R6** | Scope-creep against NG1: plan adds auto-fix / "repair" / writing a report file. | scope-guardian (creep vs Non-Goal) | | **R7** | Tests assert only the happy path; the SC2 conditional branches (esp. R2's false-positive case) are untested. | plan-critic (test coverage) / test-strategist | **Q4 score = (real risks correctly handled by the plan) + (risks the plan tripped that the review flagged) / total applicable.** A risk the plan handles correctly is *not* counted against the review (nothing to catch) but is recorded as "plan avoided it." The review's job is the residual: of the risks the plan got wrong, how many did it surface? ### C. Scoring rubric - **Q1 (plan quality):** PASS / PARTIAL / FAIL against expected-plan items 1–6, plus a one-line qualitative verdict. Independently judged by main context reading the produced `plan.md`. - **Q4 (review efficacy):** for each Ri — `plan: avoided | tripped`, and if tripped `review: caught | missed`. Headline = review recall on tripped risks (caught / tripped). Also note any **real** problems the review raised that are NOT in R1–R7 (true positives outside the pre-registered set → credit) vs. noise/false-positives (debit). - **Execute (worktree):** did `/trekexecute` produce code that (a) matches the plan and (b) passes its own tests + `node --test`? Run in isolated worktree, then discard. Records a yes/partial/no, not a quality grade. --- ## RESULTS **Run:** 2026-06-19, interactive main-context dogfood of `/trekplan --brief … → /trekexecute` on the `voyage-doctor` feature. Brief validated clean; plan written to `.claude/plans/trekplan-2026-06-19-voyage-doctor.md` (gitignored); execute ran in an isolated git worktree (`/private/tmp/claude-voyage-doctor-exec`) and was discarded — `main` HEAD unchanged at `aeee4c6`, working tree clean. **Pipeline actually exercised:** Phase 1 parse → 4b brief-reviewer (PROCEED) → Phase 5 swarm (7 agents: architecture-mapper, dependency-tracer, risk-assessor, task-finder, test-strategist, git-historian, convention-scanner; research-scout skipped — no external tech; ~345k subagent tokens) → Phase 7 synthesis → Phase 8 plan (passes `plan-validator --strict`) → Phase 9 plan-critic + scope-guardian → revise → execute (TDD) in worktree. **Not run** (bounded scope, recorded honestly, not silently skipped): the `effort: high` `gemini-bridge` 2nd-opinion pass; the full `/trekexecute` disciplined-executor harness (the plan was executed directly via TDD — the question is plan quality, not the executor's manifest ceremony). ### ⚠️ Contamination caveat (load-bearing — read first) The pre-registration (`docs/S22-happy-path-dogfood.md`) was **committed into the repo before exploration**, so the Phase-5 swarm **read the answer key**: architecture-mapper, task-finder, and risk-assessor explicitly cite "the locked Success Criteria and real-risk oracle" and echo R1–R7. **The oracle leaked into the swarm.** This inflates any "plan handled the R-risks" claim. It does **not** weaken the Q4 finding — see below, the defects review caught were *outside* R1–R7. **Lesson:** when dogfooding a planning tool on its own repo, the scorecard must live *outside* the explored tree (or be committed after exploration). This is itself a finding about how to run this experiment. ### Q1 — does the happy path produce a good plan? **Yes, but not a self-sufficient one.** - The plan was **executable and correct**: all 4 steps implementable; execute produced working code; **15/15** new tests pass; full suite **720 (718 pass / 2 skip / 0 fail)** = 705 baseline + 15, **zero regression**; CLI works (`project-doctor: PASS`, exit 0, valid `--json`). - The plan's code analysis was **unusually accurate** — plan-critic *independently verified* all 8 "load-bearing gotchas" against the source and rated every one TRUE (`combine()` first-parsed-only, `validateResearchDir` valid-on-empty, `discoverProject` un-try/catch'd, the trekreview exemption, etc.). The swarm surfaced real, code-grounded defects **beyond my oracle** (the `combine()` clobber, the crash vector, FM_MISSING guard) — a strong positive signal even net of contamination. - **But the plan was not executable as first written.** plan-critic scored it **C (71/100)** vs the planner's self-score **B+ (88)** — a ~17-pt self-inflation. The 3 majors (below) had to be fixed before a clean execute. **Verdict: the happy path produces a high-quality draft plan; the adversarial review is load-bearing, not decorative — without it the plan ships a real underspecification and a contradiction.** ### Q4 — does the adversarial review catch real bugs? **Yes — decisively, and beyond the pre-registered target.** This is the audit's "never measured by anyone" question. Measured here: **plan-critic caught 3 real majors the planner (Opus 4.8 + a 7-agent swarm) genuinely missed**, none planted, none in the R1–R7 oracle: | Finding | Real? | In R1–R7 oracle? | Consequence if shipped | |---------|-------|------------------|------------------------| | **PC-1** `discoverProject` exposes `research` as a file-array, not a dir path; Step 1 said "validateResearchDir on the research/ dir" without deriving `join(dir,'research')` | ✅ verified | ❌ no | Executor blocked / invents the path — confirmed at execute (the path derivation was genuinely absent) | | **PC-2** the plan's own #1 *Critical* risk (`PROJECT_DIR_UNREADABLE` crash-isolation) was **untested** — a missing dir returns empty via `discoverProject`'s guard, never hitting the try/catch | ✅ verified | ❌ no (R7 covered "conditional branches", not this meta-gap) | The top risk ships unverified; a meta-catch the planner missed | | **PC-3** export-name contradiction: brief SC1 `doctorProject` vs plan/manifest `diagnoseProject` | ✅ verified | ❌ no | Executor following SC1 literally fails the manifest | plan-critic also did **not** trust the plan — it re-verified the code claims itself. scope-guardian returned **ALIGNED**: every SC covered, every Non-Goal (NG1–NG4) respected, every cited `file:line` confirmed exact; one borderline minor (progress/review beyond SC2). The two converged on that single minor. **Recall vs the pre-registered R1–R7 target:** the plan *handled* R1, R2, R4, R5, R6 correctly (composed not reimplemented; correct research truth-table incl. trekreview + topics>0 guard; presence-gated missing-vs-invalid; anchored slug strip; read-only). R3/R7 were *partially* tripped — the crash-isolation risk was handled **in code** but **under-tested**, and plan-critic **caught exactly that** (PC-2). So review recall on the one tripped pre-registered sub-risk = **1/1**. **The more important result:** the oracle *under-predicted where the defects would be.* The real majors lived in **plan→execute handoff fidelity** (an unspecified path, a name contradiction, an untested top-risk), not in the algorithmic risks I anticipated. Adversarial review earned its keep precisely on the class my foresight (and the swarm's) missed. ### Execute outcome + an execute-phase finding Execute succeeded (15/15, suite green, CLI works). **One finding only execute could surface:** the revised plan's PC-2 fix prescribed `t.mock.method` to force `discoverProject` to throw — but it's an **ESM named import (read-only namespace binding)**, so `t.mock.method` can't redefine it. The executor had to substitute **dependency injection** (`opts.discover`). So even the *post-review* plan carried a residual gap that only contact with the runtime exposed — a reminder that plan review is not a substitute for execution. ### Pipeline defects the dogfood surfaced (the bonus the S14 audit could not get — it never ran the pipeline) 1. **`/trekplan` Phase 9 is broken as documented.** It instructs plan-critic + scope-guardian to "Write structured JSON output to `/tmp/…out.json`", then runs `plan-review-dedup.mjs` on those files. **Both agents' frontmatter grants only `Read/Glob/Grep` — no `Write`/`Bash`** — so the files are never created and the dedup step cannot run. Both agents fell back to returning JSON inline. **Severity: MAJOR** (a documented, wired step that cannot execute). Fix options: grant the reviewers `Write`, or have the orchestrator persist the returned JSON before calling the dedup helper. 2. **Oracle-into-swarm contamination** (see caveat) — a real trap for dogfooding planning tools on their own repo. 3. **`plan_version` not parsed.** The plan template emits `plan_version: 1.7` as prose in the "Generated by" line; `plan-validator` then warns `PLAN_NO_VERSION`. Minor template/validator mismatch — the validator looks for a frontmatter/parseable field the template doesn't emit. 4. **Version skew.** The *installed* plugin (skill the operator invokes) is cached at **v5.1.1**; the repo under development is **v5.5.0**. The dogfood ran the v5.1.1 command text against v5.5.0 `lib/`. Harmless here (the phases are stable across the bump) but worth noting: operators dogfooding the installed plugin are not testing the dev tree. ### Honest limitations - **Contamination** caps confidence in the "plan handled R1–R7" half of Q1 (the swarm saw R1–R7). The Q4 half is robust *because* the caught defects were outside the oracle. - **n = 1, one small feature, one domain** (an internal validator the planner knew well). Generalization to larger/unfamiliar features is unproven. A feature in an unfamiliar codebase would stress the swarm more and likely lower plan quality. - **No token/$ measurement** (audit Blind spot #3 remains open): ~345k Phase-5 + ~110k Phase-9 + ~26k brief-review subagent tokens observed, but not normalized to a per-run cost. Recorded, not analyzed. - **The `voyage-doctor` implementation worked** (15 passing tests, full suite green) and is genuinely useful, but was **discarded per operator scope** (deliverable = measurement). The plan persists locally at `.claude/plans/trekplan-2026-06-19-voyage-doctor.md`; productizing it is a future-session candidate, not part of S22. ### Bottom line The happy path **works** and produces high-quality, executable plans — and the adversarial review is **load-bearing**: it caught 3 real majors the full planning swarm missed, the highest-value being defects in plan→execute handoff fidelity that neither the planner nor the pre-registered oracle anticipated. The flagship "context-engineering via specialized agents + adversarial review" claim is, for this one case, **demonstrated rather than asserted** — with the honest caveats that it is n=1, the oracle leaked into the swarm, no cost was measured, and the pipeline itself shipped a broken Phase-9 dedup step (defect #1) that this very run exposed. ### Verification log (Verifiseringsplikt) | Claim | How verified | |-------|--------------| | Brief validates clean | `node lib/validators/brief-validator.mjs --json` → `{valid:true,errors:[],warnings:[]}` | | Plan passes schema | `node lib/validators/plan-validator.mjs --strict --json` → `valid:true, 4 steps, 0 errors` (1 soft `PLAN_NO_VERSION` warning = defect #3) | | 15 new tests pass; suite green | `node --test` in worktree → new file 15/15; full suite 720 (718/2/0) = 705 baseline +15 | | CLI works on live dir | `node lib/validators/project-doctor.mjs .claude/projects/2026-06-19-voyage-doctor` → `PASS`, exit 0; `--json` → valid JSON | | Execute isolated; main untouched | `git worktree remove` → `git status` clean, HEAD `aeee4c6`, `lib/validators/project-doctor.mjs` absent from main | | plan-critic's 3 majors are real | Each re-checked against source: PC-1 (`project-discovery.mjs` `research:string[]`, no dir field), PC-2 (`project-discovery.mjs:39` empty-guard precedes any throw), PC-3 (brief SC1 vs plan/manifest names) — all confirmed | | Reviewers lack Write (defect #1) | Agent registry: `voyage:plan-critic` + `voyage:scope-guardian` Tools = `Read, Glob, Grep`; both reported the write failure at runtime | | Contamination | architecture-mapper/task-finder/risk-assessor outputs explicitly cite `docs/S22-happy-path-dogfood.md` and R1–R7 |