voyage/docs/sdlc-playbook-gap.md
Kjell Tore Guttormsen 5f1ae4b4c2
docs(sdlc): measure Voyage against the playbook's 16 practices, from code
An earlier pass scored the same 16 rows from README.md and CLAUDE.md alone and
landed on full 4 / partial 5 / none 7. Re-reading the rows in commands/, agents/,
lib/, hooks/ and tests/ moves three of them and corrects five underlying claims:

- P3 full -> partial. plan-critic is prose ("if blockers are found: revise the
  plan"), not a gate. The plan is never committed by the pipeline, and execution
  starts on the word `execute`. What is hard is the plan_version 1.7 manifest.
- P9 full -> partial. Exit-code-as-truth and the Phase 7.5 audit are real, but
  trekexecute says verbatim "do not block on test-first failures", and nothing
  protects a test file from being rewritten during a fix.
- P10 none -> partial. The docs-only pass missed tests/: the config regression
  suite and a scored gold eval were there all along. The honest remaining gap is
  CI (zero workflow files) and a live-agent eval, not "no evals".

Corrected without moving a verdict: --gates is boolean; /trekreview's input is a
SHA-range diff, not a PR; no hook anywhere returns an "ask" decision; nothing in
the pipeline writes to a memory file.

Result: full 2 of 16, partial 8 of 16, none 6 of 16. Every row carries a file
pointer so a later claim of "we closed that" can be checked against a denominator.

Also records what is worth borrowing from a third-party MIT implementation of the
same playbook, with credit, and what is not: its eval runner and CI example both
invoke Claude Code from code, which this repo prohibits.

No behaviour change. Suite unchanged at 1161 (1159/0/2).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-20 16:20:07 +02:00

9.2 KiB

Voyage measured against the AI-native SDLC playbook

What this is. A self-assessment of Voyage against the 16 named practices in Anthropic's "The AI-Native SDLC playbook" (https://claude.com/blog/the-ai-native-sdlc-playbook, published 2026-08-21). It exists so the gap is a measured list with file pointers rather than an impression, and so a later claim of "we closed that" can be checked against a denominator.

Denominator N = 16. One row per named practice inside the six numbered stages — Plan 1, Design 1, Build 6, Test 2, Deploy 3, Maintain 3. Not counted: the sidebars, the worked example, and the unnumbered sections.

Method, and its limits. Every row was read in Voyage's own files (commands/, agents/, lib/, hooks/, scripts/, tests/) at 3b3f06d. An earlier pass judged the same 16 rows from README.md and CLAUDE.md alone and scored full 4 / partial 5 / none 7; re-measuring against the code moved three rows and corrected five underlying claims without moving their verdict. The verdicts are a judgement call. The file pointers are not.

Result: full 2 of 16 · partial 8 of 16 · none 6 of 16. By stage — Plan 0/1/0 · Design 0/0/1 · Build 2/3/1 · Test 0/2/0 · Deploy 0/2/1 · Maintain 0/0/3.

# Practice (stage) Verdict Pointer Why
P1 Capture intent as a committed artifact (Plan) partial commands/trekbrief.md Phase 2.5 + ## Intent; lib/validators/brief-validator.mjs The brief requires an ## Intent section and a framing: declaration with no skip path, enforced as a BLOCKER — but the artifact lands under .claude/projects/ (gitignored), the pipeline never commits it, and no approval marker exists
P2 Requirements and design as a separate spec (Design) none — No artifact between brief and plan; none of the 7 handovers in docs/HANDOVER-CONTRACTS.md is a spec, and no organisational policy is applied while writing
P3 Plan mode as the default start (Build) partial commands/trekplan.md Phase 9; commands/trekexecute.md Phase 2 (plan-validator.mjs --strict) The plan_version 1.7 manifest gate is real code that stops the run — but plan-critic is prose ("if blockers are found: revise the plan"), the plan is never committed by the pipeline, and execution starts on the word execute with no approved-plan marker
P4 Auto mode and longer autonomous sessions (Build) full lib/util/autonomy-gate.mjs; commands/trekexecute.md (waves, --resume) Auto mode, headless worktree sessions and after-the-fact artifact review all exist. Note: --gates is a boolean, not a three-way mode
P5 CLAUDE.md as institutional memory (Build) partial agents/convention-scanner.md; agents/brief-reviewer.md dim. 6 Conventions and operator memory are read; nothing in the pipeline ever writes to a memory file in any repo, so the "wrong twice → into the file" loop is absent
P6 Skills as institutional knowledge (Build) none — find . -iname 'SKILL.md' and find . -type d -name skills are both empty; the plugin ships 7 commands and 23 agent files and zero skills
P7 Hooks as build-time guardrails (Build) partial hooks/hooks.json (7 wirings/6 events); hooks/scripts/pre-write-executor.mjs, pre-bash-executor.mjs Path protection and credential blocks exist as PreToolUse hooks; there is no PostToolUse Edit/Write hook at all, so "formatter/linter after edit" is missing
P8 Parallel sessions and subagents (Build) full commands/trekexecute.md (git worktree add -b, waves, scope fence); agents/session-decomposer.md Worktree-isolated parallel headless sessions with cleanup, plus 20 spawnable subagents
P9 Give the agent a feedback loop (Test) partial lib/verification/criteria-runner.mjs; commands/trekexecute.md Phase 7 / 7.5 / 7.6 Exit-code-as-truth and an independent fresh-context manifest audit are real code — but "failing test first, protected from rewriting" is explicitly off: Test first: is optional and the file says "do not block on test-first failures"
P10 Continuous evals in CI (Test) partial tests/lib/doc-consistency.test.mjs, tests/lib/agent-frontmatter.test.mjs; tests/lib/gold-eval.test.mjs + lib/review/gold-scorer.mjs A suite does regression-test the agent configuration (agent count, frontmatter, model pin, command table, hook count — derived from the source, so a config change fells it), and a scored gold eval exists. What is missing: any CI at all (.forgejo/workflows and .github both absent), automatic triggering on config change, a merge gate on pass rate, and any live-agent eval
P11 AI in the PR review loop (Deploy) partial commands/trekreview.md; lib/review/rule-catalogue.mjs; lib/review/coordinator-contract.mjs Independent reviewers, a judge, a versioned rule catalogue and a four-level severity taxonomy are strongly implemented — but the input is a SHA-range diff, not a PR: no forge integration, no code-owner approval, and the verdict blocks nothing
P12 Hooks as approval gates (Deploy) partial lib/util/autonomy-gate.mjs; commands/trekexecute.md main-merge gate Human gates exist and the main-merge gate always pauses — but they are AskUserQuestion prose in command files; no hook returns an "ask" decision, and the hooks are binary (exit 0 / exit 2)
P13 CI/CD integration and deployment (Deploy) none — No CI workflows, no deploy/status/rollback tooling, no per-environment autonomy ladder
P14 Close the loop from maintenance (Maintain) none hooks/scripts/otel-export.mjs, lib/exporters/, lib/stats/ Observability exists but measures Voyage's own runs, not a product in production; no control bands, and the one feedback path (review → plan) is started by a human
P15 Recurring codebase scans (Maintain) none hooks/scripts/pre-bash-executor.mjs No scheduler and no periodic entry point; the only cron hits are the rule that blocks crontab persistence
P16 Agent on call in chat (Maintain) none README.md Declared an intentional omission for a solo project

Corrections the code forced

Three verdicts moved when the code was read rather than the docs:

  • P3, full → partial. "plan-critic is a hard gate" does not hold: the phase is prose with no exit code and no re-review loop. What is hard is the plan manifest.
  • P9, full → partial. Two of three sub-claims are strongly implemented; test-first is explicitly non-blocking, and nothing protects a test file from being rewritten during a fix.
  • P10, none → partial. The docs-only pass missed tests/ — both the configuration regression suite and the scored gold eval were already there. The honest remaining gap is CI and a live-agent eval, not "no evals".

Five claims were corrected without moving a verdict: --gates is boolean; /trekreview's input is a diff and not a PR; no "ask" permission decision exists anywhere in the hooks; nothing in the pipeline writes to a memory file; and the intent side of P1 is stronger than assumed while its commit/approval side is weaker.

Third-party implementation: what is worth borrowing

bashebr/ai-native-sdlc (https://github.com/bashebr/ai-native-sdlc, MIT, © 2026 AI-Native SDLC contributors) implements the same playbook as a skill bundle. It is an idea source, not a dependency — one author, 30 commits, one released version. Where an idea below is adopted, credit goes to it under MIT, and to the Anthropic article it is itself derived from.

Candidate Does Voyage already have it Verdict
Hash-chained, tamper-evident approval ledger No. No hash chain exists anywhere; lib/util/research-loop-cap.mjs is append-only but unchained Partly. The idea answers "did we build what we agreed", but a new ledger engine is out of scope; the cheap form is a prev_hash field on the event ledger that already exists
Plan as a machine-readable file list with a lock Yes — plan_version 1.7 manifest, plan-validator.mjs --strict, and the independent Phase 7.5 audit Borrow the boundary only: their named process-file exclusion list, so a state or docs update does not need a manifest entry
Short intent and spec templates Partly — the brief already requires ## Intent, ## Goal and ## Success Criteria No to a separate spec artifact (the operator is both orderer and builder). Yes to their "Gotchas" section: one fixed place for where two constraints collided and who decided
A rule → enforcement table ("advisory or enforced, and where") No such table; the doc-pins are that idea applied narrowly Yes, cheap. It makes visible which of Voyage's own rules are only prose
"A test set without a machine-checkable answer is not a test set" Yes in substance — NOT RUN can never support a passing verdict, and the gate returns "not measured" rather than zero Already held. Borrow the sentence as a written rule, not the mechanism

Not borrowed: their eval runner and CI example (both invoke Claude Code from code, which is prohibited here), the agent "organisation" with roles and queues, everything GitHub-specific, and the managed-settings model, which assumes an IT department that a single-operator machine has no room for.