An earlier pass scored the same 16 rows from README.md and CLAUDE.md alone and
landed on full 4 / partial 5 / none 7. Re-reading the rows in commands/, agents/,
lib/, hooks/ and tests/ moves three of them and corrects five underlying claims:
- P3 full -> partial. plan-critic is prose ("if blockers are found: revise the
plan"), not a gate. The plan is never committed by the pipeline, and execution
starts on the word `execute`. What is hard is the plan_version 1.7 manifest.
- P9 full -> partial. Exit-code-as-truth and the Phase 7.5 audit are real, but
trekexecute says verbatim "do not block on test-first failures", and nothing
protects a test file from being rewritten during a fix.
- P10 none -> partial. The docs-only pass missed tests/: the config regression
suite and a scored gold eval were there all along. The honest remaining gap is
CI (zero workflow files) and a live-agent eval, not "no evals".
Corrected without moving a verdict: --gates is boolean; /trekreview's input is a
SHA-range diff, not a PR; no hook anywhere returns an "ask" decision; nothing in
the pipeline writes to a memory file.
Result: full 2 of 16, partial 8 of 16, none 6 of 16. Every row carries a file
pointer so a later claim of "we closed that" can be checked against a denominator.
Also records what is worth borrowing from a third-party MIT implementation of the
same playbook, with credit, and what is not: its eval runner and CI example both
invoke Claude Code from code, which this repo prohibits.
No behaviour change. Suite unchanged at 1161 (1159/0/2).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
75 lines
9.2 KiB
Markdown
75 lines
9.2 KiB
Markdown
# Voyage measured against the AI-native SDLC playbook
|
|
|
|
**What this is.** A self-assessment of Voyage against the 16 named practices in Anthropic's
|
|
"The AI-Native SDLC playbook" (<https://claude.com/blog/the-ai-native-sdlc-playbook>,
|
|
published 2026-08-21). It exists so the gap is a measured list with file pointers rather than
|
|
an impression, and so a later claim of "we closed that" can be checked against a denominator.
|
|
|
|
**Denominator N = 16.** One row per named practice inside the six numbered stages — Plan 1,
|
|
Design 1, Build 6, Test 2, Deploy 3, Maintain 3. Not counted: the sidebars, the worked example,
|
|
and the unnumbered sections.
|
|
|
|
**Method, and its limits.** Every row was read in Voyage's own files (`commands/`, `agents/`,
|
|
`lib/`, `hooks/`, `scripts/`, `tests/`) at `3b3f06d`. An earlier pass judged the same 16 rows
|
|
from `README.md` and `CLAUDE.md` alone and scored full 4 / partial 5 / none 7; re-measuring
|
|
against the code moved three rows and corrected five underlying claims without moving their
|
|
verdict. The verdicts are a judgement call. The file pointers are not.
|
|
|
|
**Result: full 2 of 16 · partial 8 of 16 · none 6 of 16.**
|
|
By stage — Plan 0/1/0 · Design 0/0/1 · Build 2/3/1 · Test 0/2/0 · Deploy 0/2/1 · Maintain 0/0/3.
|
|
|
|
| # | Practice (stage) | Verdict | Pointer | Why |
|
|
|---|---|---|---|---|
|
|
| P1 | Capture intent as a committed artifact (Plan) | partial | `commands/trekbrief.md` Phase 2.5 + `## Intent`; `lib/validators/brief-validator.mjs` | The brief requires an `## Intent` section and a `framing:` declaration with no skip path, enforced as a BLOCKER — but the artifact lands under `.claude/projects/` (gitignored), the pipeline never commits it, and no approval marker exists |
|
|
| P2 | Requirements and design as a separate spec (Design) | none | — | No artifact between brief and plan; none of the 7 handovers in `docs/HANDOVER-CONTRACTS.md` is a spec, and no organisational policy is applied while writing |
|
|
| P3 | Plan mode as the default start (Build) | partial | `commands/trekplan.md` Phase 9; `commands/trekexecute.md` Phase 2 (`plan-validator.mjs --strict`) | The `plan_version 1.7` manifest gate is real code that stops the run — but plan-critic is prose ("if blockers are found: revise the plan"), the plan is never committed by the pipeline, and execution starts on the word `execute` with no approved-plan marker |
|
|
| P4 | Auto mode and longer autonomous sessions (Build) | full | `lib/util/autonomy-gate.mjs`; `commands/trekexecute.md` (waves, `--resume`) | Auto mode, headless worktree sessions and after-the-fact artifact review all exist. Note: `--gates` is a boolean, not a three-way mode |
|
|
| P5 | `CLAUDE.md` as institutional memory (Build) | partial | `agents/convention-scanner.md`; `agents/brief-reviewer.md` dim. 6 | Conventions and operator memory are read; nothing in the pipeline ever writes to a memory file in any repo, so the "wrong twice → into the file" loop is absent |
|
|
| P6 | Skills as institutional knowledge (Build) | none | — | `find . -iname 'SKILL.md'` and `find . -type d -name skills` are both empty; the plugin ships 7 commands and 23 agent files and zero skills |
|
|
| P7 | Hooks as build-time guardrails (Build) | partial | `hooks/hooks.json` (7 wirings/6 events); `hooks/scripts/pre-write-executor.mjs`, `pre-bash-executor.mjs` | Path protection and credential blocks exist as PreToolUse hooks; there is no PostToolUse Edit/Write hook at all, so "formatter/linter after edit" is missing |
|
|
| P8 | Parallel sessions and subagents (Build) | full | `commands/trekexecute.md` (`git worktree add -b`, waves, scope fence); `agents/session-decomposer.md` | Worktree-isolated parallel headless sessions with cleanup, plus 20 spawnable subagents |
|
|
| P9 | Give the agent a feedback loop (Test) | partial | `lib/verification/criteria-runner.mjs`; `commands/trekexecute.md` Phase 7 / 7.5 / 7.6 | Exit-code-as-truth and an independent fresh-context manifest audit are real code — but "failing test first, protected from rewriting" is explicitly off: `Test first:` is optional and the file says "do not block on test-first failures" |
|
|
| P10 | Continuous evals in CI (Test) | partial | `tests/lib/doc-consistency.test.mjs`, `tests/lib/agent-frontmatter.test.mjs`; `tests/lib/gold-eval.test.mjs` + `lib/review/gold-scorer.mjs` | A suite does regression-test the agent configuration (agent count, frontmatter, model pin, command table, hook count — derived from the source, so a config change fells it), and a scored gold eval exists. What is missing: any CI at all (`.forgejo/workflows` and `.github` both absent), automatic triggering on config change, a merge gate on pass rate, and any live-agent eval |
|
|
| P11 | AI in the PR review loop (Deploy) | partial | `commands/trekreview.md`; `lib/review/rule-catalogue.mjs`; `lib/review/coordinator-contract.mjs` | Independent reviewers, a judge, a versioned rule catalogue and a four-level severity taxonomy are strongly implemented — but the input is a SHA-range diff, not a PR: no forge integration, no code-owner approval, and the verdict blocks nothing |
|
|
| P12 | Hooks as approval gates (Deploy) | partial | `lib/util/autonomy-gate.mjs`; `commands/trekexecute.md` main-merge gate | Human gates exist and the main-merge gate always pauses — but they are `AskUserQuestion` prose in command files; no hook returns an "ask" decision, and the hooks are binary (exit 0 / exit 2) |
|
|
| P13 | CI/CD integration and deployment (Deploy) | none | — | No CI workflows, no deploy/status/rollback tooling, no per-environment autonomy ladder |
|
|
| P14 | Close the loop from maintenance (Maintain) | none | `hooks/scripts/otel-export.mjs`, `lib/exporters/`, `lib/stats/` | Observability exists but measures Voyage's own runs, not a product in production; no control bands, and the one feedback path (review → plan) is started by a human |
|
|
| P15 | Recurring codebase scans (Maintain) | none | `hooks/scripts/pre-bash-executor.mjs` | No scheduler and no periodic entry point; the only `cron` hits are the rule that blocks crontab persistence |
|
|
| P16 | Agent on call in chat (Maintain) | none | `README.md` | Declared an intentional omission for a solo project |
|
|
|
|
## Corrections the code forced
|
|
|
|
Three verdicts moved when the code was read rather than the docs:
|
|
|
|
- **P3, full → partial.** "plan-critic is a hard gate" does not hold: the phase is prose with no
|
|
exit code and no re-review loop. What *is* hard is the plan manifest.
|
|
- **P9, full → partial.** Two of three sub-claims are strongly implemented; test-first is
|
|
explicitly non-blocking, and nothing protects a test file from being rewritten during a fix.
|
|
- **P10, none → partial.** The docs-only pass missed `tests/` — both the configuration
|
|
regression suite and the scored gold eval were already there. The honest remaining gap is CI
|
|
and a live-agent eval, not "no evals".
|
|
|
|
Five claims were corrected without moving a verdict: `--gates` is boolean; `/trekreview`'s input
|
|
is a diff and not a PR; no "ask" permission decision exists anywhere in the hooks; nothing in the
|
|
pipeline writes to a memory file; and the intent side of P1 is stronger than assumed while its
|
|
commit/approval side is weaker.
|
|
|
|
## Third-party implementation: what is worth borrowing
|
|
|
|
`bashebr/ai-native-sdlc` (<https://github.com/bashebr/ai-native-sdlc>, MIT, © 2026 AI-Native SDLC
|
|
contributors) implements the same playbook as a skill bundle. It is an idea source, not a
|
|
dependency — one author, 30 commits, one released version. Where an idea below is adopted, credit
|
|
goes to it under MIT, and to the Anthropic article it is itself derived from.
|
|
|
|
| Candidate | Does Voyage already have it | Verdict |
|
|
|---|---|---|
|
|
| Hash-chained, tamper-evident approval ledger | No. No hash chain exists anywhere; `lib/util/research-loop-cap.mjs` is append-only but unchained | **Partly.** The idea answers "did we build what we agreed", but a new ledger engine is out of scope; the cheap form is a `prev_hash` field on the event ledger that already exists |
|
|
| Plan as a machine-readable file list with a lock | Yes — `plan_version 1.7` manifest, `plan-validator.mjs --strict`, and the independent Phase 7.5 audit | **Borrow the boundary only:** their named process-file exclusion list, so a state or docs update does not need a manifest entry |
|
|
| Short intent and spec templates | Partly — the brief already requires `## Intent`, `## Goal` and `## Success Criteria` | **No** to a separate spec artifact (the operator is both orderer and builder). **Yes** to their "Gotchas" section: one fixed place for where two constraints collided and who decided |
|
|
| A rule → enforcement table ("advisory or enforced, and where") | No such table; the doc-pins are that idea applied narrowly | **Yes, cheap.** It makes visible which of Voyage's own rules are only prose |
|
|
| "A test set without a machine-checkable answer is not a test set" | Yes in substance — `NOT RUN` can never support a passing verdict, and the gate returns "not measured" rather than zero | **Already held.** Borrow the sentence as a written rule, not the mechanism |
|
|
|
|
**Not borrowed:** their eval runner and CI example (both invoke Claude Code from code, which is
|
|
prohibited here), the agent "organisation" with roles and queues, everything GitHub-specific, and
|
|
the managed-settings model, which assumes an IT department that a single-operator machine has no
|
|
room for.
|