voyage/docs
Kjell Tore Guttormsen 5f1ae4b4c2
docs(sdlc): measure Voyage against the playbook's 16 practices, from code
An earlier pass scored the same 16 rows from README.md and CLAUDE.md alone and
landed on full 4 / partial 5 / none 7. Re-reading the rows in commands/, agents/,
lib/, hooks/ and tests/ moves three of them and corrects five underlying claims:

- P3 full -> partial. plan-critic is prose ("if blockers are found: revise the
  plan"), not a gate. The plan is never committed by the pipeline, and execution
  starts on the word `execute`. What is hard is the plan_version 1.7 manifest.
- P9 full -> partial. Exit-code-as-truth and the Phase 7.5 audit are real, but
  trekexecute says verbatim "do not block on test-first failures", and nothing
  protects a test file from being rewritten during a fix.
- P10 none -> partial. The docs-only pass missed tests/: the config regression
  suite and a scored gold eval were there all along. The honest remaining gap is
  CI (zero workflow files) and a live-agent eval, not "no evals".

Corrected without moving a verdict: --gates is boolean; /trekreview's input is a
SHA-range diff, not a PR; no hook anywhere returns an "ask" decision; nothing in
the pipeline writes to a memory file.

Result: full 2 of 16, partial 8 of 16, none 6 of 16. Every row carries a file
pointer so a later claim of "we closed that" can be checked against a denominator.

Also records what is worth borrowing from a third-party MIT implementation of the
same playbook, with credit, and what is not: its eval runner and CI example both
invoke Claude Code from code, which this repo prohibits.

No behaviour change. Suite unchanged at 1161 (1159/0/2).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-20 16:20:07 +02:00
..
eval-corpus feat(eval): SKAL-1·4b offline gold-scored output eval 2026-06-30 09:00:33 +02:00
agent-description-token-trim-brief.md docs(voyage): track agent-description token-trim brief (M4 input) 2026-06-26 20:15:34 +02:00
agent-return-channel-defect.md fix(review): fail-closed verdicts - an unsubstantiated finding can no longer yield ALLOW 2026-09-01 22:46:39 +02:00
architecture.md release(v5.10.1): drop gemini-bridge from the pipeline; correct the T1 §6 PoC status 2026-09-03 20:29:39 +02:00
balance-backlog-plan.md chore(voyage): S34 — V30 economy-profile self-declares experimental (uncalibrated Jaccard floor) 2026-06-20 10:18:38 +02:00
BRIEF-vurdering-v2.md docs(brief): sharpen the tiltak-3 citation 2026-08-20 23:01:45 +02:00
cc-upgrade-2.1.181-decision-matrix.md docs(voyage): S20 — CC-04/T3 verified clean (research-agent MCP degradation under --strict-mcp-config) 2026-06-19 14:08:51 +02:00
claudemd-token-trim-brief.md docs(claude-md): trim CLAUDE.md to invariants (always-loaded token trim, S53) 2026-06-29 14:49:32 +02:00
command-modes.md release(v5.10.1): drop gemini-bridge from the pipeline; correct the T1 §6 PoC status 2026-09-03 20:29:39 +02:00
deep-research-engine-brief.md docs(research): resolve deep-research-engine topic-1 (/deep-research trigging) 2026-06-30 10:39:03 +02:00
deep-research-engine-research.md docs(research): remove literal keyword tripping verify SC1 2026-08-12 20:29:35 +02:00
devils-advocate-plan.md docs(voyage): plan S14 devil's-advocate audit via Dynamic Workflow 2026-06-18 18:30:05 +02:00
devils-advocate-results.md docs(voyage): S22 — happy-path dogfood results (blind spot #1/#4 measured) 2026-06-19 20:53:21 +02:00
HANDOVER-CONTRACTS.md docs(voyage): fable-aware allowlist prose in contracts, architecture, templates, CLAUDE.md 2026-07-02 17:14:35 +02:00
observability.md docs(observability): document token-usage schema + main-context v1 scope 2026-06-26 14:47:24 +02:00
operations.md docs(voyage): add fable profile row and correct model-allowlist prose 2026-07-02 17:13:15 +02:00
profiles.md docs(voyage): add fable profile row and correct model-allowlist prose 2026-07-02 17:13:15 +02:00
S22-happy-path-dogfood.md docs(voyage): S27 — close version-skew (S22 defect #4) as no-op 2026-06-19 22:11:25 +02:00
sdlc-playbook-gap.md docs(sdlc): measure Voyage against the playbook's 16 practices, from code 2026-09-20 16:20:07 +02:00
spike-pretooluse-subagent-reach.md fix(cap-hook): shrink the inherited deny window and print the way out of it 2026-08-12 23:06:47 +02:00
storm-measurement.md fix(storm-measure): check BOTH halves of the activation SC, not just the count delta 2026-08-12 23:01:59 +02:00
subagent-delegation-audit.md feat(voyage)!: bulk content rewrite ultra -> voyage/trek prose [skip-docs] 2026-05-05 15:08:20 +02:00
T1-cc26-delegated-orchestration.md docs(t1): record the operator's decline of the section 5 head-to-head 2026-09-18 00:07:42 +02:00
T1-synthesis-poc-results.md feat(voyage): S12 — NW3 synthesis-agent built + measured → declined per measurement [skip-docs] 2026-06-18 17:58:39 +02:00
T2-bakeoff-results.md fix(voyage): S21 — close IPv4-mapped IPv6 SSRF bypass + security/safety audit (blind spot #2) 2026-06-19 20:02:56 +02:00
T2-cc27-workflow-substrate.md docs(voyage): S8 (W1/CC-27 gate) — T2 Workflow-substrate probe + measurement design 2026-06-18 13:34:05 +02:00
voyage-vs-cc-balance-analysis.md chore(voyage): S34 — V30 economy-profile self-declares experimental (uncalibrated Jaccard floor) 2026-06-20 10:18:38 +02:00
voyage-vs-cc-balance-charter.md docs(voyage): add Voyage-vs-CC balance-analysis charter (next-session launch spec) 2026-06-20 06:34:40 +02:00
W1-narrow-wins-plan.md docs(voyage): plan W1 narrow-wins implementation (NW1/NW2/NW3, S9->) 2026-06-18 13:41:20 +02:00