voyage/docs/cc-upgrade-2.1.181-decision-matrix.md
Kjell Tore Guttormsen c452f75628 docs(voyage): add CC 2.1.130->181 upgrade decision matrix
Two-track research (CC changelog digest x Voyage CC-capability surface) synthesized into a 31-entry adoption catalogue (CC-01..CC-31) across 5 workstreams: W0 correctness, W1 orchestration, W2 model/effort, W3 guardrails, W4 hygiene. Load-bearing changelog claims verified verbatim against the official changelog. Foundation for the continuous-session upgrade plan. Headline: CC 2.1.172 (sub-agents can spawn sub-agents) invalidates Voyage's inline-orchestration premise.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 11:43:16 +02:00

129 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CC-upgrade decision matrix — Claude Code 2.1.130 → 2.1.181
**Scope:** Evaluate every relevant Claude Code change shipped between 2.1.130 and 2.1.181 (latest as of 2026-06-18) against Voyage v5.1.1, and decide adoption per change.
**Precedent:** This continues the F2F14 feature-adoption process referenced in the roadmap ("ny CC-versjon med relevant feature → vurder adopsjon analogt med F2F14-prosessen"). The F-catalogue came from the extracted `ultra-cc-architect` plugin and is no longer bundled; this matrix is a fresh, self-contained catalogue (`CC-NN`) for the 2.1.130→181 window.
**Method:** Two-track research (CC changelog digest + Voyage CC-capability surface inventory), synthesized in a single context. Sources: official changelog (`code.claude.com/docs/en/changelog.md`) + canonical `anthropics/claude-code` CHANGELOG.
## Provenance & verification (verifiseringsplikt)
- **Verified verbatim against the official changelog** (load-bearing claims): `2.1.172` sub-agent nesting; `2.1.154` Opus 4.8 + dynamic workflows; `2.1.166` SendMessage authority; `2.1.178` `Tool(param:value)` syntax; `effort:` frontmatter (`2.1.154`/`2.1.152`).
- **From changelog research, not independently re-verified line-by-line:** all other entries below. Where a decision *hinges* on an unverified claim, it is marked ⚠️ and routed to EVALUATE, never SHIP.
- **Note on 2.1.1302.1.135:** these versions are not in the public changelog (public window begins 2.1.136). No user-facing entries to evaluate.
## Decision legend
| Decision | Meaning |
|----------|---------|
| **SHIP** | Adopt now — clear fit, low risk, no open design question. |
| **EVALUATE** | Needs an empirical test or a design decision before ship/skip. The hard architectural items live here — honest, not deferred-by-another-name. |
| **DEFER** | Adopt later; trigger noted. Not weekend-critical. |
| **SKIP** | Not applicable or rejected; reason noted. |
## Workstream legend
| WS | Theme | Weekend-shippable? |
|----|-------|--------------------|
| **W0** | Correctness/trust — fix now-false documentation | Yes |
| **W1** | Orchestration architecture — delegated spawning / Workflow adoption | No (multi-week, empirical) |
| **W2** | Model & effort alignment (Opus 4.8 + native `effort:`) — **gates v5.4** | Partial |
| **W3** | Guardrails & hooks | Partial |
| **W4** | Free wins & release hygiene | Yes |
---
## W0 — Correctness / trust (the headline)
| ID | Change (version) | Type | Voyage relevance | Decision | Rationale |
|----|------------------|------|------------------|----------|-----------|
| **CC-01** | Sub-agents can spawn their own sub-agents, up to 5 levels deep (**2.1.172**, verified); foreground subagents respect the same depth cap (2.1.181) | BREAKING-premise | Voyage's v2.4.0 inline-orchestration migration is justified in 4 places by the claim "the harness does not expose the Agent tool to sub-agents" → `agents/planning-orchestrator.md:511`, `agents/research-orchestrator.md:510`, `agents/review-orchestrator.md:512,:220`, `commands/trekplan.md:399406`. **That claim is now factually false.** | **SHIP** (doc correction) | Independent of the architecture decision (W1), the docs assert an impossibility that is now possible. Correct the four sites to state the factual position: as of CC 2.1.172 sub-agents *can* spawn sub-agents (≤5 deep); Voyage currently still orchestrates inline; re-architecture is under evaluation (W1/CC-26). Do **not** silently delete — replace with truth + forward pointer. The *design* response is CC-26 (EVALUATE). |
---
## W4 — Free wins & release hygiene (weekend-shippable)
| ID | Change (version) | Type | Voyage relevance | Decision | Rationale |
|----|------------------|------|------------------|----------|-----------|
| **CC-02** | MCP `tools/list` pagination fixes (2.1.144/147); `MCP_TOOL_TIMEOUT` raises remote fetch timeout (2.1.142); sub-1000ms `timeout` now ignored→default (2.1.162) | FIX | research agents (`docs-researcher`, `*-researcher`) use tavily/ms-learn/gemini MCP | **SHIP** (verify) | Automatic benefit — no code change. Verify research agents still enumerate tools correctly under paginated servers. |
| **CC-03** | `SendMessage` relayed messages lose user authority (**2.1.166**, verified) | BREAKING | Voyage uses `SendMessage` **nowhere** (confirmed by inventory) | **SKIP** | No impact. Recorded so a future `SendMessage`-based design knows the constraint up front. |
| **CC-04** | Subagent frontmatter `mcpServers` now enforce `--strict-mcp-config` + enterprise allow/deny (2.1.153) | CHANGE | Voyage ships **no `.mcp.json`** and declares MCP *tool grants* (`mcp__tavily__…`) not server *configs*; agents degrade gracefully when servers absent | **EVALUATE** (low) | Likely no impact, but confirm research agents degrade cleanly under `--strict-mcp-config`. One quick test. |
| **CC-05** | Fable 5 `[1m]` model-name suffix normalized automatically (2.1.173) | FIX | Voyage pins the `opus` alias, not raw ids | **SKIP** | No action; noted for completeness. |
| **CC-06** | `claude plugin validate` flags `skills:` pointing at a file vs dir (2.1.145); richer `claude plugin details` (2.1.143/145) | NEW | release process | **SHIP** (process) | Add `claude plugin validate` to the pre-release checklist. Voyage ships no skills, but validate also checks manifest/components. |
| **CC-07** | `fallbackModel` setting — up to 3 ordered fallbacks; `--fallback-model` now applies to interactive too (2.1.166) | NEW | `trekexecute` headless children (`claude -p`) could gain resilience to model-overload | **DEFER** | Real value for long headless waves. Trigger: when we touch `trekexecute` Phase 2.6 launch flags (or in W2 model work). Not weekend-critical. |
| **CC-08** | Hooks may not fire reliably in headless child sessions — Voyage documents GH #36071 (`trekexecute.md:341`, `templates/headless-launch-template.md:48`) | (status unknown) | safety-preamble is the headless defense if hooks don't run | **EVALUATE** ⚠️ | Changelog does **not** confirm #36071 is fixed. Verify current status before relaxing the in-prompt safety preamble. Until verified: keep the preamble. |
| **CC-09** | `--safe-mode` / `CLAUDE_CODE_SAFE_MODE` — start with all customizations disabled (2.1.169) | NEW | dev/test aid — reproduce "bare harness" behavior | **DEFER** | Useful for regression-testing Voyage against an unconfigured harness. Adopt into test tooling when convenient. |
---
## W3 — Guardrails & hooks
| ID | Change (version) | Type | Voyage relevance | Decision | Rationale |
|----|------------------|------|------------------|----------|-----------|
| **CC-10** | `Tool(param:value)` permission syntax, e.g. `Agent(model:opus)` (**2.1.178**, verified) | NEW | Voyage pins `opus` in 23 agent frontmatters *by convention*; this could enforce it as a hard permission rule | **EVALUATE** ⚠️ | **Trap:** a blanket `Agent(model:opus)`-style rule conflicts with the `balanced`/`economy` profiles, which deliberately spawn `sonnet`. Any enforcement must be **profile-aware** → tie to W2. Do NOT ship a blanket opus-lock. |
| **CC-11** | `disallowed-tools` in command/skill frontmatter (2.1.152, verified) | NEW | tighten per-command tool surfaces (e.g. ensure `trekexecute` can't spawn Agents — it already documents "no Agent tool") | **EVALUATE** (low) | Promote existing *documented* tool exclusions into *enforced* `disallowed-tools`. Small, defense-in-depth. Check interaction with `allowed-tools`. |
| **CC-12** | Stop/SubagentStop hooks can return `hookSpecificOutput.additionalContext` without being a hook error (2.1.163) | NEW | `post-compact-flush.mjs` already emits `additionalContext`; SubagentStop is a new surface | **DEFER** | Possible richer continuity injection. No current gap forces it. |
| **CC-13** | Stop/SubagentStop hook input now includes `background_tasks` + `session_crons` (2.1.145) | NEW | `otel-export.mjs` (Stop hook) observability | **DEFER** | Could enrich exported telemetry. Adopt when next touching the exporter. |
| **CC-14** | Hook `args: string[]` exec-form — no shell, no quoting (2.1.139) | NEW | Voyage's 7 hooks invoke node scripts via shell form with `${CLAUDE_PLUGIN_ROOT}` paths | **EVALUATE** (low) | Exec-form removes a class of path-quoting bugs. Low-risk robustness upgrade across `hooks/hooks.json`; verify each hook's invocation. |
| **CC-15** | Hook `if:` conditions for Read/Edit/Write paths now match reliably (2.1.139/176); reopens deferred **F2** (scope pre-bash/pre-write executors to execute sessions) | FIX | `pre-bash-executor.mjs`, `pre-write-executor.mjs` are currently universal | **EVALUATE** | F2 was deferred with "universal protection wins." The `if:` mechanism now works, so the *option* is real again — but the original rationale (brief/plan sessions also benefit from the denylist) may still hold. Re-decide explicitly, don't auto-adopt. |
| **CC-16** | `SessionStart` `reloadSkills` + `sessionTitle` (2.1.152) | NEW | Voyage sets session title via `UserPromptSubmit` (`session-title.mjs`) | **SKIP** | Current mechanism works and is command-scoped (title reflects the invoked `/trek*` command). SessionStart-title would fire before the command is known. No gain. |
| **CC-17** | Hook `terminalSequence` output — notifications/bells without a TTY (2.1.141); `continueOnBlock` for PostToolUse (2.1.139); `MessageDisplay` event (2.1.152); Stop block-cap 8 (2.1.143) | NEW | minor ergonomics; Voyage hooks are fail-open and non-interactive | **SKIP/DEFER** | No current need. `MessageDisplay`/`continueOnBlock` SKIP (no use case); `terminalSequence` DEFER (could notify on long headless waves). Block-cap is informational. |
---
## W2 — Model & effort alignment (gates v5.4)
| ID | Change (version) | Type | Voyage relevance | Decision | Rationale |
|----|------------------|------|------------------|----------|-----------|
| **CC-21** | **Opus 4.8** available; `opus` defaults to high/xhigh effort (**2.1.154**, verified) | NEW | All 23 agents + 7 commands pin `model: opus` → now resolve to 4.8 at a higher default effort | **EVALUATE** | Voyage's reasoning changed under it. Confirm behavior, update model references in docs (`CLAUDE.md`, `docs/profiles.md`, README still say things tied to older model assumptions), and decide whether the default-high-effort interacts with the `phase_signals` effort tiers. |
| **CC-22** | Native `effort:` frontmatter on agents/skills/commands + 5 levels low/medium/high/xhigh/max (**2.1.149/154**, verified) | NEW | Voyage **reinvented** effort as `phase_signals` (low/standard/high) + the `phase-signal-resolver.mjs`/`resolver.mjs` system (which carries the MAJOR doc/code-inconsistency finding) | **EVALUATE** (core W2) | **The pivotal decision.** Options: (a) map `phase_signals` effort → native `effort:` and let the harness apply it (simplifies/possibly removes resolver code); (b) keep the bespoke system; (c) hybrid. This **must resolve before v5.4 freezes the brief schema as a public contract** — freezing `phase_signals` shape now risks freezing something about to change. Absorbs/retires the resolver MAJOR finding. |
| **CC-23** | Claude Fable 5 (Mythos-class) GA (2.1.170) | NEW | Voyage is a planning/reasoning pipeline; Fable is a different model class | **SKIP** (revisit) | Not an obvious fit for deep-planning agents. Note as an option if a future profile wants a distinct model class for a specific phase. |
| **CC-24** | `enforceAvailableModels` managed setting (2.1.175); `availableModels` now constrains Default + subagent overrides (2.1.172/176) | CHANGE | enterprise deployments could constrain Voyage's `opus`/`sonnet` picks | **DEFER** | Document that profile model picks must live within any managed `availableModels` allowlist. Relevant for enterprise consumers post-open-release. |
| **CC-25** | `MAX_THINKING_TOKENS=0` / `--thinking disabled` disables thinking on think-by-default models (2.1.166) | NEW | could gate thinking in cheap `economy`-profile phases | **DEFER** | Minor cost lever; fold into W2 profile design if useful. |
---
## W1 — Orchestration architecture (multi-week, empirical)
| ID | Change (version) | Type | Voyage relevance | Decision | Rationale |
|----|------------------|------|------------------|----------|-----------|
| **CC-26** | Sub-agents spawn sub-agents ≤5 deep (**2.1.172**, verified) — the design response to CC-01 | NEW | Could restore *delegated* orchestration: an orchestrator sub-agent spawns the swarm; synthesis/writing delegated (the "missing summarizer link" in `docs/subagent-delegation-audit.md`). Frees main-context tokens | **EVALUATE** (empirical) | "Can spawn 5 deep" ≠ "Voyage's orchestrator→6-agent-swarm pattern performs well." Design + run a Q3-style measurement (reuse `scripts/q3-cache-prefix-experiment.mjs` harness pattern) before re-architecting. Depth cap (5) may bound nested pipelines. **Highest-value, highest-risk item.** |
| **CC-27** | Dynamic Workflows / Workflow tool — orchestrates tenshundreds of agents (**2.1.154**, verified); keyword `workflow``ultracode` (2.1.160); `agent()` attribution headers (2.1.174) | NEW | Voyage **hand-rolls** swarm/wave/pipeline orchestration in command prose — the Workflow tool is a native primitive for exactly this | **EVALUATE** (strategic) | The biggest identity decision: adopt Workflow as Voyage's execution substrate, or stay prose-orchestrated for portability/control? Tradeoffs: native concurrency + pipelining + budget control vs. dependency on a newer primitive + loss of fine-grained prose control + opt-in/billing semantics. Prototype one pipeline (e.g. `/trekreview`'s reviewer swarm) as a Workflow and compare. |
| **CC-28** | `TaskCreate` reliability — auto-repairs malformed input, schema in errors (2.1.163/169) | FIX | `TaskCreate`/`TaskUpdate` are in `trekplan`/orchestrator frontmatter but not actively used in command logic | **DEFER** | Becomes relevant only if W1 adopts task-graph orchestration. Tie to CC-26/27 outcome. |
| **CC-29** | `subagent_type` matching now case/separator-insensitive (2.1.140); multiple `Agent(...)` types in `tools:` no longer dropped (2.1.147); subagent transcript/backgrounding fixes (2.1.178) | FIX | improves DX of any delegated-orchestration design | **SHIP** (verify) | Free robustness. Confirm Voyage's agent `tools:` grants (none currently declare multiple `Agent(...)` types) and `subagent_type` references are unaffected. |
| **CC-31** | Worktree-isolation guard now applies in background sessions (2.1.154); `worktree.bgIsolation:"none"` (2.1.143); `EnterWorktree` switching mid-session (2.1.157) | CHANGE | `trekexecute` Phase 2.6 parallel waves + `trekplan` "execute with team" use git worktrees / `TeamCreate isolation:"worktree"` | **EVALUATE** | Verify Voyage's worktree-based parallel execution still behaves under the tightened bg-isolation guard. Affects the multi-session headless path. |
---
## Sequencing
```
Weekend (W0 + W4): CC-01 doc-truth fix · CC-06 plugin validate · CC-02/CC-29 verify-no-regression
→ ship something important AND correct.
Near-term (W2): CC-21 + CC-22 effort/model alignment decision → GATES v5.4 (per operator decision).
Resolves the resolver MAJOR finding in the same pass.
Then: v5.4 brief-schema public contract → v5.5 brief framing enforcement.
Multi-week (W1): CC-26 empirical sub-agent-nesting test + CC-27 Workflow-adoption prototype.
Outcome rewrites CC-01's "current design" statement and feeds the
docs/subagent-delegation-audit.md open problem.
Incremental (W3): CC-11/CC-14/CC-15 as small, independently-shippable hardening PRs.
```
## Open questions (need operator or empirical answer)
1. **W1 identity:** does Voyage adopt the Workflow tool as substrate, or stay prose-orchestrated? (CC-27)
2. **W1 perf:** does delegated orchestration (orchestrator sub-agent → swarm) beat inline at Voyage's scale? (CC-26 — empirical)
3. **W2 effort model:** map `phase_signals` onto native `effort:`, or keep bespoke? (CC-22 — gates v5.4)
4. **CC-08:** is GH #36071 (hooks in headless) fixed? Determines whether the safety-preamble can relax.
## Empirical tests required
- **T1 (CC-26):** orchestrator-sub-agent spawns the planning swarm vs. inline baseline — wall-time, quality, token cost, depth-cap behavior. Harness: extend `scripts/q3-cache-prefix-experiment.mjs` pattern.
- **T2 (CC-27):** reimplement `/trekreview`'s reviewer swarm as a Workflow; compare control, cost, and output fidelity vs. prose orchestration.
- **T3 (CC-04/CC-31):** research-agent MCP degradation under `--strict-mcp-config`; worktree parallel-wave behavior under tightened bg-isolation.
---
_Catalogue covers 31 evaluated changes (CC-01…CC-31; CC-18/19/20/30 folded into CC-17/CC-29 rows). Generated 2026-06-18 against Voyage v5.1.1 / CC 2.1.181._