voyage/docs/cc-upgrade-2.1.181-decision-matrix.md
Kjell Tore Guttormsen b5f3d4a932 docs(voyage): S8 (W1/CC-27 gate) — T2 Workflow-substrate probe + measurement design
Second W1 gate, same staged execution as S7 (operator-chosen): cheap live
feasibility probe + design doc; the prose-vs-Workflow bake-off specified but
NOT run.

Probe (CC 2.1.181 interactive): a minimal trekreview-shaped Workflow —
parallel([reviewerA, reviewerB]) with a findings schema -> agent(coordinator)
with a verdict schema — ran end-to-end. F1 core ports natively; F2 structured
schemas retire the JSON-parse fragility at trekreview.md:202-204; F3 result
returns to main; F4 a small purposeful fan-out did NOT trip the S7 proliferation
classifier. 3 agents / 85461 tokens / 13.8s.

Reframe: 'substrate swap' is a false binary — a /trek* command is ~80%
non-orchestration glue, so Workflow can only replace the fan-out->synthesize
core (hybrid).

CC-27 recommendation (operator gates verdict): selective hybrid, NOT wholesale
swap. Tier 1 ship a prose schema contract (the F2 win, no Workflow dep); tier 2
port trekreview Phase 5-6 to a Workflow only if the designed bake-off shows
fidelity-equivalent output + acceptable control/cost; tier 3 wholesale swap
declined (portability floor 2.1.154+, opt-in UX, mid-flow visibility loss).
Open risk inherited from S7: classifier at large fan-out under auto/bypass
still unverified.

New: docs/T2-cc27-workflow-substrate.md (gate evidence F0-F4 + bake-off design
with thresholds + no-Workflow schema-contract PoC). Matrix: CC-27 row +
S8 resolutions + open-question/T2 pointers updated. Docs-only; no code/schema.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 13:34:05 +02:00

231 lines
32 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# CC-upgrade decision matrix — Claude Code 2.1.130 → 2.1.181
**Scope:** Evaluate every relevant Claude Code change shipped between 2.1.130 and 2.1.181 (latest as of 2026-06-18) against Voyage v5.1.1, and decide adoption per change.
**Precedent:** This continues the F2F14 feature-adoption process referenced in the roadmap ("ny CC-versjon med relevant feature → vurder adopsjon analogt med F2F14-prosessen"). The F-catalogue came from the extracted `ultra-cc-architect` plugin and is no longer bundled; this matrix is a fresh, self-contained catalogue (`CC-NN`) for the 2.1.130→181 window.
**Method:** Two-track research (CC changelog digest + Voyage CC-capability surface inventory), synthesized in a single context. Sources: official changelog (`code.claude.com/docs/en/changelog.md`) + canonical `anthropics/claude-code` CHANGELOG.
## Provenance & verification (verifiseringsplikt)
- **Verified verbatim against the official changelog** (load-bearing claims): `2.1.172` sub-agent nesting; `2.1.154` Opus 4.8 + dynamic workflows; `2.1.166` SendMessage authority; `2.1.178` `Tool(param:value)` syntax; `effort:` frontmatter (`2.1.154`/`2.1.152`).
- **From changelog research, not independently re-verified line-by-line:** all other entries below. Where a decision *hinges* on an unverified claim, it is marked ⚠️ and routed to EVALUATE, never SHIP.
- **Note on 2.1.1302.1.135:** these versions are not in the public changelog (public window begins 2.1.136). No user-facing entries to evaluate.
## Decision legend
| Decision | Meaning |
|----------|---------|
| **SHIP** | Adopt now — clear fit, low risk, no open design question. |
| **EVALUATE** | Needs an empirical test or a design decision before ship/skip. The hard architectural items live here — honest, not deferred-by-another-name. |
| **DEFER** | Adopt later; trigger noted. Not weekend-critical. |
| **SKIP** | Not applicable or rejected; reason noted. |
## Workstream legend
| WS | Theme | Weekend-shippable? |
|----|-------|--------------------|
| **W0** | Correctness/trust — fix now-false documentation | Yes |
| **W1** | Orchestration architecture — delegated spawning / Workflow adoption | No (multi-week, empirical) |
| **W2** | Model & effort alignment (Opus 4.8 + native `effort:`) — **gates v5.4** | Partial |
| **W3** | Guardrails & hooks | Partial |
| **W4** | Free wins & release hygiene | Yes |
---
## W0 — Correctness / trust (the headline)
| ID | Change (version) | Type | Voyage relevance | Decision | Rationale |
|----|------------------|------|------------------|----------|-----------|
| **CC-01** | Sub-agents can spawn their own sub-agents, up to 5 levels deep (**2.1.172**, verified); foreground subagents respect the same depth cap (2.1.181) | BREAKING-premise | Voyage's v2.4.0 inline-orchestration migration is justified in 4 places by the claim "the harness does not expose the Agent tool to sub-agents" → `agents/planning-orchestrator.md:511`, `agents/research-orchestrator.md:510`, `agents/review-orchestrator.md:512,:220`, `commands/trekplan.md:399406`. **That claim is now factually false.** | **SHIP** (doc correction) | Independent of the architecture decision (W1), the docs assert an impossibility that is now possible. Correct the four sites to state the factual position: as of CC 2.1.172 sub-agents *can* spawn sub-agents (≤5 deep); Voyage currently still orchestrates inline; re-architecture is under evaluation (W1/CC-26). Do **not** silently delete — replace with truth + forward pointer. The *design* response is CC-26 (EVALUATE). |
---
## W4 — Free wins & release hygiene (weekend-shippable)
| ID | Change (version) | Type | Voyage relevance | Decision | Rationale |
|----|------------------|------|------------------|----------|-----------|
| **CC-02** | MCP `tools/list` pagination fixes (2.1.144/147); `MCP_TOOL_TIMEOUT` raises remote fetch timeout (2.1.142); sub-1000ms `timeout` now ignored→default (2.1.162) | FIX | research agents (`docs-researcher`, `*-researcher`) use tavily/ms-learn/gemini MCP | **SHIP** (verify) | Automatic benefit — no code change. Verify research agents still enumerate tools correctly under paginated servers. |
| **CC-03** | `SendMessage` relayed messages lose user authority (**2.1.166**, verified) | BREAKING | Voyage uses `SendMessage` **nowhere** (confirmed by inventory) | **SKIP** | No impact. Recorded so a future `SendMessage`-based design knows the constraint up front. |
| **CC-04** | Subagent frontmatter `mcpServers` now enforce `--strict-mcp-config` + enterprise allow/deny (2.1.153) | CHANGE | Voyage ships **no `.mcp.json`** and declares MCP *tool grants* (`mcp__tavily__…`) not server *configs*; agents degrade gracefully when servers absent | **EVALUATE** (low) | Likely no impact, but confirm research agents degrade cleanly under `--strict-mcp-config`. One quick test. |
| **CC-05** | Fable 5 `[1m]` model-name suffix normalized automatically (2.1.173) | FIX | Voyage pins the `opus` alias, not raw ids | **SKIP** | No action; noted for completeness. |
| **CC-06** | `claude plugin validate` flags `skills:` pointing at a file vs dir (2.1.145); richer `claude plugin details` (2.1.143/145) | NEW | release process | **SHIP** (process) | Add `claude plugin validate` to the pre-release checklist. Voyage ships no skills, but validate also checks manifest/components. |
| **CC-07** | `fallbackModel` setting — up to 3 ordered fallbacks; `--fallback-model` now applies to interactive too (2.1.166) | NEW | `trekexecute` headless children (`claude -p`) could gain resilience to model-overload | **DEFER** | Real value for long headless waves. Trigger: when we touch `trekexecute` Phase 2.6 launch flags (or in W2 model work). Not weekend-critical. |
| **CC-08** | Hooks may not fire reliably in headless child sessions — Voyage documents GH #36071 (`trekexecute.md:341`, `templates/headless-launch-template.md:48`) | (status unknown) | safety-preamble is the headless defense if hooks don't run | **EVALUATE** ⚠️ | Changelog does **not** confirm #36071 is fixed. Verify current status before relaxing the in-prompt safety preamble. Until verified: keep the preamble. |
| **CC-09** | `--safe-mode` / `CLAUDE_CODE_SAFE_MODE` — start with all customizations disabled (2.1.169) | NEW | dev/test aid — reproduce "bare harness" behavior | **DEFER** | Useful for regression-testing Voyage against an unconfigured harness. Adopt into test tooling when convenient. |
---
## W3 — Guardrails & hooks
| ID | Change (version) | Type | Voyage relevance | Decision | Rationale |
|----|------------------|------|------------------|----------|-----------|
| **CC-10** | `Tool(param:value)` permission syntax, e.g. `Agent(model:opus)` (**2.1.178**, verified) | NEW | Voyage pins `opus` in 23 agent frontmatters *by convention*; this could enforce it as a hard permission rule | **DECIDED (design note)** ⚠️ | **Trap:** a blanket `Agent(model:opus)`-style rule conflicts with the `balanced`/`economy` profiles, which deliberately spawn `sonnet`. Any enforcement must be **profile-aware** → tie to W2. Do NOT ship a blanket opus-lock. **S2 resolution: no rule added; deferred into W2** — see §S2 resolutions. |
| **CC-11** | `disallowed-tools` in command/skill frontmatter (2.1.152, verified) | NEW | tighten per-command tool surfaces (e.g. ensure `trekexecute` can't spawn Agents — it already documents "no Agent tool") | **SHIPPED ✅** | Promote existing *documented* tool exclusions into *enforced* `disallowed-tools`. Small, defense-in-depth. **S2: added `disallowed-tools: Agent, TeamCreate` to `trekexecute` frontmatter** (the only command with a documented exclusion). Verified `allowed-tools` omission does NOT remove a tool from the pool — `disallowed-tools` does. |
| **CC-12** | Stop/SubagentStop hooks can return `hookSpecificOutput.additionalContext` without being a hook error (2.1.163) | NEW | `post-compact-flush.mjs` already emits `additionalContext`; SubagentStop is a new surface | **DEFER** | Possible richer continuity injection. No current gap forces it. |
| **CC-13** | Stop/SubagentStop hook input now includes `background_tasks` + `session_crons` (2.1.145) | NEW | `otel-export.mjs` (Stop hook) observability | **DEFER** | Could enrich exported telemetry. Adopt when next touching the exporter. |
| **CC-14** | Hook `args: string[]` exec-form — no shell, no quoting (2.1.139) | NEW | Voyage's 7 hooks invoke node scripts via shell form with `${CLAUDE_PLUGIN_ROOT}` paths | **SHIPPED ✅** | Exec-form removes a class of path-quoting bugs. **S2: migrated all 7 hooks** in `hooks/hooks.json` to `{command:"node", args:["${CLAUDE_PLUGIN_ROOT}/…"]}`. Official hooks doc recommends exec-form *"whenever the hook references a path placeholder"* — protects consumers who install Voyage under a path containing spaces. `${CLAUDE_PLUGIN_ROOT}` interpolates in `args` (verified). |
| **CC-15** | Hook `if:` conditions for Read/Edit/Write paths now match reliably (2.1.139/176); reopens deferred **F2** (scope pre-bash/pre-write executors to execute sessions) | FIX | `pre-bash-executor.mjs`, `pre-write-executor.mjs` are currently universal | **DECIDED (keep universal)** | F2 was deferred with "universal protection wins." The `if:` mechanism now works, so the *option* is real again — but the original rationale holds. **S2 resolution: KEEP UNIVERSAL.** These guardrails (rm -rf /, fork bombs; writes to ~/.ssh, .env, .git/hooks) are session-agnostic safety; narrowing to execute-only would only weaken protection with no benefit and never interferes with brief/research/plan work. Header comments corrected to state the universal scope. |
| **CC-16** | `SessionStart` `reloadSkills` + `sessionTitle` (2.1.152) | NEW | Voyage sets session title via `UserPromptSubmit` (`session-title.mjs`) | **SKIP** | Current mechanism works and is command-scoped (title reflects the invoked `/trek*` command). SessionStart-title would fire before the command is known. No gain. |
| **CC-17** | Hook `terminalSequence` output — notifications/bells without a TTY (2.1.141); `continueOnBlock` for PostToolUse (2.1.139); `MessageDisplay` event (2.1.152); Stop block-cap 8 (2.1.143) | NEW | minor ergonomics; Voyage hooks are fail-open and non-interactive | **SKIP/DEFER** | No current need. `MessageDisplay`/`continueOnBlock` SKIP (no use case); `terminalSequence` DEFER (could notify on long headless waves). Block-cap is informational. |
---
## S2 resolutions (W3 hardening — 2026-06-18)
S2 shipped the mechanical W3 items and recorded the two W3 decisions. Schemas verified verbatim against the official slash-commands and hooks docs before any edit (the casing of `disallowed-tools` and exec-form `${CLAUDE_PLUGIN_ROOT}` interpolation are both load-bearing; a first-pass assistant claim of `disallowedTools` camelCase was caught and corrected against the doc).
- **CC-14 — SHIPPED.** All 7 hooks in `hooks/hooks.json` migrated to exec-form (`command:"node"`, `args:["${CLAUDE_PLUGIN_ROOT}/hooks/scripts/X.mjs"]`). `tests/hooks/hooks-json-stop-wired.test.mjs` updated to be form-agnostic (normalizes `command`+`args` to one invocation string). 578/580 pass, 0 fail.
- **CC-11 — SHIPPED.** `disallowed-tools: Agent, TeamCreate` added to `commands/trekexecute.md`, promoting its documented "No Agent tool, no TeamCreate, no delegation" rule from prose to enforcement. Key fact: `allowed-tools` grants auto-approval but does **not** remove tools from the pool (doc, slash-commands.md), so the prior omission left Agent callable; `disallowed-tools` removes it. trekexecute is the only command with a documented exclusion (trekplan's "never spawn more agents than warranted" is an adaptive-count guideline, not an exclusion — trekplan legitimately uses Agent).
- **CC-15 — DECIDED (keep universal).** F2 stays deferred: `pre-bash-executor.mjs` / `pre-write-executor.mjs` remain universal. Rationale re-affirmed explicitly now that `if:` works — narrowing to execute-only sessions would weaken session-agnostic safety with no benefit. Header comments corrected to state the universal scope by design.
- **CC-10 — DECIDED (design note; no code).** Do **not** add a `permissions.deny` rule like `Agent(model:opus)`. Such rules live in `settings.json` permissions (colon separator inside parens, confirmed 2.1.178), not in frontmatter, and a blanket opus-lock would break the `balanced`/`economy` profiles that deliberately spawn `sonnet`. If model-enforcement is ever wanted, it must be **profile-aware** and emitted as part of the W2 model/effort work (S3/S4) — e.g. a profile could ship its own `availableModels`/deny set. Until then, the `model: opus` convention in the 23 agent frontmatters stays advisory. Folded into W2 open question #3.
## W2 — Model & effort alignment (gates v5.4)
| ID | Change (version) | Type | Voyage relevance | Decision | Rationale |
|----|------------------|------|------------------|----------|-----------|
| **CC-21** | **Opus 4.8** available; `opus` defaults to high/xhigh effort (**2.1.154**, verified) | NEW | All 23 agents + 7 commands pin `model: opus` → now resolve to 4.8 at a higher default effort | **DECIDED (S3)** | Accept Opus-4.8-`high` as the new baseline (verified: opus default = `high`). The moderation lever is native `effort:` (CC-22 option C). Doc follow-up: update model/effort references in `CLAUDE.md`/`docs/profiles.md`/README. See §S3 resolutions. |
| **CC-22** | Native `effort:` frontmatter on agents/skills/commands + 5 levels low/medium/high/xhigh/max (**2.1.149/154**, verified) | NEW | Voyage **reinvented** effort as `phase_signals` (low/standard/high) + the `phase-signal-resolver.mjs`/`resolver.mjs` system (which carries the MAJOR doc/code-inconsistency finding) | **DECIDED (S3): option C** | **Resolved.** The (a)-map premise was false: Voyage `effort` = *orchestration-shape* (which agents/passes/gates run, consumed by command prose), native `effort:` = *per-spawn reasoning budget***different axes, same name.** Decision: **freeze `phase_signals.effort` 3-level (low/standard/high) as-is** (unblocks v5.4) and adopt native `effort:` **additively at the agent/profile layer, outside the brief contract.** Field name kept `effort` (no breaking rename) + distinction documented. Does NOT retire the resolver MAJOR (that's model-gating, independent — S4). See §S3 resolutions. |
| **CC-23** | Claude Fable 5 (Mythos-class) GA (2.1.170) | NEW | Voyage is a planning/reasoning pipeline; Fable is a different model class | **SKIP** (revisit) | Not an obvious fit for deep-planning agents. Note as an option if a future profile wants a distinct model class for a specific phase. |
| **CC-24** | `enforceAvailableModels` managed setting (2.1.175); `availableModels` now constrains Default + subagent overrides (2.1.172/176) | CHANGE | enterprise deployments could constrain Voyage's `opus`/`sonnet` picks | **DEFER** | Document that profile model picks must live within any managed `availableModels` allowlist. Relevant for enterprise consumers post-open-release. |
| **CC-25** | `MAX_THINKING_TOKENS=0` / `--thinking disabled` disables thinking on think-by-default models (2.1.166) | NEW | could gate thinking in cheap `economy`-profile phases | **DEFER** | Minor cost lever; fold into W2 profile design if useful. |
---
## S3 resolutions (W2 effort/model alignment — 2026-06-18, operator-gated)
S3 was a decision-gate. Two-track evidence (codebase map of `phase_signals`/`phase_models`/profiles/resolver + verbatim-cited CC native-`effort:` semantics) overturned CC-22's framing hypothesis and the operator confirmed the path.
**The load-bearing finding — two axes, one name.** Voyage's `phase_signals.effort` (low/standard/high) is consumed only by **command prose** to pick *orchestration shape*: skip/add swarm passes, gate strictness, sequential-vs-parallel (`trekplan.md:925957`, `trekresearch.md:472503`, `trekreview.md:385417`, `trekexecute.md:15421577`). CC native `effort:` (low/medium/high/xhigh/max) is a *per-spawn reasoning-token budget* applied by the harness. `low` in Voyage means "run fewer agents," not "think less." A remap (option a) would conflate the two and silently delete orchestration behavior — and would **not** remove the resolver (it also carries the model half). Verified native semantics: `effort:` settable in agent/skill/command frontmatter; precedence env (`CLAUDE_CODE_EFFORT_LEVEL`) > frontmatter > session > model default; Opus 4.8 default = `high`.
**Operator decisions (2026-06-18):**
1. **CC-22 → option C.** Freeze `phase_signals.effort` 3-level as-is (it is an orchestration axis, deliberately not the native 5-level vocab) → **unblocks the v5.4 brief-schema freeze**. Adopt native `effort:` additively, at the agent/profile layer, outside the brief contract.
2. **Field name kept `effort`** (no breaking rename at v5.4); the orchestration-vs-reasoning distinction is documented loudly instead.
**Dispositions feeding S4 / v5.4 / profile design:**
- **CC-21 — DECIDED.** Opus-4.8-`high` is the accepted baseline; native `effort:` is the moderation lever. Doc-truth follow-up: model/effort references in `CLAUDE.md`/`docs/profiles.md`/README.
- **CC-24 — DEFER (confirmed).** `availableModels`/`enforceAvailableModels` constrains **model only**, not effort (verified). Document that profile model picks must stay within any enterprise allowlist.
- **CC-25 — DEFER (confirmed).** `MAX_THINKING_TOKENS=0` overrides effort (disables thinking) — the most aggressive `economy` cost lever; fold into profile design.
- **Resolver MAJOR (S4, unchanged).** `phase-signal-resolver.mjs:40` copies `model` ungated while `:39` gates `effort`. Independent of the effort decision; option C keeps `phase_signals.model`, so the fix (import + apply `BASE_ALLOWED_MODELS`) stays on S4.
**S4 scope (now unblocked):** (1) resolver MAJOR model-gate fix; (2) additive native `effort:` — verified-safe minimum is static per-agent frontmatter (retrieval agents `task-finder`/`git-historian`/`dependency-tracer`/`architecture-mapper` lower; reasoning agents `plan-critic`/`risk-assessor`/`contrarian-researcher`/`review-coordinator` stay high); (3) doc-truth model/effort updates; (4) document the effort-axis distinction. **Open (non-blocking):** profile-driven *dynamic* effort needs verification — per-spawn `effort` param is **unverified** (only `model` is documented per-spawn); dynamic would otherwise need env-var injection.
## S4 resolutions (W2 implementation — 2026-06-18)
S4 implemented the four scope items above. TDD: a failing test was written first for the resolver gate, then the fix; full suite green throughout.
- **Resolver MAJOR — FIXED.** `lib/profiles/phase-signal-resolver.mjs` now imports `BASE_ALLOWED_MODELS` from `profile-validator.mjs` and gates `model` (`if ('model' in entry && BASE_ALLOWED_MODELS.includes(entry.model))`), mirroring the `EFFORT_LEVELS` gate one line above. Out-of-allowlist models (e.g. `gpt-4`, `haiku`) are now dropped (treated as absent) instead of handed through to an agent spawn — defense-in-depth behind `brief-validator`'s validation-time check. 2 new tests in `tests/lib/phase-signal-resolver.test.mjs` (drops-invalid + keeps-valid). No circular import (`brief-validator` already imports the same symbol; `profile-validator` depends only on `util/`).
- **Native `effort:` — SHIPPED (static, additive).** Added `effort:` frontmatter to the 8 named agents: retrieval (`task-finder`, `git-historian`, `dependency-tracer`, `architecture-mapper`) → `medium`; adversarial-reasoning (`plan-critic`, `risk-assessor`, `contrarian-researcher`, `review-coordinator`) → `high`. The other 15 agents stay unset → inherit Opus-4.8 default (`high`). `claude plugin validate` passes (effort is a valid frontmatter field; only the pre-existing root-CLAUDE.md warning remains).
- **Doc-truth + axis distinction — DONE.** New canonical section `docs/profiles.md` §Model & effort axes (opus→Opus 4.8 default-`high`; orchestration `phase_signals.effort` vs native reasoning `effort:` table; native-effort precedence; which agents carry which level). Short note added to `CLAUDE.md` (after Agents table) and `README.md` (Cost profile section), both pointing to the profiles.md section.
- **Open (non-blocking, unchanged).** Only *static* per-agent effort shipped — the verified-safe minimum. Profile-driven *dynamic* effort still needs verification of the per-spawn `effort` param (only `model` is documented per-spawn) or env-var injection (`CLAUDE_CODE_EFFORT_LEVEL`). Deferred to W2 profile design.
- **Tests:** 582 total / 580 pass / 0 fail / 2 skip (was 578 pass; +2 new resolver tests).
## W1 — Orchestration architecture (multi-week, empirical)
| ID | Change (version) | Type | Voyage relevance | Decision | Rationale |
|----|------------------|------|------------------|----------|-----------|
| **CC-26** | Sub-agents spawn sub-agents ≤5 deep (**2.1.172**, verified) — the design response to CC-01 | NEW | Could restore *delegated* orchestration: an orchestrator sub-agent spawns the swarm; synthesis/writing delegated (the "missing summarizer link" in `docs/subagent-delegation-audit.md`). Frees main-context tokens | **EVALUATE → lean NO (S7)** | "Can spawn 5 deep" ≠ "Voyage's orchestrator→6-agent-swarm pattern performs well." **S7 (2026-06-18, GATE):** feasibility probed cheaply — depth-2 nesting works, no degradation; depth cap moot (Voyage needs depth 2); NEW finding = auto-mode proliferation classifier polices agent fan-out (risk unique to delegation). On cost/benefit, **wholesale delegation NOT recommended**; the only defensible path is a narrow opt-in synthesis-agent PoC, proven by Δ main-context tokens. Operator gates verdict. Full bake-off designed but NOT run. See `docs/T1-cc26-delegated-orchestration.md` + §S7. |
| **CC-27** | Dynamic Workflows / Workflow tool — orchestrates tenshundreds of agents (**2.1.154**, verified); keyword `workflow``ultracode` (2.1.160); `agent()` attribution headers (2.1.174) | NEW | Voyage **hand-rolls** swarm/wave/pipeline orchestration in command prose — the Workflow tool is a native primitive for exactly this | **EVALUATE → selective hybrid (S8)** | The biggest identity decision: adopt Workflow as substrate, or stay prose-orchestrated? **S8 (2026-06-18, GATE):** probed cheaply — a minimal trekreview-shaped `parallel()``agent()` Workflow ran end-to-end (F1), structured schemas retire the JSON-parse fragility at `trekreview.md:202204` (F2), result returns to main (F3), and a small purposeful fan-out did **not** trip the S7 proliferation classifier (F4). Reframe: "substrate swap" is a false binary — a `/trek*` command is ~80% non-orchestration glue, so Workflow can only ever replace the fan-out→synthesize *core* (hybrid). **Recommendation: selective hybrid, NOT wholesale swap** — tier 1 ship a prose schema contract (the F2 win, no Workflow dep); tier 2 port trekreview Phase 56 to a Workflow only if the designed bake-off shows fidelity-equivalent output + acceptable control/cost; tier 3 wholesale swap declined (portability floor 2.1.154+, opt-in UX, visibility loss). Operator gates verdict. Full bake-off designed, NOT run. See `docs/T2-cc27-workflow-substrate.md` + §S8. |
| **CC-28** | `TaskCreate` reliability — auto-repairs malformed input, schema in errors (2.1.163/169) | FIX | `TaskCreate`/`TaskUpdate` are in `trekplan`/orchestrator frontmatter but not actively used in command logic | **DEFER** | Becomes relevant only if W1 adopts task-graph orchestration. Tie to CC-26/27 outcome. |
| **CC-29** | `subagent_type` matching now case/separator-insensitive (2.1.140); multiple `Agent(...)` types in `tools:` no longer dropped (2.1.147); subagent transcript/backgrounding fixes (2.1.178) | FIX | improves DX of any delegated-orchestration design | **SHIP** (verify) | Free robustness. Confirm Voyage's agent `tools:` grants (none currently declare multiple `Agent(...)` types) and `subagent_type` references are unaffected. |
| **CC-31** | Worktree-isolation guard now applies in background sessions (2.1.154); `worktree.bgIsolation:"none"` (2.1.143); `EnterWorktree` switching mid-session (2.1.157) | CHANGE | `trekexecute` Phase 2.6 parallel waves + `trekplan` "execute with team" use git worktrees / `TeamCreate isolation:"worktree"` | **EVALUATE** | Verify Voyage's worktree-based parallel execution still behaves under the tightened bg-isolation guard. Affects the multi-session headless path. |
---
## S7 resolutions (W1 / CC-26 gate — 2026-06-18, operator-gated)
S7 was the first W1 gate. Operator chose a **staged** execution: cheap live feasibility probe +
measurement-design doc; the expensive head-to-head comparison was specified but **not run**.
- **Feasibility (measured).** A recursive `general-purpose` agent chain in this CC 2.1.181
interactive session confirmed: depth-2 nesting works (`main → L1 → L2`; both children report the
Agent tool available), and the nested sub-agent returned a real, well-formed result — **no silent
degradation** at depth 2. The v2.4.0 premise ("harness does not expose Agent to sub-agents") is
confirmed false at the interactive sub-agent level.
- **Depth cap moot.** The ≤5 cap was never reached; the recursion stopped at L2→L3 via the
**auto-mode permission classifier** (policy denial: "uncontrolled agent proliferation"), not the
nesting limit. Voyage's needed pattern is depth 2, so the cap does not bind this gate.
- **NEW finding — proliferation classifier.** Auto/bypass modes actively deny purposeless agent
fan-out. A delegated orchestrator spawning a 610-agent swarm from inside a sub-agent under
`auto`/`bypassPermissions` is exactly that shape — a classifier-interference risk **unique to
delegation** that inline orchestration does not carry, and a new silent-degradation surface if a
mid-pipeline spawn is denied. Must be in any future delegated-spawn test matrix.
- **CC-26 recommendation (operator gates verdict).** Lean **NO** on wholesale delegated
orchestration: feasibility is no longer the blocker, so the gate turns on cost/benefit, which is
unfavourable (delegation's only upside is main-context relief, against wall-time loss, context
re-delivery cost, the audit's iteration/adversarial-review/debuggability tradeoffs, and the new
classifier risk). The only defensible win is a **narrow opt-in synthesis-agent** (delegate just
trekplan Phase 7's heaviest inline read), adopted only if a measured Δ main-context ≥ 30% with no
quality loss materialises. CC-27 (Workflow tool, S8) is the more promising substrate question and
is untouched.
- **Artifact.** `docs/T1-cc26-delegated-orchestration.md` — full gate evidence, the §5 full
bake-off design (thresholds POSITIVE/NEGATIVE/INCONCLUSIVE), and the §6 synthesis-agent PoC, all
ready to run if the operator greenlights pursuing delegation.
## S8 resolutions (W1 / CC-27 gate — 2026-06-18, operator-gated)
S8 was the second W1 gate — the orchestration-substrate identity decision. Operator chose the same
**staged** execution as S7: cheap live feasibility probe + measurement-design doc; the head-to-head
prose-vs-Workflow bake-off was specified but **not run**.
- **Feasibility (measured).** A minimal trekreview-shaped Workflow — `parallel([reviewerA,
reviewerB])` with a findings schema → `agent(coordinator)` with a verdict schema, trivial agents,
synthetic input — ran end-to-end in this CC 2.1.181 interactive session: both reviewers returned
(F1), structured schemas delivered typed findings with no JSON-parse step (F2), the synthesizer
reproduced Phase-6 dedup+verdict behavior and the result returned to main (F3), and the small
purposeful fan-out did **not** trip the S7 proliferation classifier (F4). _3 agents · 85 461
tokens · 13.8 s._
- **Reframe — "substrate swap" is a false binary.** A `/trek*` command is ~80% non-orchestration
glue (mode parsing, triage, validators, stats, HTML) and ~20% agent fan-out. The Workflow tool
can only ever replace the fan-out→synthesize *core* (trekreview Phase 56), so the real decision
is a **scoped hybrid per core**, not a wholesale identity swap. (Mirrors S7's "wall-time is not
the gate metric" reframe.)
- **The one concrete win is schema robustness (F2)** — and it is capturable *without* the Workflow
tool: codify reviewer-output JSON as a validated prose schema contract, retiring the fragile
"collect trailing JSON / re-ask on parse error" at `trekreview.md:202204`.
- **CC-27 recommendation (operator gates verdict).** **Selective hybrid, NOT wholesale swap**, in
three tiers: (1) ship a prose schema contract regardless — the F2 win, zero new dependency; (2)
port trekreview Phase 56 to a Workflow *only if* the designed bake-off shows fidelity-equivalent
output + acceptable control/cost — best-case first port; (3) wholesale substrate swap **declined**
(portability floor 2.1.154+, opt-in/billing UX on every invocation, loss of mid-flow operator
visibility for an operator-gated review tool). CC-26 (S7) and CC-27 (S8) are now both resolved to
"narrow/selective, operator-gated, not wholesale."
- **Open risk inherited from S7.** Classifier behavior at *large* fan-out (trekplan's 610-agent
swarm) under `auto`/`bypass` is still unverified — trekreview's 23 agents are below threshold,
but a later trekplan port must measure it first.
- **Artifact.** `docs/T2-cc27-workflow-substrate.md` — full gate evidence (F0F4), §5 bake-off
design (thresholds POSITIVE/NEGATIVE/INCONCLUSIVE), §6 no-Workflow schema-contract PoC, all ready
to run if the operator greenlights tier 1 and/or tier 2.
## Sequencing
```
Weekend (W0 + W4): CC-01 doc-truth fix · CC-06 plugin validate · CC-02/CC-29 verify-no-regression
→ ship something important AND correct.
Near-term (W2): CC-21 + CC-22 effort/model alignment decision → GATES v5.4 (per operator decision).
Resolves the resolver MAJOR finding in the same pass.
Then: v5.4 brief-schema public contract → v5.5 brief framing enforcement.
Multi-week (W1): CC-26 empirical sub-agent-nesting test + CC-27 Workflow-adoption prototype.
Outcome rewrites CC-01's "current design" statement and feeds the
docs/subagent-delegation-audit.md open problem.
Incremental (W3): CC-11/CC-14/CC-15 as small, independently-shippable hardening PRs.
```
## Open questions (need operator or empirical answer)
1. **W1 identity:** does Voyage adopt the Workflow tool as substrate, or stay prose-orchestrated? (CC-27) — **S8: false binary; lean selective hybrid (port the fan-out→synthesize core only), NOT wholesale swap; see §S8.**
2. **W1 perf:** does delegated orchestration (orchestrator sub-agent → swarm) beat inline at Voyage's scale? (CC-26 — empirical) — **S7: feasibility YES, but lean NO on wholesale adoption; see §S7.**
3. **W2 effort model:** map `phase_signals` onto native `effort:`, or keep bespoke? (CC-22 — gates v5.4)
4. **CC-08:** is GH #36071 (hooks in headless) fixed? Determines whether the safety-preamble can relax.
## Empirical tests required
- **T1 (CC-26):** orchestrator-sub-agent spawns the planning swarm vs. inline baseline — wall-time, quality, token cost, depth-cap behavior. Harness: extend `scripts/q3-cache-prefix-experiment.mjs` pattern. — **S7: feasibility half RUN** (depth-2 works, no degradation, cap moot, proliferation-classifier risk found); **perf half DESIGNED, NOT run** (full bake-off + cheaper synthesis-agent PoC specified in `docs/T1-cc26-delegated-orchestration.md` §5/§6).
- **T2 (CC-27):** reimplement `/trekreview`'s reviewer swarm as a Workflow; compare control, cost, and output fidelity vs. prose orchestration. — **S8: feasibility RUN** (probe: core ports natively, schemas retire JSON fragility, small fan-out classifier-clean); **fidelity/control/cost bake-off DESIGNED, NOT run** (`docs/T2-cc27-workflow-substrate.md` §5/§6).
- **T3 (CC-04/CC-31):** research-agent MCP degradation under `--strict-mcp-config`; worktree parallel-wave behavior under tightened bg-isolation.
---
_Catalogue covers 31 evaluated changes (CC-01…CC-31; CC-18/19/20/30 folded into CC-17/CC-29 rows). Generated 2026-06-18 against Voyage v5.1.1 / CC 2.1.181._