Commit graph

8 commits

Author SHA1 Message Date
e7a58b0082 docs(voyage): S20 — CC-04/T3 verified clean (research-agent MCP degradation under --strict-mcp-config)
T3's CC-04 half: confirmed research agents degrade cleanly under
--strict-mcp-config. Verified semantics (CC cli-reference.md +
sub-agents.md, v2.1.153) via claude-code-guide: the flag strips all MCP
servers absent --mcp-config; absent MCP tool grants are skipped with a
warning and the subagent launches with its remaining tools (no hard
spawn error).

4/5 MCP-granting research agents (docs/community/security/contrarian)
keep a native WebSearch/WebFetch fallback and degrade by losing MCP
enhancement only. gemini-bridge is the lone MCP-only agent: under
--strict-mcp-config it spawns tool-less (no-op) but does not hard-fail,
is conditionally gated (--local skips; only high-effort forces it
always-on), and the existing graceful-degradation rule already names
Gemini. Disposition: VERIFIED clean, no code change (mirrors CC-08/29/31
prose dispositions; CC-31 worktree half already VERIFIED aligned -> T3
fully evaluated).

Recorded as §S20 resolution + inline T3 status; matrix rows and
generation stamp untouched. Optional hardening (gate gemini-bridge spawn
on gemini-server availability; pin the fallback invariant in
agent-frontmatter.test.mjs) recorded as a forward pointer, not done.

node --test green: 713/711 pass/2 skip/0 fail. This edit = zero test
delta (verified via git stash vs HEAD). Note: actual runtime count is
713, not the 698 recorded in S19's body/STATE — stale figure, not a
regression; census top-level split (behavior=601, pins=69, total=670)
still holds.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-19 14:08:51 +02:00
056720e7f3 feat(voyage): S13 — RELEASE v5.5.0 (framing 2.2 badge + NW1–NW3 + W2/W3 roll-up)
The coordinated release held since S6. Bumps 5.1.1 → 5.5.0 across plugin.json,
package.json, package-lock.json, README badge, and CHANGELOG (operator-confirmed
version; matches the codebase-wide v5.4/v5.5 milestone labels). Additive — no
breaking change for existing consumers; new brief requirements gate only on
briefs that declare brief_version 2.2.

Lands:
- brief_version 2.2 framing enforcement (framing enum + memory-alignment dim +
  obligatory TL;DR), held since S6 — badge now bumped
- Handover 1 PUBLIC CONTRACT formalization (the unreleased "v5.4")
- W2: Opus 4.8 baseline + native effort: on 8 agents + resolver model-gate fix
- W3: exec-form hooks (CC-14) + disallowed-tools on trekexecute (CC-11)
- NW1 reviewer-output schema contract; NW2 --workflow opt-in (bake-off POSITIVE);
  NW3 synthesis-agent dormant (declined per measurement)

CC 2.1.130→181 dispositions (docs/cc-upgrade-2.1.181-decision-matrix.md §S13):
- CC-08: GH#36071 (hooks in headless) CLOSED AS NOT PLANNED, not fixed →
  in-prompt safety preamble retained
- CC-29/31: verified clean / aligned, no code change
- CC-07/12/13: DEFER confirmed (recorded, no code)
- CLAUDE.md root warning: ACCEPTED BY DESIGN (universal across marketplace,
  advisory only; repo/maintainer context, not shipped consumer context)

TDD: version-consistency + v5.5.0-entry tests added first (RED→GREEN).
Tests 695 → 697 (695 pass / 2 skip / 0 fail). `claude plugin validate` passes
(one accepted advisory warning).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 18:12:51 +02:00
b5f3d4a932 docs(voyage): S8 (W1/CC-27 gate) — T2 Workflow-substrate probe + measurement design
Second W1 gate, same staged execution as S7 (operator-chosen): cheap live
feasibility probe + design doc; the prose-vs-Workflow bake-off specified but
NOT run.

Probe (CC 2.1.181 interactive): a minimal trekreview-shaped Workflow —
parallel([reviewerA, reviewerB]) with a findings schema -> agent(coordinator)
with a verdict schema — ran end-to-end. F1 core ports natively; F2 structured
schemas retire the JSON-parse fragility at trekreview.md:202-204; F3 result
returns to main; F4 a small purposeful fan-out did NOT trip the S7 proliferation
classifier. 3 agents / 85461 tokens / 13.8s.

Reframe: 'substrate swap' is a false binary — a /trek* command is ~80%
non-orchestration glue, so Workflow can only replace the fan-out->synthesize
core (hybrid).

CC-27 recommendation (operator gates verdict): selective hybrid, NOT wholesale
swap. Tier 1 ship a prose schema contract (the F2 win, no Workflow dep); tier 2
port trekreview Phase 5-6 to a Workflow only if the designed bake-off shows
fidelity-equivalent output + acceptable control/cost; tier 3 wholesale swap
declined (portability floor 2.1.154+, opt-in UX, mid-flow visibility loss).
Open risk inherited from S7: classifier at large fan-out under auto/bypass
still unverified.

New: docs/T2-cc27-workflow-substrate.md (gate evidence F0-F4 + bake-off design
with thresholds + no-Workflow schema-contract PoC). Matrix: CC-27 row +
S8 resolutions + open-question/T2 pointers updated. Docs-only; no code/schema.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 13:34:05 +02:00
cccc535a13 docs(voyage): S7 (W1/CC-26 gate) — T1 feasibility probe + measurement design
Staged gate execution (operator-chosen): cheap live feasibility probe + design
doc; expensive head-to-head specified but NOT run.

Probe (CC 2.1.181 interactive): depth-2 sub-agent nesting works (main->L1->L2,
both have Agent tool), no degradation; v2.4.0 'no Agent for sub-agents' premise
confirmed false. Depth cap (<=5) moot for Voyage (needs depth 2). NEW finding:
auto-mode proliferation classifier polices agent fan-out — a classifier-
interference risk unique to delegation.

CC-26 recommendation (operator gates verdict): lean NO on wholesale delegated
orchestration; only defensible path is a narrow opt-in synthesis-agent PoC
proven by delta main-context tokens. CC-27 (Workflow, S8) untouched.

New: docs/T1-cc26-delegated-orchestration.md (gate evidence + full bake-off
design with thresholds + cheaper synthesis-agent PoC). Matrix: CC-26 row +
S7 resolutions + open-question/T1 pointers updated. Docs-only; no code/schema.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 13:21:09 +02:00
fdd3ad80d7 feat(voyage): W2 impl (S4) — gate resolver model, adopt native effort:, doc-truth
S4 of the 2.1.181 upgrade — implementation, not a gate. TDD: failing test
written first for the resolver gate, then the fix; suite green throughout.

- Resolver MAJOR (FIX): lib/profiles/phase-signal-resolver.mjs now imports
  BASE_ALLOWED_MODELS from profile-validator and gates `model`
  (if 'model' in entry && BASE_ALLOWED_MODELS.includes(entry.model)),
  mirroring the EFFORT_LEVELS gate one line up. Out-of-allowlist models
  (gpt-4, haiku) are dropped instead of handed to an agent spawn —
  defense-in-depth behind brief-validator's validation-time check. No
  circular import (brief-validator already imports the same symbol).
  +2 tests (drops-invalid / keeps-valid).
- Native effort: (SHIP, static additive): effort: frontmatter on 8 agents —
  retrieval (task-finder, git-historian, dependency-tracer,
  architecture-mapper) = medium; adversarial-reasoning (plan-critic,
  risk-assessor, contrarian-researcher, review-coordinator) = high. The
  other 15 stay unset -> inherit Opus-4.8 default (high). This per-spawn
  REASONING effort is a different axis from brief phase_signals.effort
  (ORCHESTRATION shape) per the S3 decision.
- Doc-truth + axis distinction: new canonical docs/profiles.md
  §Model & effort axes (opus->Opus 4.8 default-high; orchestration vs
  reasoning effort table; native-effort precedence; per-agent levels).
  Short notes in CLAUDE.md (after Agents table) and README.md (Cost
  profile), both pointing to profiles.md.
- Open (non-blocking, unchanged): only STATIC effort shipped — the
  verified-safe minimum. Profile-driven DYNAMIC effort still needs
  verification of the per-spawn effort param or env-var injection.

Matrix: new "S4 resolutions" section. Tests 582 total / 580 pass / 0 fail /
2 skip (was 578 pass; +2). claude plugin validate passes (only pre-existing
root-CLAUDE.md warning).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 12:27:07 +02:00
fbf83f2271 docs(voyage): record S3 (W2) effort/model decision — option C, freeze brief-effort, keep field name
S3 was a decision-gate. Two-track evidence (codebase map of phase_signals/
phase_models/profiles/resolver + verbatim-cited CC native-effort: semantics)
overturned CC-22's framing and the operator confirmed the path.

Load-bearing finding: Voyage phase_signals.effort (low/standard/high) is an
ORCHESTRATION-SHAPE axis consumed by command prose (which agents/passes/gates
run), while CC native effort: is a per-spawn REASONING budget. Same name,
different axes. A remap would conflate them and silently delete orchestration
behavior, and would not remove the resolver (it also carries the model half).

Operator decisions (2026-06-18):
- CC-22 -> option C: freeze phase_signals.effort 3-level as-is (unblocks the
  v5.4 brief-schema freeze); adopt native effort: additively at the
  agent/profile layer, OUTSIDE the brief contract.
- Keep field name `effort` (no breaking rename); document the
  orchestration-vs-reasoning distinction loudly instead.

Dispositions: CC-21 DECIDED (Opus-4.8-high baseline accepted; native effort:
is the moderation lever; doc-truth follow-up). CC-24/CC-25 DEFER confirmed
(availableModels constrains model only, not effort; MAX_THINKING_TOKENS=0 is
an economy lever). Resolver MAJOR (phase-signal-resolver.mjs:40 ungated model)
stays on S4 — independent of the effort decision.

Matrix: CC-21/CC-22 rows flipped to DECIDED + new "S3 resolutions" section with
S4 scope and the non-blocking open (per-spawn effort param unverified).
Tests 578 pass / 0 fail / 2 skip.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 12:14:59 +02:00
66b3b15fb6 chore(voyage): W3 hardening (S2) — exec-form hooks, enforced disallowed-tools, F2 decision
S2 of the 2.1.181 upgrade. Schemas verified verbatim against the official
slash-commands and hooks docs before editing (a first-pass camelCase
'disallowedTools' claim was caught and corrected to kebab-case against the doc).

- CC-14 (SHIP): migrate all 7 hooks in hooks/hooks.json to exec-form
  {command:"node", args:["${CLAUDE_PLUGIN_ROOT}/hooks/scripts/X.mjs"]}. Doc
  recommends exec-form whenever a hook references a path placeholder; protects
  consumers installing under a path with spaces. ${CLAUDE_PLUGIN_ROOT}
  interpolates inside args (verified). hooks-json-stop-wired test made
  form-agnostic (normalizes command+args to one invocation string).
- CC-11 (SHIP): add `disallowed-tools: Agent, TeamCreate` to trekexecute
  frontmatter, enforcing its documented "No Agent tool, no TeamCreate" rule.
  allowed-tools grants auto-approval but does NOT remove tools from the pool,
  so the prior omission left Agent callable; disallowed-tools removes it.
  trekexecute is the only command with a documented exclusion.
- CC-15 (DECIDE: keep universal): re-affirm F2 deferral. pre-bash/pre-write
  executors stay universal -- session-agnostic safety (rm -rf /, ~/.ssh, .env)
  that narrowing to execute-only would only weaken. Header comments corrected.
- CC-10 (DECIDE: design note, no code): no blanket Agent(model:opus) deny rule
  -- would break balanced/economy profiles; any model-enforcement must be
  profile-aware, deferred into W2. Folded into open question #3.

Matrix updated with S2 resolutions section. Tests 578 pass / 0 fail / 2 skip;
claude plugin validate passes (only pre-existing root-CLAUDE.md warning).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 12:03:11 +02:00
c452f75628 docs(voyage): add CC 2.1.130->181 upgrade decision matrix
Two-track research (CC changelog digest x Voyage CC-capability surface) synthesized into a 31-entry adoption catalogue (CC-01..CC-31) across 5 workstreams: W0 correctness, W1 orchestration, W2 model/effort, W3 guardrails, W4 hygiene. Load-bearing changelog claims verified verbatim against the official changelog. Foundation for the continuous-session upgrade plan. Headline: CC 2.1.172 (sub-agents can spawn sub-agents) invalidates Voyage's inline-orchestration premise.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 11:43:16 +02:00