Commit graph

6 commits

Author SHA1 Message Date
b5f3d4a932 docs(voyage): S8 (W1/CC-27 gate) — T2 Workflow-substrate probe + measurement design
Second W1 gate, same staged execution as S7 (operator-chosen): cheap live
feasibility probe + design doc; the prose-vs-Workflow bake-off specified but
NOT run.

Probe (CC 2.1.181 interactive): a minimal trekreview-shaped Workflow —
parallel([reviewerA, reviewerB]) with a findings schema -> agent(coordinator)
with a verdict schema — ran end-to-end. F1 core ports natively; F2 structured
schemas retire the JSON-parse fragility at trekreview.md:202-204; F3 result
returns to main; F4 a small purposeful fan-out did NOT trip the S7 proliferation
classifier. 3 agents / 85461 tokens / 13.8s.

Reframe: 'substrate swap' is a false binary — a /trek* command is ~80%
non-orchestration glue, so Workflow can only replace the fan-out->synthesize
core (hybrid).

CC-27 recommendation (operator gates verdict): selective hybrid, NOT wholesale
swap. Tier 1 ship a prose schema contract (the F2 win, no Workflow dep); tier 2
port trekreview Phase 5-6 to a Workflow only if the designed bake-off shows
fidelity-equivalent output + acceptable control/cost; tier 3 wholesale swap
declined (portability floor 2.1.154+, opt-in UX, mid-flow visibility loss).
Open risk inherited from S7: classifier at large fan-out under auto/bypass
still unverified.

New: docs/T2-cc27-workflow-substrate.md (gate evidence F0-F4 + bake-off design
with thresholds + no-Workflow schema-contract PoC). Matrix: CC-27 row +
S8 resolutions + open-question/T2 pointers updated. Docs-only; no code/schema.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 13:34:05 +02:00
cccc535a13 docs(voyage): S7 (W1/CC-26 gate) — T1 feasibility probe + measurement design
Staged gate execution (operator-chosen): cheap live feasibility probe + design
doc; expensive head-to-head specified but NOT run.

Probe (CC 2.1.181 interactive): depth-2 sub-agent nesting works (main->L1->L2,
both have Agent tool), no degradation; v2.4.0 'no Agent for sub-agents' premise
confirmed false. Depth cap (<=5) moot for Voyage (needs depth 2). NEW finding:
auto-mode proliferation classifier polices agent fan-out — a classifier-
interference risk unique to delegation.

CC-26 recommendation (operator gates verdict): lean NO on wholesale delegated
orchestration; only defensible path is a narrow opt-in synthesis-agent PoC
proven by delta main-context tokens. CC-27 (Workflow, S8) untouched.

New: docs/T1-cc26-delegated-orchestration.md (gate evidence + full bake-off
design with thresholds + cheaper synthesis-agent PoC). Matrix: CC-26 row +
S7 resolutions + open-question/T1 pointers updated. Docs-only; no code/schema.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 13:21:09 +02:00
fdd3ad80d7 feat(voyage): W2 impl (S4) — gate resolver model, adopt native effort:, doc-truth
S4 of the 2.1.181 upgrade — implementation, not a gate. TDD: failing test
written first for the resolver gate, then the fix; suite green throughout.

- Resolver MAJOR (FIX): lib/profiles/phase-signal-resolver.mjs now imports
  BASE_ALLOWED_MODELS from profile-validator and gates `model`
  (if 'model' in entry && BASE_ALLOWED_MODELS.includes(entry.model)),
  mirroring the EFFORT_LEVELS gate one line up. Out-of-allowlist models
  (gpt-4, haiku) are dropped instead of handed to an agent spawn —
  defense-in-depth behind brief-validator's validation-time check. No
  circular import (brief-validator already imports the same symbol).
  +2 tests (drops-invalid / keeps-valid).
- Native effort: (SHIP, static additive): effort: frontmatter on 8 agents —
  retrieval (task-finder, git-historian, dependency-tracer,
  architecture-mapper) = medium; adversarial-reasoning (plan-critic,
  risk-assessor, contrarian-researcher, review-coordinator) = high. The
  other 15 stay unset -> inherit Opus-4.8 default (high). This per-spawn
  REASONING effort is a different axis from brief phase_signals.effort
  (ORCHESTRATION shape) per the S3 decision.
- Doc-truth + axis distinction: new canonical docs/profiles.md
  §Model & effort axes (opus->Opus 4.8 default-high; orchestration vs
  reasoning effort table; native-effort precedence; per-agent levels).
  Short notes in CLAUDE.md (after Agents table) and README.md (Cost
  profile), both pointing to profiles.md.
- Open (non-blocking, unchanged): only STATIC effort shipped — the
  verified-safe minimum. Profile-driven DYNAMIC effort still needs
  verification of the per-spawn effort param or env-var injection.

Matrix: new "S4 resolutions" section. Tests 582 total / 580 pass / 0 fail /
2 skip (was 578 pass; +2). claude plugin validate passes (only pre-existing
root-CLAUDE.md warning).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 12:27:07 +02:00
fbf83f2271 docs(voyage): record S3 (W2) effort/model decision — option C, freeze brief-effort, keep field name
S3 was a decision-gate. Two-track evidence (codebase map of phase_signals/
phase_models/profiles/resolver + verbatim-cited CC native-effort: semantics)
overturned CC-22's framing and the operator confirmed the path.

Load-bearing finding: Voyage phase_signals.effort (low/standard/high) is an
ORCHESTRATION-SHAPE axis consumed by command prose (which agents/passes/gates
run), while CC native effort: is a per-spawn REASONING budget. Same name,
different axes. A remap would conflate them and silently delete orchestration
behavior, and would not remove the resolver (it also carries the model half).

Operator decisions (2026-06-18):
- CC-22 -> option C: freeze phase_signals.effort 3-level as-is (unblocks the
  v5.4 brief-schema freeze); adopt native effort: additively at the
  agent/profile layer, OUTSIDE the brief contract.
- Keep field name `effort` (no breaking rename); document the
  orchestration-vs-reasoning distinction loudly instead.

Dispositions: CC-21 DECIDED (Opus-4.8-high baseline accepted; native effort:
is the moderation lever; doc-truth follow-up). CC-24/CC-25 DEFER confirmed
(availableModels constrains model only, not effort; MAX_THINKING_TOKENS=0 is
an economy lever). Resolver MAJOR (phase-signal-resolver.mjs:40 ungated model)
stays on S4 — independent of the effort decision.

Matrix: CC-21/CC-22 rows flipped to DECIDED + new "S3 resolutions" section with
S4 scope and the non-blocking open (per-spawn effort param unverified).
Tests 578 pass / 0 fail / 2 skip.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 12:14:59 +02:00
66b3b15fb6 chore(voyage): W3 hardening (S2) — exec-form hooks, enforced disallowed-tools, F2 decision
S2 of the 2.1.181 upgrade. Schemas verified verbatim against the official
slash-commands and hooks docs before editing (a first-pass camelCase
'disallowedTools' claim was caught and corrected to kebab-case against the doc).

- CC-14 (SHIP): migrate all 7 hooks in hooks/hooks.json to exec-form
  {command:"node", args:["${CLAUDE_PLUGIN_ROOT}/hooks/scripts/X.mjs"]}. Doc
  recommends exec-form whenever a hook references a path placeholder; protects
  consumers installing under a path with spaces. ${CLAUDE_PLUGIN_ROOT}
  interpolates inside args (verified). hooks-json-stop-wired test made
  form-agnostic (normalizes command+args to one invocation string).
- CC-11 (SHIP): add `disallowed-tools: Agent, TeamCreate` to trekexecute
  frontmatter, enforcing its documented "No Agent tool, no TeamCreate" rule.
  allowed-tools grants auto-approval but does NOT remove tools from the pool,
  so the prior omission left Agent callable; disallowed-tools removes it.
  trekexecute is the only command with a documented exclusion.
- CC-15 (DECIDE: keep universal): re-affirm F2 deferral. pre-bash/pre-write
  executors stay universal -- session-agnostic safety (rm -rf /, ~/.ssh, .env)
  that narrowing to execute-only would only weaken. Header comments corrected.
- CC-10 (DECIDE: design note, no code): no blanket Agent(model:opus) deny rule
  -- would break balanced/economy profiles; any model-enforcement must be
  profile-aware, deferred into W2. Folded into open question #3.

Matrix updated with S2 resolutions section. Tests 578 pass / 0 fail / 2 skip;
claude plugin validate passes (only pre-existing root-CLAUDE.md warning).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 12:03:11 +02:00
c452f75628 docs(voyage): add CC 2.1.130->181 upgrade decision matrix
Two-track research (CC changelog digest x Voyage CC-capability surface) synthesized into a 31-entry adoption catalogue (CC-01..CC-31) across 5 workstreams: W0 correctness, W1 orchestration, W2 model/effort, W3 guardrails, W4 hygiene. Load-bearing changelog claims verified verbatim against the official changelog. Foundation for the continuous-session upgrade plan. Headline: CC 2.1.172 (sub-agents can spawn sub-agents) invalidates Voyage's inline-orchestration premise.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LqBYc8Ltrk7LipyJmGxXiB
2026-06-18 11:43:16 +02:00