Compare commits
No commits in common. "main" and "v5.7.0" have entirely different histories.
95 changed files with 665 additions and 6329 deletions
|
|
@ -1,23 +1,12 @@
|
|||
{
|
||||
"name": "voyage",
|
||||
"description": "Voyage — brief, research, plan, execute, review, continue. Contract-driven Claude Code pipeline. /trekbrief, /trekplan, and /trekreview each end by building a self-contained operator-annotation HTML (scripts/annotate.mjs, modelled on claude-code-100x): select text or click any element, pick intent (Fiks/Endre/Spørsmål), write comment, copy structured prompt, paste back, Claude revises the .md.",
|
||||
"version": "5.9.1",
|
||||
"version": "5.7.0",
|
||||
"author": {
|
||||
"name": "Kjell Tore Guttormsen"
|
||||
},
|
||||
"homepage": "https://git.fromaitochitta.com/open/ktg-plugin-marketplace/src/branch/main/plugins/voyage",
|
||||
"repository": "https://git.fromaitochitta.com/open/ktg-plugin-marketplace.git",
|
||||
"license": "MIT",
|
||||
"keywords": [
|
||||
"voyage",
|
||||
"trek",
|
||||
"planning",
|
||||
"implementation",
|
||||
"research",
|
||||
"context-engineering",
|
||||
"agents",
|
||||
"adversarial-review",
|
||||
"headless",
|
||||
"execution"
|
||||
]
|
||||
"keywords": ["voyage", "trek", "planning", "implementation", "research", "context-engineering", "agents", "adversarial-review", "headless", "execution"]
|
||||
}
|
||||
|
|
|
|||
8
.gitignore
vendored
8
.gitignore
vendored
|
|
@ -19,10 +19,8 @@ blob-report/
|
|||
# Local configuration / session files
|
||||
*.local.*
|
||||
|
||||
# STATE.md — current state-of-play. LOCAL-ONLY per ~/.claude/CLAUDE.md:
|
||||
# origin is open/ (an OFFENTLIG/public mirror) → STATE must NEVER be pushed there.
|
||||
# Kept local for continuity only. History scrubbed 2026-06-26 (was wrongly tracked 23 commits).
|
||||
STATE.md
|
||||
# STATE.md — current state-of-play. TRACKED continuity per ~/.claude/CLAUDE.md
|
||||
# (overrides the polyrepo gitignore convention); pushed to private Forgejo, never public.
|
||||
|
||||
# Local planning docs (briefs, design notes, observations) — never committed.
|
||||
# Existing tracked files in docs/ predate this rule; new planning docs stay local.
|
||||
|
|
@ -32,7 +30,7 @@ docs/ultracontinue-design-notes.md
|
|||
# Ultraplan project directories — briefs, research, plans, progress all local.
|
||||
.claude/projects/
|
||||
|
||||
# --- session/local state (gitignored) — STATE.md is LOCAL-ONLY (open/ = public), se ~/.claude/CLAUDE.md ---
|
||||
# --- session/local state (gitignored) — STATE.md er nå tracked, se ~/.claude/CLAUDE.md ---
|
||||
REMEMBER.md
|
||||
ROADMAP.md
|
||||
TODO.md
|
||||
|
|
|
|||
77
CHANGELOG.md
77
CHANGELOG.md
|
|
@ -4,83 +4,6 @@ All notable changes to this project will be documented in this file.
|
|||
|
||||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).
|
||||
|
||||
## v5.9.1 — 2026-07-03 — Fix /trekendsession load-time crash (eager-exec placeholders)
|
||||
|
||||
Patch, no functional additions.
|
||||
|
||||
### Fixed
|
||||
|
||||
- `/trekendsession` was unusable in every invocation: two of its three `` !`...` `` eager-exec blocks (Phase 3 atomic-write, Phase 4 validator call) contained unresolved runtime placeholders (`<project-dir>` etc.). The harness executes eager-exec blocks at command LOAD time, so zsh parsed `<project-dir>` as input redirection and the command aborted before the model saw a single instruction. Both blocks are now plain runtime Bash fences with the `{curly}` placeholder convention (shell-inert), matching `trekplan.md`/`trekresearch.md`. The Phase 1 discovery block (self-contained) keeps its legitimate eager-exec prefix; `trekcontinue.md`'s discovery block was runtime-verified unaffected.
|
||||
- Latent secondary bug in the same blocks: cwd-relative plugin paths (`lib/validators/...`, `./lib/util/atomic-write.mjs`) would have failed with `ERR_MODULE_NOT_FOUND` even after substitution, since the Bash cwd is the user's repo. Both now use absolute `${CLAUDE_PLUGIN_ROOT}` paths per the existing command convention (Node ESM accepts absolute-path import specifiers — verified on Node 18+).
|
||||
|
||||
### Added
|
||||
|
||||
- Regression guard `tests/commands/trekendsession.test.mjs`: scans every `` !` ``-block in `commands/*.md` for unresolved `<angle>`/`{curly}` placeholders (this bug class is silent until first invocation), plus structure tests pinning Phase 3/4 as runtime Bash with `${CLAUDE_PLUGIN_ROOT}` paths and exactly one surviving eager block. Suite baseline 828 → 832 (830 pass / 0 fail / 2 skip).
|
||||
|
||||
## v5.9.0 — 2026-07-02 — Fable model tier + deep-research engine
|
||||
|
||||
Additive, plus one behavior alignment: profile `phase_models` now reach sub-agent spawn sites (previously documented but never wired), and the seven command orchestrators no longer pin `model: opus` — frontmatter omits `model:`, so the orchestrator follows the session model.
|
||||
|
||||
### Fable model tier
|
||||
|
||||
- `fable` (→ Claude Fable 5, Mythos-class, positioned above Opus) is an accepted model value throughout the validation chain: `BASE_ALLOWED_MODELS` widened to `['sonnet', 'opus', 'fable']` in `lib/validators/profile-validator.mjs` — the single source imported by brief-validator and phase-signal-resolver (two-layer gate preserved; accept-fable AND reject-unknown-model covered at both layers). No env gate — haiku's `VOYAGE_ALLOW_HAIKU` opt-in stays as-is.
|
||||
- `/trekbrief` Phase 3.5 tier loop offers a 4th option: `fable → {effort: high, model: fable}`. AskUserQuestion's 4-option maximum is now fully used — a 5th tier requires a loop redesign. The fable tier reuses `effort: high` orchestration semantics; `EFFORT_LEVELS` is unchanged.
|
||||
- New built-in profile `lib/profiles/fable.yaml` (all six phases on `fable`, modeled on premium; registered in `BUILTIN_NAMES` with a `loadProfile('fable')` canary test so a registry regression fails loudly instead of silently resolving premium). Premium stays the default.
|
||||
- Reasoning effort is inherited from the session: Fable 5's default effort is `high`, NOT xhigh, and switching model resets effort — set xhigh at session level (`/effort xhigh`, the `effortLevel` setting, or `CLAUDE_CODE_EFFORT_LEVEL`). Documented canonically in `docs/profiles.md` §Model & effort axes.
|
||||
- `claude-fable-5` added to the cost `PRICE_TABLE` ($10/MTok input, $50/MTok output; cache write 5m $12.50 / 1h $20; cache read $1 — verified 2026-07-02 against the official platform pricing docs). `PRICE_TABLE_VERSION` bumped to `2026-07-02`. Without the entry, every fable run would report `cost_usd: null` in the observability export.
|
||||
- Profile tables + allowlist prose updated across README, `docs/profiles.md`, `docs/operations.md`, `docs/HANDOVER-CONTRACTS.md`, `docs/architecture.md`, `docs/command-modes.md`, templates, and CLAUDE.md; the S15 doc pins now machine-check the fable row cell-for-cell against `fable.yaml`. The `^(opus|sonnet)…` regex claim in two docs was corrected — validation is an exact string match against `BASE_ALLOWED_MODELS`; the regex never existed in code.
|
||||
|
||||
### Behavior alignment: profile `phase_models` now reach sub-agent spawns
|
||||
|
||||
- Pre-existing wiring gap (found in exploration): all four pipeline commands invoked only the brief-only `phase-signal-resolver.mjs`, so the `?? profile.phase_models[<phase>]` half of the documented composition rule never executed — `--profile <x>` never reached sub-agent spawns.
|
||||
- Fixed with a single composed resolver: `resolver.mjs --resolve-phase-model` now returns `{effort, model, source}` (brief > profile > default, with effort passed through atomically) and is the one CLI the four pipeline commands invoke. A doc-consistency pin requires the composed invocation and forbids the brief-only CLI in command Bash blocks.
|
||||
- **Behavior change (contract alignment):** `--profile economy/balanced` now genuinely reaches sub-agent spawn sites for the first time — behavior aligns with what the docs have long claimed. The premium default is unaffected in practice (premium resolves `opus`, which equals the frontmatter fallback).
|
||||
- Command frontmatter: the `model: opus` line is DELETED from all seven commands — omission (not the disputed `inherit` literal) is the spelling both official surfaces document as session-inheritance, guarded by a frontmatter-absence doc pin. Accepted tradeoff: in a sonnet session the orchestrator runs on sonnet; re-add a frontmatter pin for deterministic orchestrator choice. The 24 `agents/*.md` `model: opus` pins are untouched (spawn-time injection wins; frontmatter is the fallback). `/trekcontinue`/`/trekendsession` spawn no exploration swarm and get no spawn-site injection; the continue phase is covered at resolver level and follows the session model.
|
||||
|
||||
### Bundled unreleased work (since v5.8.0)
|
||||
|
||||
- `/trekresearch --engine {swarm|deep-research}` (`581489a..9374820`): opt-in delegation of the external research phase to Claude Code's built-in `/deep-research` workflow, with in-context adapter + self-check, availability fallback to swarm (never hard-fails), and doc-consistency pins across surfaces.
|
||||
- brief-validator CLI no-flag invocation fix (`926b768`).
|
||||
- Deep-research engine research notes + docs (`60e9e7a`, `9d8e043`).
|
||||
|
||||
### Operator + consume-side notes
|
||||
|
||||
- The operator-global CLAUDE.md policy "Opus 4.8 default for all subagents" predates the fable tier; updating it is an operator action outside this repo.
|
||||
- `/plugin update` compares against a stale local marketplace clone and can report "already at the latest version" after this release (Claude Code issues #35752 / #38271, both closed-not-planned). Reliable refresh: remove + re-add the marketplace, or `git pull --ff-only` in the marketplace clone. Cross-version skew consequence: a stale cached v5.8 brief-validator REJECTS fable-bearing briefs with `BRIEF_INVALID_MODEL` — enum widening is safe for new readers of old data, not old readers of new data.
|
||||
- Org `availableModels` with `enforceAvailableModels: true` can make an inheriting orchestrator silently fall back to the first allowed model.
|
||||
|
||||
### Release hygiene
|
||||
|
||||
- Suite (measured with bare `npm test` at release): **828 (826 pass / 0 fail / 2 skipped)** — +17 over the pre-release ground-truth baseline of 811 measured 2026-07-02 (allowlist/gate coverage, fable profile pins, composed-resolver + frontmatter-absence doc pins, PRICE_TABLE case).
|
||||
- Version sync: `plugin.json`, `package.json`, `package-lock.json`, README badge, CHANGELOG top entry all at `5.9.0`, guarded by `doc-consistency.test.mjs`.
|
||||
|
||||
## v5.8.0 — 2026-06-30 — offline gold-scored output eval (SKAL-1·4b)
|
||||
|
||||
Additive — no behavior change, no breaking change. Internal eval infrastructure only (`lib/` + `tests/` + docs); no command, agent, profile, or Handover contract touched.
|
||||
|
||||
### Gold-scored output eval (SKAL-1·4b)
|
||||
|
||||
- The review-coordinator self-eval gains its **scoring run**, building on the deterministic coordinator contract + golden corpus shipped as the 4a foundation in v5.7.0.
|
||||
- `lib/review/gold-scorer.mjs`: `scoreFindings(runFindings, goldFindings)` matches at **`(file, rule_key)` granularity** (line + severity deliberately ignored) → `{ tp, fp, fn, precision, recall, f1, matched, missed, spurious }`; `scoreVerdict` checks exact verdict match. Pure — no I/O, no LLM, no network. Vacuous-set conventions (empty run → recall 0; empty gold → precision 0; f1 collapses to 0) are documented in the module header.
|
||||
- `tests/fixtures/bakeoff-rich/runs/run-perfect.json`: the first **committed run** — reviewer payloads that, fed through `runContract`, reproduce all 5 seeded gold findings. The regression guard: any future contract change that silently suppresses or skips a seeded finding breaks the eval.
|
||||
- `tests/lib/gold-eval.test.mjs`: the **scoring run**, wired into `node --test`. Asserts `run-perfect` scores precision/recall/f1 = 1.0 and `verdict == expected_verdict` (BLOCK). Offline: committed payloads, no live agent spawn (the LLM-in-the-loop grading is the separate 4c tier).
|
||||
- **Third test-census category.** `lib/util/test-census.mjs` now reports a `goldEval` bucket (matched by `GOLD_EVAL_FILE_RE`) separately from `behavior` and `docPins` — a scoring run is neither behavior coverage nor a prose pin, so the honest-count invariant is a 3-way sum.
|
||||
- `docs/eval-corpus/README.md`: SKAL-1·4b moved from "Future hardening" to an implemented section documenting the scorer, run format, and census category.
|
||||
- Suite: **822 (820 pass / 0 fail / 2 skip)**, up from 809 (+10 scorer behavior tests, +3 eval scoring-run tests). The scorer test covers the discriminating paths (false negatives + false positives) and the degenerate empty-run / empty-gold cases, not merely the all-match case.
|
||||
- Version sync: `plugin.json`, `package.json`, `package-lock.json`, README badge, CHANGELOG top entry all at `5.8.0`, guarded by `doc-consistency.test.mjs`.
|
||||
|
||||
## v5.7.1 — 2026-06-29 — relocate agent `<example>` blocks to body (always-loaded token trim)
|
||||
|
||||
Performance/packaging change — no behavior change, no breaking change. (M4)
|
||||
|
||||
### Always-loaded token trim
|
||||
|
||||
- The 17 example-bearing agents each carried two `<example>` blocks (34 total) inside their `description:` frontmatter. voyage launches its agents **by name** from the orchestrator commands, so those auto-selection examples were paying a per-turn token tax for a path the pipeline never uses. They are now relocated **verbatim** into each agent's body under a `## When to use — examples` section — preserved, not deleted.
|
||||
- Frontmatter `description` chars across all 24 agents: **17,672 → 4,945** — roughly **3,180 fewer always-loaded tokens per turn** for every session with voyage enabled. The 7 zero-example agents are untouched; no system prompt, tools, model, or `name` changed.
|
||||
- New invariant test in `tests/lib/agent-frontmatter.test.mjs`: no `<example>` in any frontmatter `description`, and ≥ 34 `<example>` retained across agent bodies (relocation moves, never deletes). Suite: **807 pass / 0 fail / 2 skip**.
|
||||
- **The saving only reaches a machine after `/plugin marketplace update` + reload** — the delta is not instant on the machine that ships the release.
|
||||
- Version sync: `plugin.json`, `package.json`, `package-lock.json`, README badge, CHANGELOG top entry all at `5.7.1`, guarded by `doc-consistency.test.mjs`.
|
||||
|
||||
## v5.7.0 — 2026-06-26 — opt-in token/cost metering (SKAL-2) + eval foundation (SKAL-1·4a)
|
||||
|
||||
Additive — no breaking change. Two unreleased work-streams land together.
|
||||
|
|
|
|||
30
CLAUDE.md
30
CLAUDE.md
|
|
@ -2,20 +2,20 @@
|
|||
|
||||
Voyage — a contract-driven Claude Code pipeline: brief, research, plan, execute, review, continue. Deep implementation planning and research with specialized agent swarms, external research, adversarial review, session decomposition, disciplined execution, and headless support.
|
||||
|
||||
**Design principle: Context Engineering** — build the right context by orchestrating specialized agents. Each step in the pipeline (brief → research → plan → execute) produces a structured artifact that the next step consumes. The load-bearing benefit is the parallel wall-clock + structured artifact handoffs; main-context relief is asserted-by-design, not measured (Δ ≈ 0 in the one PoC — see `docs/T1-synthesis-poc-results.md`).
|
||||
**Design principle: Context Engineering** — build the right context by orchestrating specialized agents. Each step in the pipeline (brief → research → plan → execute) produces a structured artifact that the next step consumes. (The claim that fanning out to a swarm *relieves* the main context is asserted-by-design, not measured end-to-end: the one delegation actually measured — the Phase-7 synthesis PoC — found Δ main-context ≈ 0, see `docs/T1-synthesis-poc-results.md`. The parallel wall-clock + structured artifact handoffs are the load-bearing benefit; main-context relief is not yet demonstrated.)
|
||||
|
||||
> **Architecture slot.** The plan command auto-discovers `architecture/overview.md` if present — any compatible producer plugs in (the architect plugin is no longer publicly distributed). Migration history (v3.0.0 extraction) → [CHANGELOG.md](CHANGELOG.md).
|
||||
> **v3.0.0 — architect step extracted from this plugin.** The plan command still auto-discovers `architecture/overview.md` if present, so any compatible producer (architect plugin no longer publicly distributed; the architecture/overview.md slot remains available for any compatible producer) plugs into the same slot. See [CHANGELOG.md](CHANGELOG.md) for migration history.
|
||||
|
||||
> **Trinity context (informational).** Voyage is Tier 1 (per-task) of a three-tier architecture. **Asymmetry is a hard invariant:** Voyage stays unaware of Tier 2/3; Handover 1 (brief format) is the only integration point, no producer is privileged, and brief-schema changes are breaking for downstream consumers (formalized as a public contract in v5.5.0). Tier 2/3 producer detail + the public contract → `docs/HANDOVER-CONTRACTS.md` §Handover 1 (PUBLIC CONTRACT).
|
||||
> **Trinity context (2026-05-13, informational).** Voyage is Tier 1 (per-task) of a three-tier architecture in active design under the author's private marketplace: Tier 2 `app-creator` (per-app — "what does the app need, what's the next brief?") produces briefs Voyage consumes; Tier 3 `app-factory` (per-portfolio — "which app needs me now?") aggregates state across multiple app-creator instances. Both are pre-implementation and will ship to Forgejo when ready. **Asymmetry is a hard invariant:** Voyage stays unaware of Tier 2/3. Handover 1 (brief format) is the only integration point — any compatible producer can feed Voyage, app-creator is not privileged. Brief-schema changes are therefore breaking changes for downstream consumers, formalized as a public contract in v5.5.0 — see `docs/HANDOVER-CONTRACTS.md` §Handover 1 (PUBLIC CONTRACT).
|
||||
|
||||
> **Cross-cutting invariant: brief framing must match operator intent.** The brief is the pipeline's source of truth; operator intent lives in memory files — the pipeline must not polish a wrong premise. Enforced as the `brief_version 2.2` gate (v5.5), all BLOCKER for briefs declaring ≥ 2.2: (1) explicit `framing: preserve|refine|replace|new-direction` frontmatter, `AskUserQuestion`-validated in `/trekbrief` Phase 2.5; (2) memory-alignment as `brief-reviewer` dimension 6; (3) mandatory `## TL;DR`. Existing 2.0/2.1 briefs stay valid; `trekreview` briefs are exempt. Full implementation + contract evolution → `docs/HANDOVER-CONTRACTS.md` §Handover 1 (PUBLIC CONTRACT).
|
||||
> **Cross-cutting invariant: brief framing must match operator intent (2026-05-15).** Etablert etter residiv. Briefen er pipelinens source of truth; operatørens intent lever i hodet + i memory-filer (`feedback_*`, `project_*`); pipelinen tvinger ikke alignment. Høyere reasoning-kraft polerer feil premiss istedenfor å utfordre det. **Tre lag av forsvar (input-siden), implementert i S6 som `brief_version 2.2`-gate (v5.5), alle BLOCKER ved brudd for briefer som deklarerer ≥ 2.2:** (1) eksplisitt `framing: preserve|refine|replace|new-direction` i brief-frontmatter, `AskUserQuestion`-validert i `/trekbrief` Phase 2.5 før brief-prosa skrives (ikke-skippbar, også i `--quick`); enum-feil → `BRIEF_INVALID_FRAMING` (alle versjoner), fravær ved ≥ 2.2 → `BRIEF_MISSING_FRAMING`; (2) memory-alignment som dimensjon 6 i `brief-reviewer` — sammenlikner brief-prosa + `framing` mot relevante memory-filer, rapporterer kun eksplisitte motsigelser (degraderer til score 5 N/A uten memory-kontekst), wired til Phase 4e-gate (`memory_alignment.score ≥ 4`); (3) obligatorisk `## TL;DR`-seksjon (≤ 5 linjer, soft-cap → `BRIEF_TLDR_TOO_LONG`) øverst i `brief.md`. Eksisterende 2.0/2.1-briefer forblir gyldige (forward+backward-compat, speiler `phase_signals ≥ 2.1`-presedensen); `trekreview`-briefer er unntatt. **Schema-akse bumpet (2.1→2.2); plugin-versjon-badge + CHANGELOG bumpes ved den koordinerte releasen (S10).** Kontrakt-evolusjon dokumentert i `docs/HANDOVER-CONTRACTS.md` §Handover 1. Den gamle manuelle stopgap-sjekken er dermed retired for ≥ 2.2-briefer.
|
||||
|
||||
## Commands
|
||||
|
||||
| Command | Description | Model |
|
||||
|---------|-------------|-------|
|
||||
| `/trekbrief` | Brief — interactive interview produces a task brief with explicit research plan; optionally orchestrates the pipeline | opus |
|
||||
| `/trekresearch` | Research — deep local + external research, produces structured research brief. Opt-in `--engine {swarm\|deep-research}` delegates the external phase to Claude Code's built-in `/deep-research` workflow (swarm default) | opus |
|
||||
| `/trekresearch` | Research — deep local + external research, produces structured research brief | opus |
|
||||
| `/trekplan` | Plan — brief-reviewer, explore, plan, review. Requires `--brief` or `--project`. Auto-discovers `architecture/overview.md` if present | opus |
|
||||
| `/trekexecute` | Execute — disciplined plan/session-spec executor with failure recovery | opus |
|
||||
| `/trekreview` | Review — independent post-hoc review of delivered code against the brief. Produces `review.md` with severity-tagged findings (Handover 6) | opus |
|
||||
|
|
@ -24,20 +24,6 @@ Voyage — a contract-driven Claude Code pipeline: brief, research, plan, execut
|
|||
|
||||
Full flag reference for each command (modes, `--gates`, `--profile`, breaking changes): see `docs/command-modes.md`.
|
||||
|
||||
> **STORM bounded loop — default-off, env-gated.** `/trekresearch` Phase 4.5
|
||||
> (dimension discovery, under the existing `maxDimensions: 8` ceiling) and
|
||||
> Phase 5 (bounded multi-turn follow-up) run only at `effort: high` **and** only
|
||||
> when `VOYAGE_STORM_ENABLED=1`; unset, both are inert — Phase 5 because
|
||||
> `lib/util/research-loop-cap.mjs` grants a budget of 0, Phase 4.5 because its
|
||||
> skip-guard reads the flag directly (it never calls the cap). `TREKRESEARCH_MAX_CONV_TURNS` (default `3`,
|
||||
> invalid values fall back to `3`) sets turns per dimension; the budget is that
|
||||
> × `maxDimensions`. `VOYAGE_DISABLE_CAP_HOOK=1` switches off
|
||||
> `hooks/scripts/pre-agent-cap.mjs`, the `PreToolUse` gate that enforces the
|
||||
> budget in the harness rather than trusting prose. The cap counts turns from
|
||||
> its own append-only ledger. Adoption as a default is gated on the
|
||||
> pre-registered measurement in `docs/storm-measurement.md` — decline is a
|
||||
> no-op, adopt is one constant.
|
||||
|
||||
## Agents
|
||||
|
||||
| Agent | Model | Role |
|
||||
|
|
@ -67,15 +53,15 @@ Full flag reference for each command (modes, `--gates`, `--profile`, breaking ch
|
|||
| contrarian-researcher | opus | Counter-evidence, overlooked alternatives |
|
||||
| gemini-bridge | opus | Gemini Deep Research second opinion (conditional) |
|
||||
|
||||
> **Inventory (S33 reconcile).** 24 agent files = **21 spawnable** (one, `synthesis-agent`, ships **dormant** — Δ≈0, wired to nothing) **+ 3 orchestrator reference docs** (`planning-/research-/review-orchestrator` document the inline `/trek*` workflow, not spawnable capabilities). All 24 stay `model: opus` (operator pin `40d8742`); the glue/mechanical/retrieval/dormant roles were reconsidered for a sonnet downgrade and **kept opus** — decision record: `docs/voyage-vs-cc-balance-analysis.md` §10.
|
||||
> **Inventory (S33 reconcile).** 24 agent files = **21 spawnable** (one of which, `synthesis-agent`, ships **dormant** — Δ≈0, wired to nothing, see `docs/T1-synthesis-poc-results.md`) **+ 3 orchestrator reference docs** (`planning-/research-/review-orchestrator` document the inline `/trek*` workflow; they are reference documentation, not spawnable capabilities). All 24 stay `model: opus` (operator pin `40d8742`); the glue/mechanical/retrieval/dormant roles (V09 gemini-bridge, V16 session-decomposer, V11 retrieval agents, V08 researchers, V35 synthesis) were reconsidered for a sonnet downgrade and **kept opus** — document-only, no frontmatter change (the decision record lives in `docs/voyage-vs-cc-balance-analysis.md` §10).
|
||||
|
||||
> **Model & effort.** `opus` = Opus 4.8 (default reasoning effort `high`); `sonnet` = Sonnet 4.6; `fable` = Fable 5 (Mythos-class, above Opus — reasoning effort inherits from the session; xhigh requires a session-level setting). Select agents carry native per-spawn `effort:` (retrieval → `medium`, adversarial-reasoning → `high`) — a different axis from brief `phase_signals.effort` (orchestration shape: which agents/passes run). Per-agent table + axes → `docs/profiles.md` §Model & effort axes.
|
||||
> **Model & effort (CC 2.1.154+).** `opus` resolves to **Opus 4.8** (default reasoning effort `high`); `sonnet` to Sonnet 4.6. Select agents carry native `effort:` — retrieval agents (`task-finder`, `git-historian`, `dependency-tracer`, `architecture-mapper`) at `medium`, adversarial-reasoning agents (`plan-critic`, `risk-assessor`, `contrarian-researcher`, `review-coordinator`) at `high`. This native per-spawn **reasoning** effort is a different axis from brief `phase_signals.effort` (orchestration shape — which agents/passes run). See `docs/profiles.md` §Model & effort axes.
|
||||
|
||||
## Reference docs (read on demand)
|
||||
|
||||
- **Architecture, workflows, project-directory contract, state, terminology:** `docs/architecture.md`
|
||||
- **Quality infrastructure (`lib/` validators, parsers, autonomy primitives, hooks):** `docs/architecture.md` §Quality infrastructure
|
||||
- **Autonomy gates (`--gates`), Path A/B/C decision:** `docs/operations.md`
|
||||
- **Profile system (`--profile economy/balanced/premium/fable`), lookup order, custom profiles:** `docs/operations.md`
|
||||
- **Profile system (`--profile economy/balanced/premium`), lookup order, custom profiles:** `docs/operations.md`
|
||||
- **Observability (Stop hook, OTLP/textfile export, SSRF mitigation):** `docs/operations.md`
|
||||
- **Handover contracts (the 7 pipeline handovers):** `docs/HANDOVER-CONTRACTS.md`
|
||||
|
|
|
|||
131
GOVERNANCE.md
Normal file
131
GOVERNANCE.md
Normal file
|
|
@ -0,0 +1,131 @@
|
|||
# Governance
|
||||
|
||||
How this marketplace is maintained, what you can expect from upstream, and how it's meant to be used.
|
||||
|
||||
## TL;DR
|
||||
|
||||
- Solo-maintained, AI-assisted development, MIT licensed.
|
||||
- **Fork-and-own is the default model.** Upstream is a starting point, not a vendor.
|
||||
- Issues welcome as signals. Pull requests are not accepted — see [Why no PRs](#pull-requests--no).
|
||||
- No SLA. Best-effort bug fixes and security advisories. Breaking changes happen and are noted in each plugin's CHANGELOG.
|
||||
|
||||
---
|
||||
|
||||
## Can I trust this?
|
||||
|
||||
Be honest with yourself about what you're adopting:
|
||||
|
||||
- **One maintainer.** If I get hit by a bus, the bus wins. The repos stay up under MIT, but no one owes you a fix.
|
||||
- **AI-generated code with human review.** Every plugin is built through dialog-driven development with Claude Code. I read, test, and judge the output before it ships, but I'm not auditing every line the way a security firm would. Treat it accordingly.
|
||||
- **No commercial interests.** I'm not selling a SaaS, not steering you toward a paid tier, not collecting telemetry. The plugins run locally in your Claude Code installation.
|
||||
- **MIT licensed.** Fork it, modify it, ship it under your own name.
|
||||
|
||||
If you work somewhere that needs vendor accountability, support contracts, or signed assurances — **this isn't that.** Use it as a reference implementation, fork it into your own organization, and own the result.
|
||||
|
||||
---
|
||||
|
||||
## How this is meant to be used
|
||||
|
||||
### Fork-and-own
|
||||
|
||||
The intended workflow:
|
||||
|
||||
1. **Fork** the marketplace (or a single plugin) into your own organization or namespace.
|
||||
2. **Tailor** it to your context — terminology, integrations, cycle lengths, regulatory framing, whatever doesn't fit out of the box.
|
||||
3. **Maintain it yourself.** Treat your fork as the canonical version for your team.
|
||||
4. **Watch upstream selectively.** Cherry-pick changes that help, ignore changes that don't. There's no obligation to stay in sync.
|
||||
|
||||
This isn't a workaround for not accepting PRs. It's the actual recommended adoption pattern, especially for plugins like `okr` and `ms-ai-architect` where every Norwegian public sector organization will need its own tildelingsbrev mappings, terminology, and integrations. A central "one true plugin" would be wrong for everyone.
|
||||
|
||||
### What to change first when you fork
|
||||
|
||||
Each plugin differs, but the common edits are:
|
||||
|
||||
- **Identity** — rename the plugin, replace authorship, update README.
|
||||
- **External integrations** — issue trackers, knowledge bases, dashboards, observability backends. The plugins ship as starting points, not pre-wired. Every organization must configure its own integrations.
|
||||
- **Norwegian-specific framing** — relevant for `okr` and `ms-ai-architect`. Other plugins are jurisdiction-neutral. Rewrite for your jurisdiction if you're outside Norway.
|
||||
- **Reference docs** — the knowledge base in each plugin reflects my reading. Replace with your organization's authoritative sources.
|
||||
- **Hooks and policies** — security thresholds, blocked commands, and audit gates are tuned to my taste. Tune them to yours.
|
||||
|
||||
### Staying current with upstream
|
||||
|
||||
If you want to pull in upstream changes later:
|
||||
|
||||
- **Cherry-pick, don't merge.** Each plugin moves independently and breaking changes land without ceremony.
|
||||
- **Read the CHANGELOG first.** Every plugin has one.
|
||||
- **Keep your customizations in clearly-named files.** The harder upstream is to merge cleanly, the more painful staying current becomes. A `local/` directory or `*.local.md` convention helps.
|
||||
|
||||
---
|
||||
|
||||
## What upstream provides
|
||||
|
||||
| | What I do | What I don't |
|
||||
|---|---|---|
|
||||
| **Bug fixes** | Best-effort when I notice or get a clear report | No SLA, no triage commitment |
|
||||
| **Security issues** | Investigate within reasonable time, document in CHANGELOG | No CVE process, no embargo coordination |
|
||||
| **New features** | When they fit my own usage | Not on request |
|
||||
| **Norwegian public sector context** | Kept current as long as the project lives | If I lose interest or change jobs, the framing freezes |
|
||||
| **Breaking changes** | Documented in CHANGELOG | They happen — version pin if you need stability |
|
||||
| **Compatibility** | Tracked against current Claude Code releases | No long-term support branches |
|
||||
|
||||
If any of this is a dealbreaker — fork now, version-pin, and stop reading upstream.
|
||||
|
||||
---
|
||||
|
||||
## How to contribute
|
||||
|
||||
### Issues — yes, please
|
||||
|
||||
Issues are the most valuable thing you can send me:
|
||||
|
||||
- **Bug reports** with reproduction steps. Even a screenshot helps.
|
||||
- **Use-case feedback.** "I tried to use this in my organization and X didn't fit" is genuinely useful, even if I can't fix it for you.
|
||||
- **Pointers to better sources.** If you know a DFØ veileder, an NSM guideline, or an academic paper that contradicts what's in a knowledge base, tell me.
|
||||
- **Security findings.** See each plugin's `SECURITY.md` for disclosure preference where one exists; otherwise email rather than open a public issue.
|
||||
|
||||
### Pull requests — no
|
||||
|
||||
This is deliberate, not laziness:
|
||||
|
||||
- **Solo review is a bottleneck.** Honest PR review takes me longer than rewriting from scratch. The math doesn't work.
|
||||
- **Forks are where the value is.** The fork-and-own model means upstream consolidation isn't the point. Your organization's adaptations belong in your fork, not mine.
|
||||
- **AI-generated code complicates provenance.** Every line here is produced through dialog with Claude Code, with me as the judge. Mixing in PRs from contributors with different processes and licensing assumptions creates a mess I'd rather not untangle.
|
||||
|
||||
If you've built something useful on top of a fork, **publish it under your own name and link back.** I'll happily list notable forks here once they exist.
|
||||
|
||||
### Notable forks
|
||||
|
||||
*(To be populated as forks emerge. If you've forked one of these plugins for production use, open an issue and I'll add a link.)*
|
||||
|
||||
---
|
||||
|
||||
## Relationship between plugins
|
||||
|
||||
These plugins are **independent**. Install one without the others, fork one without the others. They share conventions (slash command naming, hook patterns, AI-generated disclosure) but no runtime dependencies.
|
||||
|
||||
The marketplace is a **catalog**, not a suite. Don't fork the whole repo unless you actually want to maintain everything.
|
||||
|
||||
---
|
||||
|
||||
## Versioning and stability
|
||||
|
||||
- **Semantic versioning per plugin.** Each plugin has its own `CHANGELOG.md` and version number.
|
||||
- **Breaking changes happen.** I bump the major version when they do, but I don't run an LTS branch.
|
||||
- **Pin your version.** If stability matters more than features, install a specific version and stay there until you choose to upgrade.
|
||||
|
||||
---
|
||||
|
||||
## Public sector adoption notes
|
||||
|
||||
For Norwegian etater specifically:
|
||||
|
||||
- **DPIA-relevant data flows are documented in the relevant plugin README where applicable.** Read them before installation.
|
||||
- **No data leaves your machine** beyond what Claude Code itself sends to Anthropic. The plugins themselves do not call external services unless you configure an integration.
|
||||
- **Drøftingsplikt and ledelsesansvar** are not replaced by these tools. The `okr` plugin coaches; it does not decide. The `ms-ai-architect` plugin advises; it does not approve.
|
||||
- **Choose your Claude deployment carefully.** claude.ai vs. API direct vs. Bedrock in EU region have different data residency profiles. The plugins don't choose for you.
|
||||
|
||||
---
|
||||
|
||||
## License
|
||||
|
||||
MIT for all plugins in this marketplace. See each plugin's `LICENSE` file.
|
||||
|
|
@ -6,7 +6,7 @@ v3.x → v4.0.0 is a rebrand. All command names changed:
|
|||
`/ultrareview-local` → `/trekreview`, `/ultracontinue-local` → `/trekcontinue`,
|
||||
`/ultraplan-end-session-local` → `/trekendsession`. The plugin is now
|
||||
named `voyage`. Re-fork from main if upgrading. There is no migration
|
||||
path — see [`GOVERNANCE.md`](https://git.fromaitochitta.com/open/repo-standard/src/branch/main/GOVERNANCE.md) for the fork-and-own model.
|
||||
path — see `GOVERNANCE.md` for the fork-and-own model.
|
||||
|
||||
Prior version migration notes (v1→v2, v2→v3) are preserved in
|
||||
`CHANGELOG.md` only.
|
||||
|
|
|
|||
99
README.md
99
README.md
|
|
@ -1,17 +1,17 @@
|
|||
# voyage
|
||||
# trekplan — Brief, Research, Plan, Execute, Review, Continue
|
||||
|
||||
Contract-driven Claude Code pipeline: brief, research, plan, execute, review. Agent swarms, research triangulation, adversarial review, multi-session resumption.
|
||||
|
||||

|
||||

|
||||

|
||||

|
||||
|
||||
> **Solo-maintained, fork-and-own.** This plugin is a starting point, not a vendor product. Issues are welcome as signals; pull requests are not accepted. See [GOVERNANCE.md](https://git.fromaitochitta.com/open/repo-standard/src/branch/main/GOVERNANCE.md) for the full model and what upstream provides.
|
||||
> **Solo-maintained, fork-and-own.** This plugin is a starting point, not a vendor product. Issues are welcome as signals; pull requests are not accepted. See [GOVERNANCE.md](GOVERNANCE.md) for the full model and what upstream provides.
|
||||
|
||||
*AI-generated: all code produced by Claude Code through dialog-driven development. Every change is human-directed, reviewed, and validated before commit. Per Anthropic Consumer Terms §4, ownership of outputs is assigned to the user; this plugin is licensed MIT.*
|
||||
*AI-generated: all code produced by Claude Code through dialog-driven development. [Full disclosure →](../../README.md#ai-generated-code-disclosure)*
|
||||
|
||||
A [Claude Code](https://docs.anthropic.com/en/docs/claude-code) plugin for deep implementation planning, multi-source research, autonomous execution, independent post-hoc review, and zero-friction multi-session resumption. Six commands, one pipeline:
|
||||
|
||||
> **What's new — v5.6.1: leaner always-loaded agent listing.** The four reference/dormant agents (the three `*-orchestrator` reference docs + the dormant `synthesis-agent`) now carry one-line `description:` frontmatter — their full rationale already lives in each file's body — trimming ~700 tokens from the agent listing injected into every session, with no behavior change. **v5.6.0: `/trekexecute` loop hardening.** Execution-loop termination is now machine-verifiable (completion gate, Hard Rule 18) and recovery is bounded by an explicit `TREKEXECUTE_MAX_RECOVERY_ITERATIONS` budget (default 25, Hard Rule 20). The prior coordinated release **v5.5.0** added brief **framing** enforcement (`brief_version 2.2`) + a `/trekreview` reviewer-schema contract. Additive — no breaking changes. **Full version history → [CHANGELOG.md](CHANGELOG.md).**
|
||||
|
||||
| Command | What it does |
|
||||
|---------|-------------|
|
||||
| **`/trekbrief`** | Brief — interactive interview produces a task brief with explicit research plan |
|
||||
|
|
@ -21,23 +21,6 @@ A [Claude Code](https://docs.anthropic.com/en/docs/claude-code) plugin for deep
|
|||
| **`/trekreview`** | Review — independent post-hoc review of delivered code against the brief, severity-tagged findings |
|
||||
| **`/trekcontinue`** | Continue — read `.session-state.local.json` and resume the next session in a multi-session project |
|
||||
|
||||
## Install
|
||||
|
||||
```bash
|
||||
claude plugin marketplace add https://git.fromaitochitta.com/open/ktg-plugin-marketplace.git
|
||||
claude plugin install voyage@ktg-plugin-marketplace
|
||||
```
|
||||
|
||||
Or enable directly in `~/.claude/settings.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"enabledPlugins": {
|
||||
"voyage@ktg-plugin-marketplace": true
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`/trekbrief`, `/trekplan`, and `/trekreview` each end by running `scripts/annotate.mjs` against the just-written artifact and printing the resulting `file://<abs path>` link. The operator opens the HTML in a browser, clicks any line of the document, writes their own note in the inline textarea, watches a sidebar of all notes (editable, deletable, persisted in browser `localStorage`), and clicks "Copy Prompt" to get one structured prompt that they paste back into Claude — Claude then revises the `.md` from the notes. **The operator drives every annotation.** See [Reviewing and annotating artifacts](#reviewing-and-annotating-artifacts-v502).
|
||||
|
||||
Every artifact lives in one project directory: `.claude/projects/{YYYY-MM-DD}-{slug}/` contains `brief.md`, `research/NN-*.md`, `plan.md`, `sessions/`, `progress.json`, and `review.md`.
|
||||
|
|
@ -85,6 +68,9 @@ Under the hood, `lib/util/autonomy-gate.mjs` runs a small state machine (`idle
|
|||
## Quick start
|
||||
|
||||
```bash
|
||||
# Install the marketplace, then browse and enable plugins with /plugin
|
||||
claude plugin marketplace add https://git.fromaitochitta.com/open/ktg-plugin-marketplace.git
|
||||
|
||||
# Capture intent (interactive)
|
||||
/trekbrief Add user authentication with JWT tokens
|
||||
# → .claude/projects/2026-04-18-jwt-auth/brief.md
|
||||
|
|
@ -141,10 +127,7 @@ Concrete capabilities, observable in the code — not aspirations.
|
|||
|
||||
**Virksomhet / regulated environment.** Defense-in-depth security across four layers (plugin hooks, prompt-level denylist, pre-execution plan scan, scoped tool access). `disableSkillShellExecution: true` recommendation for fork-ers handling untrusted briefs. No cloud dependency, no GitHub requirement. Validators are plain-Node CLIs — invocable from CI, custom hooks, or external tools, not just from voyage commands.
|
||||
|
||||
## Non-goals
|
||||
|
||||
What this pipeline does **not** solve — read this before adopting it:
|
||||
|
||||
**What it doesn't solve:**
|
||||
- LLM output truthfulness. Validators check shape, not facts. A plan with hallucinated paths passes schema but fails in execute. Plan-critic catches some, not all.
|
||||
- Multi-user concurrency on a single project directory. Two simultaneous executors will clobber `progress.json`.
|
||||
- Cost management. Opus on the orchestrator layer is expensive; documented in [Cost profile](#cost-profile), no automatic model downgrade.
|
||||
|
|
@ -168,7 +151,7 @@ Output: `.claude/projects/{YYYY-MM-DD}-{slug}/brief.md`
|
|||
|------|-------|----------|
|
||||
| **Default** | `/trekbrief <task>` | Dynamic interview until quality gates pass. No question cap. |
|
||||
| **Quick** | `/trekbrief --quick <task>` | Starts compact (optional sections get at most one probe), still escalates on weak required sections or failed review gate. |
|
||||
| **Profile** | `/trekbrief --profile <name> <task>` | (v4.1.0) Pin model profile for the brief phase: `economy` / `balanced` / `premium` / `fable` / `<custom>`. See [Profile system](#profile-system-v410) below. |
|
||||
| **Profile** | `/trekbrief --profile <name> <task>` | (v4.1.0) Pin model profile for the brief phase: `economy` / `balanced` / `premium` / `<custom>`. See [Profile system](#profile-system-v410) below. |
|
||||
|
||||
`/trekbrief` is **always interactive**. There is no foreground/background mode — the interview requires user input.
|
||||
|
||||
|
|
@ -212,28 +195,9 @@ Output:
|
|||
| **External** | `/trekresearch --external <question>` | Only external research agents (skip codebase analysis) |
|
||||
| **Foreground** | `/trekresearch --fg <question>` | No-op alias (foreground is default since v2.4.0) |
|
||||
| **Profile** | `/trekresearch --profile <name> <question>` | (v4.1.0) Pin model profile for the research phase. See [Profile system](#profile-system-v410). |
|
||||
| **Engine** | `/trekresearch --external --engine deep-research <question>` | Delegate the external phase to Claude Code's built-in `/deep-research` workflow; falls back to `swarm` if unavailable. Default `swarm`. |
|
||||
|
||||
Flags combine: `--project <dir> --external`.
|
||||
|
||||
#### Bounded conversation loop (default-off)
|
||||
|
||||
At `effort: high`, research can discover additional dimensions (Phase 4.5) and
|
||||
run a bounded multi-turn follow-up loop on under-illuminated ones (Phase 5).
|
||||
Both are **default-off** and env-gated rather than flag-gated, because enabling
|
||||
them costs turns:
|
||||
|
||||
| Env-var | Default | Behavior |
|
||||
|---------|---------|----------|
|
||||
| `VOYAGE_STORM_ENABLED` | _(unset — default-off)_ | `=1` gives the Phase 5 loop a non-zero turn budget and lets Phase 4.5 discovery run. Unset, the budget is 0 and **both** phases are inert. |
|
||||
| `TREKRESEARCH_MAX_CONV_TURNS` | `3` | Turns per under-illuminated dimension. Invalid values fall back to `3`, never to unbounded. |
|
||||
| `VOYAGE_DISABLE_CAP_HOOK` | _(unset)_ | `=1` disables the `PreToolUse` hook that enforces the turn budget. |
|
||||
|
||||
The cap counts turns itself from an append-only ledger — it never asks the loop
|
||||
how many turns it has used. Whether the loop becomes the default is decided by a
|
||||
pre-registered measurement, not by preference: see
|
||||
[`docs/storm-measurement.md`](docs/storm-measurement.md).
|
||||
|
||||
Research uses up to 5 local agents (architecture-mapper, dependency-tracer, task-finder, git-historian, convention-scanner) and 4 external agents (docs-researcher, community-researcher, security-researcher, contrarian-researcher) plus the optional Gemini bridge for an independent second opinion. Per-agent details in [`agents/`](agents/).
|
||||
|
||||
---
|
||||
|
|
@ -792,16 +756,35 @@ The `pre-compact-flush.mjs` hook directly fixes the documented P0 in `docs/treke
|
|||
|
||||
**Annotation HTML requires a desktop browser.** `scripts/annotate.mjs` produces a single self-contained `.html` file you open with `file://` in any modern browser (Chrome / Safari / Firefox / Edge — last two versions). No CDN, no server, no npm runtime deps. State persists in `localStorage` so closing and re-opening the tab keeps your work, but it's local to one browser on one machine — not synced anywhere. If you want to annotate without a browser, paste the `.md` into Claude with "comments inline below" and write notes in chat — same end result, just without the visual surface.
|
||||
|
||||
## Installation
|
||||
|
||||
Add the marketplace and browse plugins with `/plugin`:
|
||||
|
||||
```bash
|
||||
claude plugin marketplace add https://git.fromaitochitta.com/open/ktg-plugin-marketplace.git
|
||||
```
|
||||
|
||||
Or enable directly in `~/.claude/settings.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"enabledPlugins": {
|
||||
"voyage@ktg-plugin-marketplace": true
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
An optional architect step between research and plan was previously available via a separate plugin; that architect plugin is no longer publicly distributed. The `architecture/overview.md` filesystem slot remains supported by `/trekplan` for any compatible producer.
|
||||
|
||||
## Profile system (v4.1.0)
|
||||
|
||||
Four built-in model profiles plus operator-defined `<custom>.yaml` (drop in `lib/profiles/`). Each profile pins `phase_models` for the six pipeline phases. The active profile is recorded in plan.md frontmatter as `profile: <name>` and emitted to JSONL stats for cost-attribution.
|
||||
Three built-in model profiles plus operator-defined `<custom>.yaml` (drop in `lib/profiles/`). Each profile pins `phase_models` for the six pipeline phases. The active profile is recorded in plan.md frontmatter as `profile: <name>` and emitted to JSONL stats for cost-attribution.
|
||||
|
||||
| Profile | Brief | Research | Plan | Execute | Review | Continue | Use case |
|
||||
|---------|-------|----------|------|---------|--------|----------|----------|
|
||||
| `economy` | sonnet | sonnet | sonnet | sonnet | sonnet | sonnet | ⚠ **Experimental** (uncalibrated Jaccard floor) — lowest cost; high-confidence small-scope tasks (opt-in via `--profile economy`) |
|
||||
| `balanced` | sonnet | sonnet | opus | sonnet | opus | sonnet | Mixed — opus where reasoning depth pays off (opt-in via `--profile balanced`) |
|
||||
| `premium` (default) | opus | opus | opus | opus | opus | opus | Maximum quality — Opus on every phase (default since the 2026-05-13 operator decision) |
|
||||
| `fable` | fable | fable | fable | fable | fable | fable | Max quality — Fable 5 (Mythos-class, above Opus) on every phase (opt-in via `--profile fable`); reasoning effort inherits from the session |
|
||||
|
||||
Lookup order:
|
||||
|
||||
|
|
@ -826,9 +809,9 @@ Default JSONL stats stream (`${CLAUDE_PLUGIN_DATA}/trek*-stats.jsonl`) is unchan
|
|||
|
||||
## Cost profile
|
||||
|
||||
The default `premium` profile runs **Opus on every phase** of the pipeline's agent work — the exploration and review swarms (5–10 sub-agents per command; spawn sites inject the composed brief > profile > frontmatter resolution, with `agents/*.md` `model: opus` pins as the fallback) and the executor (one per plan session). The command orchestrator itself is not profile-controlled: as of v5.9, command frontmatter omits `model:`, so the orchestrator follows the session model. For cheaper runs, opt into `--profile balanced` (Sonnet on brief/research/execute/continue, Opus on plan + review) or `--profile economy` (Sonnet everywhere); for maximum quality, `--profile fable` (Fable 5 on every phase). Per-command cost is published in `${CLAUDE_PLUGIN_DATA}/trek*-stats.jsonl` if you want exact numbers.
|
||||
The default `premium` profile runs **Opus on every phase** — the orchestrator (one per command), the exploration and review swarms (5–10 sub-agents per command, all `model: opus`-pinned in `agents/*.md`), and the executor (one per plan session). The model is **uniform per phase**: there is no "Opus orchestrates, Sonnet runs the swarms" split — a phase resolves to one model and both the orchestrator and its sub-agents use it. For cheaper runs, opt into `--profile balanced` (Sonnet on brief/research/execute/continue, Opus on plan + review) or `--profile economy` (Sonnet everywhere). Per-command cost is published in `${CLAUDE_PLUGIN_DATA}/trek*-stats.jsonl` if you want exact numbers.
|
||||
|
||||
The `opus` alias resolves to **Opus 4.8** (default reasoning effort `high`), `sonnet` to Sonnet 4.6, and `fable` to **Fable 5** (Mythos-class, above Opus; default reasoning effort `high` — xhigh requires a session-level setting, see [`docs/profiles.md`](docs/profiles.md)). Note two distinct effort axes that share the word "effort": brief `phase_signals.effort` (low/standard/high) tunes *orchestration shape* — how many agents and passes run — while native `effort:` on selected agents (retrieval at `medium`, adversarial-reasoning at `high`) tunes the *per-spawn reasoning budget*. See [`docs/profiles.md`](docs/profiles.md) § Model & effort axes.
|
||||
The `opus` alias resolves to **Opus 4.8** (default reasoning effort `high`) and `sonnet` to Sonnet 4.6. Note two distinct effort axes that share the word "effort": brief `phase_signals.effort` (low/standard/high) tunes *orchestration shape* — how many agents and passes run — while native `effort:` on selected agents (retrieval at `medium`, adversarial-reasoning at `high`) tunes the *per-spawn reasoning budget*. See [`docs/profiles.md`](docs/profiles.md) § Model & effort axes.
|
||||
|
||||
For per-profile cost estimates, see [`docs/profiles.md`](docs/profiles.md).
|
||||
|
||||
|
|
@ -849,7 +832,7 @@ trekplan/
|
|||
│ └ 21 spawnable (1 dormant: synthesis-agent, Δ≈0) + 3 orchestrator reference docs (not spawned)
|
||||
├── commands/ 6 slash commands (trekbrief, trekresearch, trekplan, trekexecute, trekreview, trekcontinue) + trekendsession helper
|
||||
├── templates/ Frontmatter templates for brief, research, plan, session, launch
|
||||
├── hooks/ 8 hooks (pre-bash, pre-write, pre-agent-cap, session-title, post-bash-stats, pre-compact-flush, post-compact-flush, otel-export)
|
||||
├── hooks/ 7 hooks (pre-bash, pre-write, session-title, post-bash-stats, pre-compact-flush, post-compact-flush, otel-export)
|
||||
├── lib/ Zero-dep parsers and validators (CLI shims under lib/validators/)
|
||||
├── tests/ comprehensive node:test suite — `npm test` is the fork-readiness gate
|
||||
├── docs/ HANDOVER-CONTRACTS.md + architect-bridge-test.md
|
||||
|
|
@ -936,18 +919,6 @@ suppress this, leave the `architecture/` directory absent from your
|
|||
project directory. Discovery is additive — missing file is fine, no
|
||||
error.
|
||||
|
||||
## Changelog
|
||||
|
||||
Full version history → [CHANGELOG.md](CHANGELOG.md).
|
||||
|
||||
Recent, in one line each:
|
||||
|
||||
- **v5.8.0** — offline gold-scored output eval (SKAL-1·4b): `lib/review/gold-scorer.mjs` grades a committed agent-run fixture against the golden corpus at `(file, rule_key)` granularity; suite census gains a `goldEval` category. Offline + deterministic, no live agent spawn.
|
||||
- **v5.7.1** — leaner always-loaded agent listing (`<example>` blocks moved into agent bodies, ~3,180 tok/turn, no behavior change).
|
||||
- **v5.7.0** — opt-in per-session token/cost metering (SKAL-2) + eval foundation (SKAL-1·4a).
|
||||
- **v5.6.1** — one-line `description:` for the four reference/dormant agents (~700 tok).
|
||||
- **v5.5.0** — brief **framing** enforcement (`brief_version 2.2`) + a `/trekreview` reviewer-schema contract. Additive, no breaking changes.
|
||||
|
||||
## Contributing
|
||||
|
||||
See [CONTRIBUTING.md](CONTRIBUTING.md).
|
||||
|
|
|
|||
|
|
@ -3,6 +3,24 @@ name: architecture-mapper
|
|||
description: |
|
||||
Use this agent when you need deep architecture analysis of a codebase — structure,
|
||||
tech stack, patterns, anti-patterns, and key abstractions.
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase needs architecture overview
|
||||
user: "/trekplan Add authentication to the API"
|
||||
assistant: "Launching architecture-mapper to analyze codebase structure and patterns."
|
||||
<commentary>
|
||||
Phase 5 of trekplan triggers this agent for every codebase size.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to understand an unfamiliar codebase
|
||||
user: "Map out the architecture of this project"
|
||||
assistant: "I'll use the architecture-mapper agent to analyze the codebase structure."
|
||||
<commentary>
|
||||
Direct architecture analysis request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
effort: medium
|
||||
color: cyan
|
||||
|
|
@ -86,23 +104,3 @@ Structure your report with clear sections matching the 7 areas above. Include:
|
|||
- A brief "Architecture Summary" paragraph at the top (3-4 sentences)
|
||||
|
||||
Do NOT include raw file listings — synthesize and organize the information.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase needs architecture overview
|
||||
user: "/trekplan Add authentication to the API"
|
||||
assistant: "Launching architecture-mapper to analyze codebase structure and patterns."
|
||||
<commentary>
|
||||
Phase 5 of trekplan triggers this agent for every codebase size.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to understand an unfamiliar codebase
|
||||
user: "Map out the architecture of this project"
|
||||
assistant: "I'll use the architecture-mapper agent to analyze the codebase structure."
|
||||
<commentary>
|
||||
Direct architecture analysis request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -5,6 +5,24 @@ description: |
|
|||
checks completeness, consistency, testability, scope clarity, and
|
||||
research-plan validity. Catches problems early to avoid wasting tokens on
|
||||
exploration with a flawed brief.
|
||||
|
||||
<example>
|
||||
Context: Voyage runs brief review before exploration
|
||||
user: "/trekplan --project .claude/projects/2026-04-18-notifications"
|
||||
assistant: "Reviewing brief quality before launching exploration agents."
|
||||
<commentary>
|
||||
Orchestrator Phase 1b triggers this agent after the brief is available.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to validate a brief before planning
|
||||
user: "Review this brief for completeness"
|
||||
assistant: "I'll use the brief-reviewer agent to check brief quality."
|
||||
<commentary>
|
||||
Brief review request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
color: magenta
|
||||
tools: ["Read", "Glob", "Grep"]
|
||||
|
|
@ -284,23 +302,3 @@ information that would strengthen the brief. List only if actionable.}
|
|||
files or technologies exist, but deep code analysis is not your job.
|
||||
- **Research-plan checks are load-bearing.** A brief with `research_status: pending`
|
||||
and missing research files is a scope hazard — flag it as a major risk.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage runs brief review before exploration
|
||||
user: "/trekplan --project .claude/projects/2026-04-18-notifications"
|
||||
assistant: "Reviewing brief quality before launching exploration agents."
|
||||
<commentary>
|
||||
Orchestrator Phase 1b triggers this agent after the brief is available.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to validate a brief before planning
|
||||
user: "Review this brief for completeness"
|
||||
assistant: "I'll use the brief-reviewer agent to check brief quality."
|
||||
<commentary>
|
||||
Brief review request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -4,6 +4,26 @@ description: |
|
|||
Use this agent when the research task requires practical, real-world experience rather
|
||||
than official documentation — community sentiment, production war stories, known gotchas,
|
||||
and what developers actually encounter when using a technology.
|
||||
|
||||
<example>
|
||||
Context: trekresearch needs real-world experience data on a database migration
|
||||
user: "/trekresearch What's the real-world experience with migrating from MongoDB to PostgreSQL?"
|
||||
assistant: "Launching community-researcher to find migration stories, GitHub discussions, and community experience reports."
|
||||
<commentary>
|
||||
Official docs won't cover migration regrets or production war stories. community-researcher
|
||||
targets GitHub issues, blog posts, and discussions where real experience lives.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: trekresearch is building a technology comparison
|
||||
user: "/trekresearch Research community sentiment around adopting SvelteKit vs Next.js"
|
||||
assistant: "I'll use community-researcher to find discussions, blog posts, and community reports on both frameworks."
|
||||
<commentary>
|
||||
Framework comparisons live in community discourse, not official docs. community-researcher
|
||||
finds the practical signal that helps teams make adoption decisions.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
color: green
|
||||
tools: ["WebSearch", "WebFetch", "mcp__tavily__tavily_search", "mcp__tavily__tavily_research"]
|
||||
|
|
@ -113,25 +133,3 @@ End with a summary table:
|
|||
Do not pick a side — report the split.
|
||||
- **Flag if a "problem" has since been fixed.** Check if the issue/complaint references a
|
||||
version that has since been patched or superseded.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: trekresearch needs real-world experience data on a database migration
|
||||
user: "/trekresearch What's the real-world experience with migrating from MongoDB to PostgreSQL?"
|
||||
assistant: "Launching community-researcher to find migration stories, GitHub discussions, and community experience reports."
|
||||
<commentary>
|
||||
Official docs won't cover migration regrets or production war stories. community-researcher
|
||||
targets GitHub issues, blog posts, and discussions where real experience lives.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: trekresearch is building a technology comparison
|
||||
user: "/trekresearch Research community sentiment around adopting SvelteKit vs Next.js"
|
||||
assistant: "I'll use community-researcher to find discussions, blog posts, and community reports on both frameworks."
|
||||
<commentary>
|
||||
Framework comparisons live in community discourse, not official docs. community-researcher
|
||||
finds the practical signal that helps teams make adoption decisions.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -4,6 +4,26 @@ description: |
|
|||
Use this agent when the research task has an emerging conclusion that needs adversarial
|
||||
stress-testing — find counter-evidence, overlooked alternatives, and reasons the leading
|
||||
answer might be wrong.
|
||||
|
||||
<example>
|
||||
Context: trekresearch has found evidence favoring a technology and needs the other side
|
||||
user: "/trekresearch We're leaning toward adopting Kafka for our event streaming needs"
|
||||
assistant: "Launching contrarian-researcher to find the strongest arguments against Kafka and what alternatives might serve better."
|
||||
<commentary>
|
||||
The research equivalent of plan-critic. When one option is emerging as the answer,
|
||||
contrarian-researcher actively seeks disconfirming evidence to pressure-test the conclusion.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: trekresearch is comparing options and needs the downsides of the leading candidate
|
||||
user: "/trekresearch Compare Redis vs Memcached — initial research favors Redis"
|
||||
assistant: "I'll use contrarian-researcher to find the strongest case against Redis and scenarios where Memcached wins."
|
||||
<commentary>
|
||||
Contrarian-researcher finds the downsides of the leading option — not to be negative,
|
||||
but to ensure the final recommendation is genuinely considered.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
effort: high
|
||||
color: red
|
||||
|
|
@ -132,25 +152,3 @@ Followed by a **Verdict** section:
|
|||
apply to a read-heavy workload. Assess relevance before reporting.
|
||||
- **Check recency.** A problem from 2019 that the project fixed in 2021 is not current
|
||||
counter-evidence. Flag whether issues are current or historical.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: trekresearch has found evidence favoring a technology and needs the other side
|
||||
user: "/trekresearch We're leaning toward adopting Kafka for our event streaming needs"
|
||||
assistant: "Launching contrarian-researcher to find the strongest arguments against Kafka and what alternatives might serve better."
|
||||
<commentary>
|
||||
The research equivalent of plan-critic. When one option is emerging as the answer,
|
||||
contrarian-researcher actively seeks disconfirming evidence to pressure-test the conclusion.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: trekresearch is comparing options and needs the downsides of the leading candidate
|
||||
user: "/trekresearch Compare Redis vs Memcached — initial research favors Redis"
|
||||
assistant: "I'll use contrarian-researcher to find the strongest case against Redis and scenarios where Memcached wins."
|
||||
<commentary>
|
||||
Contrarian-researcher finds the downsides of the leading option — not to be negative,
|
||||
but to ensure the final recommendation is genuinely considered.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -5,6 +5,24 @@ description: |
|
|||
Produces a structured conventions report covering naming, directory layout,
|
||||
import style, error handling, test patterns, git commit style, and
|
||||
documentation patterns. Uses concrete examples from the codebase.
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase for a medium+ codebase
|
||||
user: "/trekplan Add authentication to the API"
|
||||
assistant: "Launching convention-scanner to discover coding patterns."
|
||||
<commentary>
|
||||
Phase 5 of trekplan triggers this agent for medium+ codebases (50+ files).
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to understand a project's conventions before contributing
|
||||
user: "What are the coding conventions in this project?"
|
||||
assistant: "I'll use the convention-scanner agent to analyze the codebase."
|
||||
<commentary>
|
||||
Direct convention discovery request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
color: yellow
|
||||
tools: ["Read", "Glob", "Grep", "Bash"]
|
||||
|
|
@ -141,23 +159,3 @@ Based on existing conventions, new code should:
|
|||
than scanning everything.
|
||||
- **Stay focused.** This is about conventions — not architecture, dependencies, or risks.
|
||||
Those are handled by other agents.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase for a medium+ codebase
|
||||
user: "/trekplan Add authentication to the API"
|
||||
assistant: "Launching convention-scanner to discover coding patterns."
|
||||
<commentary>
|
||||
Phase 5 of trekplan triggers this agent for medium+ codebases (50+ files).
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to understand a project's conventions before contributing
|
||||
user: "What are the coding conventions in this project?"
|
||||
assistant: "I'll use the convention-scanner agent to analyze the codebase."
|
||||
<commentary>
|
||||
Direct convention discovery request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -3,6 +3,24 @@ name: dependency-tracer
|
|||
description: |
|
||||
Use this agent when you need to trace import chains, map data flow, or understand
|
||||
how modules connect and what side effects they produce.
|
||||
|
||||
<example>
|
||||
Context: Voyage needs to understand module relationships for a task
|
||||
user: "/trekplan Refactor the payment processing pipeline"
|
||||
assistant: "Launching dependency-tracer to map module connections and data flow."
|
||||
<commentary>
|
||||
Phase 5 of trekplan triggers this agent to trace dependencies relevant to the task.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User needs to understand impact of changing a module
|
||||
user: "What would break if I change the User model?"
|
||||
assistant: "I'll use the dependency-tracer agent to trace all dependents of the User model."
|
||||
<commentary>
|
||||
Impact analysis request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
effort: medium
|
||||
color: blue
|
||||
|
|
@ -75,23 +93,3 @@ Structure as:
|
|||
6. **Risk Flags** — circular deps, tight coupling, hidden side effects
|
||||
|
||||
Include file paths and line numbers for every finding.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage needs to understand module relationships for a task
|
||||
user: "/trekplan Refactor the payment processing pipeline"
|
||||
assistant: "Launching dependency-tracer to map module connections and data flow."
|
||||
<commentary>
|
||||
Phase 5 of trekplan triggers this agent to trace dependencies relevant to the task.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User needs to understand impact of changing a module
|
||||
user: "What would break if I change the User model?"
|
||||
assistant: "I'll use the dependency-tracer agent to trace all dependents of the User model."
|
||||
<commentary>
|
||||
Impact analysis request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -3,6 +3,26 @@ name: docs-researcher
|
|||
description: |
|
||||
Use this agent when the research task requires authoritative information from official
|
||||
documentation, RFCs, vendor specifications, or Microsoft/Azure documentation.
|
||||
|
||||
<example>
|
||||
Context: trekresearch needs to ground an OAuth2 implementation in official specs
|
||||
user: "/trekresearch Research OAuth2 PKCE flow for our SPA"
|
||||
assistant: "Launching docs-researcher to find the official RFC and vendor documentation for OAuth2 PKCE."
|
||||
<commentary>
|
||||
docs-researcher targets authoritative sources — RFCs, specs, official vendor docs —
|
||||
not community opinions. This is the right agent for protocol and standards questions.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: trekresearch encounters an Azure-specific technology
|
||||
user: "/trekresearch How should we configure Azure Service Bus for our event pipeline?"
|
||||
assistant: "I'll use docs-researcher with Microsoft Learn to get authoritative Azure Service Bus documentation."
|
||||
<commentary>
|
||||
Microsoft/Azure technologies have dedicated MCP tools (microsoft_docs_search,
|
||||
microsoft_docs_fetch) that docs-researcher uses for higher-quality results.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
color: blue
|
||||
tools: ["WebSearch", "WebFetch", "Read", "mcp__tavily__tavily_search", "mcp__tavily__tavily_research", "mcp__microsoft-learn__microsoft_docs_search", "mcp__microsoft-learn__microsoft_docs_fetch"]
|
||||
|
|
@ -99,25 +119,3 @@ End with a summary table:
|
|||
- **Flag conflicts between official sources.** When vendor docs and the spec disagree, report both.
|
||||
- **Stay focused.** Research only what the research question asks. Do not explore tangentially.
|
||||
- **Official sources only.** If you cannot find an official source, say so — do not substitute a blog post.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: trekresearch needs to ground an OAuth2 implementation in official specs
|
||||
user: "/trekresearch Research OAuth2 PKCE flow for our SPA"
|
||||
assistant: "Launching docs-researcher to find the official RFC and vendor documentation for OAuth2 PKCE."
|
||||
<commentary>
|
||||
docs-researcher targets authoritative sources — RFCs, specs, official vendor docs —
|
||||
not community opinions. This is the right agent for protocol and standards questions.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: trekresearch encounters an Azure-specific technology
|
||||
user: "/trekresearch How should we configure Azure Service Bus for our event pipeline?"
|
||||
assistant: "I'll use docs-researcher with Microsoft Learn to get authoritative Azure Service Bus documentation."
|
||||
<commentary>
|
||||
Microsoft/Azure technologies have dedicated MCP tools (microsoft_docs_search,
|
||||
microsoft_docs_fetch) that docs-researcher uses for higher-quality results.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -5,6 +5,25 @@ description: |
|
|||
needed on a technology choice, architectural question, or complex research topic.
|
||||
Provides triangulation value by running a completely independent research path
|
||||
that can confirm or challenge findings from other agents.
|
||||
|
||||
<example>
|
||||
Context: trekresearch launches gemini-bridge for an independent second opinion on a technology choice
|
||||
user: "/trekplan Should we use Kafka or NATS for our event streaming layer?"
|
||||
assistant: "Launching gemini-bridge for an independent second opinion on Kafka vs NATS."
|
||||
<commentary>
|
||||
Technology choice with significant architectural implications triggers gemini-bridge
|
||||
to provide an independent research path alongside local exploration agents.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: user wants deep research via Gemini on a complex architectural question
|
||||
user: "Get me a Gemini deep research on event sourcing patterns for distributed systems"
|
||||
assistant: "I'll use the gemini-bridge agent to run a deep research on event sourcing patterns."
|
||||
<commentary>
|
||||
Direct request for Gemini research on a complex architectural question triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
color: magenta
|
||||
tools: ["mcp__gemini-mcp__gemini_deep_research", "mcp__gemini-mcp__gemini_get_research_status", "mcp__gemini-mcp__gemini_get_research_result", "mcp__gemini-mcp__gemini_research_followup"]
|
||||
|
|
@ -128,24 +147,3 @@ and other external agents:*
|
|||
- **Graceful degradation at every step.** Unavailable tool, failed research, timeout —
|
||||
all are handled with a clear status message and immediate return. Never leave the
|
||||
pipeline hanging.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: trekresearch launches gemini-bridge for an independent second opinion on a technology choice
|
||||
user: "/trekplan Should we use Kafka or NATS for our event streaming layer?"
|
||||
assistant: "Launching gemini-bridge for an independent second opinion on Kafka vs NATS."
|
||||
<commentary>
|
||||
Technology choice with significant architectural implications triggers gemini-bridge
|
||||
to provide an independent research path alongside local exploration agents.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: user wants deep research via Gemini on a complex architectural question
|
||||
user: "Get me a Gemini deep research on event sourcing patterns for distributed systems"
|
||||
assistant: "I'll use the gemini-bridge agent to run a deep research on event sourcing patterns."
|
||||
<commentary>
|
||||
Direct request for Gemini research on a complex architectural question triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -3,6 +3,24 @@ name: git-historian
|
|||
description: |
|
||||
Use this agent to analyze git history for planning context — recent changes,
|
||||
code ownership, hot files, and active branches relevant to the task.
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase needs git context
|
||||
user: "/trekplan Refactor the database layer"
|
||||
assistant: "Launching git-historian to check recent changes and ownership of DB code."
|
||||
<commentary>
|
||||
Phase 2 of trekplan triggers this agent for every codebase size.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to understand change history before modifying code
|
||||
user: "Who has been changing the auth module recently?"
|
||||
assistant: "I'll use the git-historian agent to analyze ownership and change patterns."
|
||||
<commentary>
|
||||
Git history analysis request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
effort: medium
|
||||
color: yellow
|
||||
|
|
@ -104,23 +122,3 @@ Run `git status --short` to check for:
|
|||
are risks the planner needs to know about.
|
||||
- **Use relative time.** "2 days ago" is more useful than a raw timestamp.
|
||||
- **Never expose email addresses.** Use author names only.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase needs git context
|
||||
user: "/trekplan Refactor the database layer"
|
||||
assistant: "Launching git-historian to check recent changes and ownership of DB code."
|
||||
<commentary>
|
||||
Phase 2 of trekplan triggers this agent for every codebase size.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to understand change history before modifying code
|
||||
user: "Who has been changing the auth module recently?"
|
||||
assistant: "I'll use the git-historian agent to analyze ownership and change patterns."
|
||||
<commentary>
|
||||
Git history analysis request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -3,6 +3,24 @@ name: plan-critic
|
|||
description: |
|
||||
Use this agent when an implementation plan needs adversarial review — it finds
|
||||
problems, never praises.
|
||||
|
||||
<example>
|
||||
Context: Voyage adversarial review phase
|
||||
user: "/trekplan Implement WebSocket real-time updates"
|
||||
assistant: "Launching plan-critic to stress-test the implementation plan."
|
||||
<commentary>
|
||||
Phase 9 of trekplan triggers this agent to review the generated plan.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants a plan reviewed before execution
|
||||
user: "Review this plan and find problems"
|
||||
assistant: "I'll use the plan-critic agent to perform adversarial review."
|
||||
<commentary>
|
||||
Plan review request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
effort: high
|
||||
color: red
|
||||
|
|
@ -279,23 +297,3 @@ short, stable id for the finding class (used for exact-match dedup). Emit
|
|||
|
||||
Be specific. Reference exact plan sections, step numbers, and file paths.
|
||||
Never use "generally" or "usually" — cite the specific problem in this specific plan.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage adversarial review phase
|
||||
user: "/trekplan Implement WebSocket real-time updates"
|
||||
assistant: "Launching plan-critic to stress-test the implementation plan."
|
||||
<commentary>
|
||||
Phase 9 of trekplan triggers this agent to review the generated plan.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants a plan reviewed before execution
|
||||
user: "Review this plan and find problems"
|
||||
assistant: "I'll use the plan-critic agent to perform adversarial review."
|
||||
<commentary>
|
||||
Plan review request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -7,21 +7,12 @@ tools: ["Read", "Glob", "Grep", "Write", "Edit", "Bash"]
|
|||
---
|
||||
|
||||
<!-- Phase mapping: orchestrator → command
|
||||
Corrected in v5.10: every row below was off by one, and the old last row
|
||||
pointed at a ninth command phase that does not exist — the command ends
|
||||
at Phase 8.
|
||||
Orchestrator Phase 1 = Command Phase 4 (Agent group selection)
|
||||
Orchestrator Phase 2 = Command Phase 4 (Parallel research — same
|
||||
command phase; the orchestrator
|
||||
splits selection from launch)
|
||||
(no orchestrator phase)= Command Phase 4.5 (Dimension discovery, high
|
||||
effort AND VOYAGE_STORM_ENABLED
|
||||
=1 only — v5.10)
|
||||
Orchestrator Phase 3 = Command Phase 5 (Targeted follow-ups; bounded
|
||||
conversation loop at high effort)
|
||||
Orchestrator Phase 4 = Command Phase 6 (Triangulation)
|
||||
Orchestrator Phase 5 = Command Phase 7 (Synthesis + write brief)
|
||||
Orchestrator Phase 6 = Command Phase 8 (Present and track / completion)
|
||||
Orchestrator Phase 2 = Command Phase 5 (Parallel research)
|
||||
Orchestrator Phase 3 = Command Phase 6 (Targeted follow-ups)
|
||||
Orchestrator Phase 4 = Command Phase 7 (Triangulation)
|
||||
Orchestrator Phase 5 = Command Phase 8 (Synthesis + write brief)
|
||||
Orchestrator Phase 6 = Command Phase 9 (Completion)
|
||||
As of v2.4.0, /trekresearch runs these phases inline in main
|
||||
context instead of spawning this agent. Keep this file as the canonical
|
||||
reference for what those phases do. -->
|
||||
|
|
|
|||
|
|
@ -3,6 +3,24 @@ name: research-scout
|
|||
description: |
|
||||
Use this agent when the implementation task involves unfamiliar technologies, external
|
||||
APIs, or libraries where official documentation and known issues should be checked.
|
||||
|
||||
<example>
|
||||
Context: Voyage detects external technology in the task
|
||||
user: "/trekplan Integrate Stripe payment processing"
|
||||
assistant: "Launching research-scout to find Stripe documentation and best practices."
|
||||
<commentary>
|
||||
Phase 5 of trekplan conditionally triggers this agent when external tech is detected.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User needs research before implementation
|
||||
user: "Research the best approach for WebSocket scaling"
|
||||
assistant: "I'll use the research-scout agent to find documentation and best practices."
|
||||
<commentary>
|
||||
Research request for external technology triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
color: blue
|
||||
tools: ["WebSearch", "WebFetch", "Read"]
|
||||
|
|
@ -100,23 +118,3 @@ End with a summary table:
|
|||
- **Date everything.** Documentation ages — the reader needs to judge freshness.
|
||||
- **Flag conflicts.** If official docs and community advice disagree, report both.
|
||||
- **Stay focused.** Research only what the task needs. Do not explore tangentially.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage detects external technology in the task
|
||||
user: "/trekplan Integrate Stripe payment processing"
|
||||
assistant: "Launching research-scout to find Stripe documentation and best practices."
|
||||
<commentary>
|
||||
Phase 5 of trekplan conditionally triggers this agent when external tech is detected.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User needs research before implementation
|
||||
user: "Research the best approach for WebSocket scaling"
|
||||
assistant: "I'll use the research-scout agent to find documentation and best practices."
|
||||
<commentary>
|
||||
Research request for external technology triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -3,6 +3,24 @@ name: risk-assessor
|
|||
description: |
|
||||
Use this agent when you need to identify risks, edge cases, failure modes, and
|
||||
technical debt that could affect an implementation task.
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase identifies potential risks
|
||||
user: "/trekplan Migrate database from PostgreSQL to MongoDB"
|
||||
assistant: "Launching risk-assessor to identify failure modes and edge cases for this migration."
|
||||
<commentary>
|
||||
Phase 5 of trekplan triggers this agent to find risks before planning begins.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to understand risks before a change
|
||||
user: "What could go wrong with this refactor?"
|
||||
assistant: "I'll use the risk-assessor agent to map risks and failure modes."
|
||||
<commentary>
|
||||
Risk analysis request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
effort: high
|
||||
color: yellow
|
||||
|
|
@ -88,23 +106,3 @@ Produce a prioritized risk list:
|
|||
**Low** = minor concerns worth noting
|
||||
|
||||
Follow with a narrative section expanding on each Critical and High risk.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase identifies potential risks
|
||||
user: "/trekplan Migrate database from PostgreSQL to MongoDB"
|
||||
assistant: "Launching risk-assessor to identify failure modes and edge cases for this migration."
|
||||
<commentary>
|
||||
Phase 5 of trekplan triggers this agent to find risks before planning begins.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to understand risks before a change
|
||||
user: "What could go wrong with this refactor?"
|
||||
assistant: "I'll use the risk-assessor agent to map risks and failure modes."
|
||||
<commentary>
|
||||
Risk analysis request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -3,6 +3,24 @@ name: scope-guardian
|
|||
description: |
|
||||
Use this agent when you need to verify that an implementation plan matches its
|
||||
requirements — catches scope creep and scope gaps.
|
||||
|
||||
<example>
|
||||
Context: Voyage adversarial review phase checks scope alignment
|
||||
user: "/trekplan Add caching to the API layer"
|
||||
assistant: "Launching scope-guardian to verify plan matches requirements."
|
||||
<commentary>
|
||||
Phase 9 of trekplan triggers this agent alongside plan-critic.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to verify plan doesn't do too much or too little
|
||||
user: "Does this plan match what I asked for?"
|
||||
assistant: "I'll use the scope-guardian agent to check scope alignment."
|
||||
<commentary>
|
||||
Scope verification request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
color: magenta
|
||||
tools: ["Read", "Glob", "Grep"]
|
||||
|
|
@ -126,23 +144,3 @@ One object per scope-creep / gap / dependency finding. `file`/`line` point at th
|
|||
plan step (or the brief requirement) the finding concerns; `rule_key` is a short,
|
||||
stable id for the finding class (used for exact-match dedup against plan-critic).
|
||||
Emit `"findings": []` when the plan is fully aligned.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage adversarial review phase checks scope alignment
|
||||
user: "/trekplan Add caching to the API layer"
|
||||
assistant: "Launching scope-guardian to verify plan matches requirements."
|
||||
<commentary>
|
||||
Phase 9 of trekplan triggers this agent alongside plan-critic.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to verify plan doesn't do too much or too little
|
||||
user: "Does this plan match what I asked for?"
|
||||
assistant: "I'll use the scope-guardian agent to check scope alignment."
|
||||
<commentary>
|
||||
Scope verification request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -3,6 +3,26 @@ name: security-researcher
|
|||
description: |
|
||||
Use this agent when the research task requires security investigation of a technology,
|
||||
dependency, or library — CVEs, audit history, supply chain risks, and OWASP relevance.
|
||||
|
||||
<example>
|
||||
Context: trekresearch is evaluating whether a dependency is safe to adopt
|
||||
user: "/trekresearch Research whether we should trust the `node-fetch` library"
|
||||
assistant: "Launching security-researcher to check CVE history, supply chain risk, and audit reports for node-fetch."
|
||||
<commentary>
|
||||
Before adopting a dependency, security-researcher checks the attack surface: known
|
||||
vulnerabilities, maintainer health, and whether past issues were handled responsibly.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: trekresearch is assessing the security posture of a technology choice
|
||||
user: "/trekresearch Evaluate the security implications of using JWT for session management"
|
||||
assistant: "I'll use security-researcher to check known JWT vulnerabilities, OWASP guidance, and community security reports."
|
||||
<commentary>
|
||||
Technology choices have security tradeoffs. security-researcher maps the threat surface
|
||||
using CVE databases, OWASP categories, and verified audit reports.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
color: red
|
||||
tools: ["WebSearch", "WebFetch", "mcp__tavily__tavily_search", "mcp__tavily__tavily_research"]
|
||||
|
|
@ -120,25 +140,3 @@ End with an overall security summary table:
|
|||
risks from incomplete information.
|
||||
- **Severity matters.** A CVSS 9.8 is not equivalent to a CVSS 3.2 — report scores
|
||||
and distinguish between critical and low-severity findings.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: trekresearch is evaluating whether a dependency is safe to adopt
|
||||
user: "/trekresearch Research whether we should trust the `node-fetch` library"
|
||||
assistant: "Launching security-researcher to check CVE history, supply chain risk, and audit reports for node-fetch."
|
||||
<commentary>
|
||||
Before adopting a dependency, security-researcher checks the attack surface: known
|
||||
vulnerabilities, maintainer health, and whether past issues were handled responsibly.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: trekresearch is assessing the security posture of a technology choice
|
||||
user: "/trekresearch Evaluate the security implications of using JWT for session management"
|
||||
assistant: "I'll use security-researcher to check known JWT vulnerabilities, OWASP guidance, and community security reports."
|
||||
<commentary>
|
||||
Technology choices have security tradeoffs. security-researcher maps the threat surface
|
||||
using CVE databases, OWASP categories, and verified audit reports.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -4,6 +4,24 @@ description: |
|
|||
Use this agent to decompose an trekplan into self-contained headless sessions.
|
||||
Reads a plan file, analyzes step dependencies, groups steps into sessions,
|
||||
identifies parallelism, and generates session specs + dependency graph + launch script.
|
||||
|
||||
<example>
|
||||
Context: User wants to run a plan across multiple headless sessions
|
||||
user: "/trekplan --decompose .claude/plans/trekplan-2026-04-06-auth-refactor.md"
|
||||
assistant: "Launching session-decomposer to split the plan into headless sessions."
|
||||
<commentary>
|
||||
The --decompose flag triggers this agent to analyze and split the plan.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User has a large plan and wants parallel execution
|
||||
user: "Split this plan into sessions I can run in parallel"
|
||||
assistant: "I'll use the session-decomposer to identify parallel session groups."
|
||||
<commentary>
|
||||
Plan decomposition request for parallel headless execution.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
color: green
|
||||
tools: ["Read", "Glob", "Grep", "Write"]
|
||||
|
|
@ -292,23 +310,3 @@ After all sessions complete, run:
|
|||
wrong sequentiality only costs time.
|
||||
- **Verify file existence.** Use Glob to confirm that files referenced in the
|
||||
plan actually exist before assigning them to sessions.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: User wants to run a plan across multiple headless sessions
|
||||
user: "/trekplan --decompose .claude/plans/trekplan-2026-04-06-auth-refactor.md"
|
||||
assistant: "Launching session-decomposer to split the plan into headless sessions."
|
||||
<commentary>
|
||||
The --decompose flag triggers this agent to analyze and split the plan.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User has a large plan and wants parallel execution
|
||||
user: "Split this plan into sessions I can run in parallel"
|
||||
assistant: "I'll use the session-decomposer to identify parallel session groups."
|
||||
<commentary>
|
||||
Plan decomposition request for parallel headless execution.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -4,6 +4,24 @@ description: |
|
|||
Use this agent to find all files, functions, types, and interfaces directly
|
||||
related to the planning task. Replaces generic Explore agents with targeted,
|
||||
structured code discovery.
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase needs task-relevant code
|
||||
user: "/trekplan Add authentication to the API"
|
||||
assistant: "Launching task-finder to locate auth-related code, endpoints, and models."
|
||||
<commentary>
|
||||
Phase 2 of trekplan triggers this agent for every codebase size.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to find code related to a specific feature
|
||||
user: "Find all code related to payment processing"
|
||||
assistant: "I'll use the task-finder agent to locate payment-related code."
|
||||
<commentary>
|
||||
Direct code discovery request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
effort: medium
|
||||
color: green
|
||||
|
|
@ -128,23 +146,3 @@ Structure your report using three tiers:
|
|||
- **Stay focused on the task.** Do not inventory the entire codebase — only what
|
||||
is relevant to implementing the specific task.
|
||||
- **Never read file contents that look like secrets or credentials.**
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase needs task-relevant code
|
||||
user: "/trekplan Add authentication to the API"
|
||||
assistant: "Launching task-finder to locate auth-related code, endpoints, and models."
|
||||
<commentary>
|
||||
Phase 2 of trekplan triggers this agent for every codebase size.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to find code related to a specific feature
|
||||
user: "Find all code related to payment processing"
|
||||
assistant: "I'll use the task-finder agent to locate payment-related code."
|
||||
<commentary>
|
||||
Direct code discovery request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -3,6 +3,24 @@ name: test-strategist
|
|||
description: |
|
||||
Use this agent when you need to design a test strategy for an implementation task —
|
||||
discovers existing patterns, maps coverage gaps, and recommends what tests to write.
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase for medium+ codebase
|
||||
user: "/trekplan Add rate limiting to the API"
|
||||
assistant: "Launching test-strategist to analyze existing test patterns and design test coverage."
|
||||
<commentary>
|
||||
Phase 5 of trekplan triggers this agent for medium and large codebases.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to know how to test a feature
|
||||
user: "What tests should I write for this new feature?"
|
||||
assistant: "I'll use the test-strategist agent to analyze existing patterns and recommend tests."
|
||||
<commentary>
|
||||
Test planning request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
model: opus
|
||||
color: green
|
||||
tools: ["Read", "Glob", "Grep", "Bash"]
|
||||
|
|
@ -77,23 +95,3 @@ For each test, provide:
|
|||
5. **Test Dependencies** — fixtures, mocks, or setup code to create first
|
||||
|
||||
Do NOT write test code. Describe what each test should verify and which patterns to follow.
|
||||
|
||||
## When to use — examples
|
||||
|
||||
<example>
|
||||
Context: Voyage exploration phase for medium+ codebase
|
||||
user: "/trekplan Add rate limiting to the API"
|
||||
assistant: "Launching test-strategist to analyze existing test patterns and design test coverage."
|
||||
<commentary>
|
||||
Phase 5 of trekplan triggers this agent for medium and large codebases.
|
||||
</commentary>
|
||||
</example>
|
||||
|
||||
<example>
|
||||
Context: User wants to know how to test a feature
|
||||
user: "What tests should I write for this new feature?"
|
||||
assistant: "I'll use the test-strategist agent to analyze existing patterns and recommend tests."
|
||||
<commentary>
|
||||
Test planning request triggers the agent.
|
||||
</commentary>
|
||||
</example>
|
||||
|
|
|
|||
|
|
@ -2,6 +2,7 @@
|
|||
name: trekbrief
|
||||
description: Interactive interview that produces a task brief with explicit research plan. Feeds /trekresearch and /trekplan. Optionally orchestrates the full pipeline end-to-end.
|
||||
argument-hint: "[--quick] <task description>"
|
||||
model: opus
|
||||
allowed-tools: Agent, Read, Glob, Grep, Write, Edit, Bash, AskUserQuestion
|
||||
---
|
||||
|
||||
|
|
@ -366,14 +367,13 @@ in the question body so the operator sees why it was picked.
|
|||
### The loop — 4 tier-coupled AskUserQuestion calls
|
||||
|
||||
Loop over `[research, plan, execute, review]` in order. For each phase,
|
||||
issue one `AskUserQuestion` with 4 options:
|
||||
issue one `AskUserQuestion` with 3 options:
|
||||
|
||||
| Option | Maps to phase_signals entry |
|
||||
|--------|----------------------------|
|
||||
| **Low effort** | `{phase: <name>, effort: low, model: sonnet}` |
|
||||
| **Standard (default)** | `{phase: <name>, effort: standard}` *(model omitted — composition falls through to profile)* |
|
||||
| **High effort** | `{phase: <name>, effort: high, model: opus}` |
|
||||
| **Fable (max quality)** | `{phase: <name>, effort: high, model: fable}` |
|
||||
|
||||
The proposed tier per phase (from the default-derivation heuristic) MUST be
|
||||
labelled `(default)` in the option list so the operator can one-click
|
||||
|
|
@ -384,15 +384,6 @@ The mapping table is canonical:
|
|||
- `low → {effort: low, model: sonnet}` (force sonnet for the low-cost path)
|
||||
- `standard → {effort: standard}` (model omitted; composition rule resolves via profile)
|
||||
- `high → {effort: high, model: opus}` (force opus for the high-confidence path)
|
||||
- `fable → {effort: high, model: fable}` (force Fable 5 for the max-quality path)
|
||||
|
||||
The fable tier reuses `effort: high` semantics — full swarm, contrarian +
|
||||
gemini always-on; `EFFORT_LEVELS` is unchanged (Voyage effort is orchestration
|
||||
shape, not model reasoning effort). Model reasoning effort is inherited from
|
||||
the session: Fable 5's default effort is `high`, NOT xhigh. To run xhigh, the
|
||||
operator sets it at session level via `/effort xhigh`, the `effortLevel`
|
||||
setting, or `CLAUDE_CODE_EFFORT_LEVEL` — switching model resets effort to the
|
||||
model default, so it does not follow the model.
|
||||
|
||||
### Force-stop handling
|
||||
|
||||
|
|
@ -900,7 +891,7 @@ Never let stats failures block the workflow.
|
|||
## Profile (v4.1)
|
||||
|
||||
Accepts `--profile <name>` where `<name>` is one of `economy`, `balanced`,
|
||||
`premium`, `fable`, or a custom profile under `voyage-profiles/`. Default: `premium`.
|
||||
`premium`, or a custom profile under `voyage-profiles/`. Default: `premium`.
|
||||
|
||||
Resolution order (per `lib/profiles/resolver.mjs`):
|
||||
1. `--profile` flag (source: `flag`)
|
||||
|
|
|
|||
|
|
@ -2,6 +2,7 @@
|
|||
name: trekcontinue
|
||||
description: Resume the next session in a multi-session trekplan project. Reads .session-state.local.json and immediately begins the next session.
|
||||
argument-hint: "[<project-dir> | --help]"
|
||||
model: opus
|
||||
---
|
||||
|
||||
# Ultracontinue Local v1.0
|
||||
|
|
|
|||
|
|
@ -2,6 +2,7 @@
|
|||
name: trekendsession
|
||||
description: Mark the current session as complete and write session-state pointing at the next session. Helper for informal multi-session flows.
|
||||
argument-hint: "<next-brief-path> <next-label> | --help"
|
||||
model: opus
|
||||
---
|
||||
|
||||
# Voyage End-Session Local v1.0
|
||||
|
|
@ -90,16 +91,16 @@ want an interactive flow, use `/trekcontinue --help` to see the full pipeline.
|
|||
|
||||
## Phase 3 — Atomically write `.session-state.local.json` + sibling NEXT-SESSION-PROMPT.local.md
|
||||
|
||||
Write `{project_dir}/.session-state.local.json` with the schema-v1 object:
|
||||
Write `<project-dir>/.session-state.local.json` with the schema-v1 object:
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 1,
|
||||
"project": "{project_dir}",
|
||||
"next_session_brief_path": "{arg 1}",
|
||||
"next_session_label": "{arg 2}",
|
||||
"project": "<project-dir>",
|
||||
"next_session_brief_path": "<arg 1>",
|
||||
"next_session_label": "<arg 2>",
|
||||
"status": "in_progress",
|
||||
"updated_at": "{now, ISO-8601}"
|
||||
"updated_at": "<now, ISO-8601>"
|
||||
}
|
||||
```
|
||||
|
||||
|
|
@ -114,22 +115,14 @@ Under `node --input-type=module -e "<script>" arg1 arg2 arg3`, Node sets
|
|||
|
||||
This phase ALSO writes a sibling `NEXT-SESSION-PROMPT.local.md` in the
|
||||
project directory with YAML frontmatter (`produced_by: trekendsession`,
|
||||
`produced_at: {ISO-8601}`, `project: {project_dir}`). Both files are written
|
||||
in a single ESM block so the writes succeed or fail together.
|
||||
|
||||
Run the block below via the Bash tool at runtime, substituting the resolved
|
||||
values for the `{curly}` placeholders (Phase 1 gives `{project_dir}`, Phase 2
|
||||
gives `{next_brief_path}` and `{next_label}`). This is NOT an eager-exec
|
||||
block — the values do not exist at command-load time. The import path must
|
||||
stay absolute via `${CLAUDE_PLUGIN_ROOT}` — your Bash cwd is the user's
|
||||
repo, not the plugin root, so a cwd-relative import throws
|
||||
`ERR_MODULE_NOT_FOUND`:
|
||||
`produced_at: <ISO-8601>`, `project: <project-dir>`). Both files are written
|
||||
in a single ESM block so the writes succeed or fail together:
|
||||
|
||||
```bash
|
||||
node --input-type=module -e "
|
||||
!`node --input-type=module -e "
|
||||
import path from 'node:path';
|
||||
import { writeFileSync } from 'node:fs';
|
||||
import { atomicWriteJson } from '${CLAUDE_PLUGIN_ROOT}/lib/util/atomic-write.mjs';
|
||||
import { atomicWriteJson } from './lib/util/atomic-write.mjs';
|
||||
const [, dir, brief, label] = process.argv;
|
||||
const now = new Date().toISOString();
|
||||
const stateObj = { schema_version: 1, project: dir, next_session_brief_path: brief, next_session_label: label, status: 'in_progress', updated_at: now };
|
||||
|
|
@ -140,28 +133,26 @@ const promptBody = '---\\nproduced_by: trekendsession\\nproduced_at: ' + now + '
|
|||
writeFileSync(promptFile, promptBody);
|
||||
console.log(stateFile);
|
||||
console.log(promptFile);
|
||||
" '{project_dir}' '{next_brief_path}' '{next_label}'
|
||||
" '<project-dir>' '<next-brief-path>' '<next-label>'`
|
||||
```
|
||||
|
||||
## Phase 4 — Validate + narrate
|
||||
|
||||
Validate the freshly-written state file via the Bash tool at runtime,
|
||||
substituting the resolved `{project_dir}` (NOT eager-exec — the file does
|
||||
not exist at command-load time):
|
||||
Validate the freshly-written state file:
|
||||
|
||||
```bash
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/validators/session-state-validator.mjs --json {project_dir}/.session-state.local.json
|
||||
!`node lib/validators/session-state-validator.mjs --json <project-dir>/.session-state.local.json`
|
||||
```
|
||||
|
||||
If `valid: true`, print the success block matching `/trekcontinue` Phase 3
|
||||
narration (SC-8 cross-project consistency — same template both sides):
|
||||
|
||||
```
|
||||
Session state written: {project_dir}/.session-state.local.json
|
||||
Session state written: <project-dir>/.session-state.local.json
|
||||
|
||||
Project: {project_dir}
|
||||
Next session: {next_label}
|
||||
Brief: {next_brief_path}
|
||||
Project: <project-dir>
|
||||
Next session: <next-label>
|
||||
Brief: <next-brief-path>
|
||||
|
||||
In a fresh Claude session, run /trekcontinue to resume.
|
||||
```
|
||||
|
|
|
|||
|
|
@ -2,6 +2,7 @@
|
|||
name: trekexecute
|
||||
description: Disciplined plan executor — single-session or multi-session with parallel orchestration, failure recovery, and headless support
|
||||
argument-hint: "[--project <dir>] [--fg | --resume | --dry-run | --validate | --step N | --session N] [plan.md]"
|
||||
model: opus
|
||||
allowed-tools: Read, Write, Edit, Bash, Glob, Grep, AskUserQuestion
|
||||
disallowed-tools: Agent, TeamCreate
|
||||
---
|
||||
|
|
@ -1576,7 +1577,7 @@ Never let stats failures block the workflow.
|
|||
## Profile (v4.1)
|
||||
|
||||
Accepts `--profile <name>` where `<name>` is `economy`, `balanced`, `premium`,
|
||||
`fable`, or a custom profile under `voyage-profiles/`. Default: `premium`.
|
||||
or a custom profile under `voyage-profiles/`. Default: `premium`.
|
||||
|
||||
Resolution order (per `lib/profiles/resolver.mjs`):
|
||||
1. `--profile` flag (source: `flag`)
|
||||
|
|
@ -1608,25 +1609,12 @@ model_for_phase = brief.phase_signals[<phase>]?.model ?? profile.phase_models[
|
|||
```
|
||||
|
||||
The brief signal wins per-phase when present; the profile fills any
|
||||
gaps. Both fields are mechanically resolved by the single composed CLI,
|
||||
invoked in Phase 2.4 alongside the sequencing-gate brief-validator call:
|
||||
|
||||
```bash
|
||||
# v5.9 — composed phase-model resolution (brief > profile > default) for the
|
||||
# execute phase. ONE call returns {effort, model, source}; captured as
|
||||
# phase_signal_result. Append --profile {profile} when the operator passed
|
||||
# --profile.
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/resolver.mjs --resolve-phase-model --phase execute --brief-path "{dir}/brief.md" [--profile {profile}] --json
|
||||
```
|
||||
|
||||
`phase_signal_result.effort` is consumed when picking the execution
|
||||
strategy (gates auto-escalation, parallel-wave choice — see High-effort
|
||||
behavior below). The resolver does NOT control the orchestrator's own
|
||||
model — that is fixed at invocation time (command frontmatter omits
|
||||
`model:`, so it follows the session model) and cannot be switched mid-turn.
|
||||
`/trekexecute` spawns no sub-agent swarm (Hard Rule 10), so
|
||||
`phase_signal_result.model` has no spawn site here; it is returned for
|
||||
cross-command uniformity and stats.
|
||||
gaps. Composition is mechanically resolved via
|
||||
`node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/phase-signal-resolver.mjs`
|
||||
invoked in Phase 2.4; the resolved JSON is captured as `phase_signal_result`
|
||||
and consumed when picking the orchestration model + parallel-wave
|
||||
strategy. The resolver controls only the orchestrator — sub-agents read
|
||||
`model:` from their own `agents/*.md` frontmatter (still pinned to `opus`).
|
||||
|
||||
For `/trekexecute` specifically: `effort == 'low'` activates `--gates open`
|
||||
+ sequential-only execution (no worktree-isolated parallel waves — runs
|
||||
|
|
|
|||
|
|
@ -2,6 +2,7 @@
|
|||
name: trekplan
|
||||
description: Deep implementation planning from a task brief. Requires --brief or --project. Runs parallel specialized agents, optional external research, and adversarial review.
|
||||
argument-hint: "--brief <path> | --project <dir> [--fg | --quick | --research <brief> | --decompose <plan> | --export headless <plan>]"
|
||||
model: opus
|
||||
allowed-tools: Agent, Read, Glob, Grep, Write, Edit, Bash, AskUserQuestion, TaskCreate, TaskUpdate, TeamCreate, TeamDelete
|
||||
---
|
||||
|
||||
|
|
@ -74,11 +75,10 @@ Parse `$ARGUMENTS` for mode flags. Order of precedence:
|
|||
# older brief that sidesteps framing enforcement.
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/validators/brief-validator.mjs --soft --json [--min-version {min_brief_version}] "{dir}/brief.md"
|
||||
|
||||
# v5.9 — composed phase-model resolution (brief > profile > default) for
|
||||
# the plan phase. ONE call returns {effort, model, source}; captured as
|
||||
# phase_signal_result and injected at Agent-spawn sites below.
|
||||
# Append --profile {profile} when the operator passed --profile.
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/resolver.mjs --resolve-phase-model --phase plan --brief-path "{dir}/brief.md" [--profile {profile}] --json
|
||||
# v5.1.1 — resolve per-phase brief-signal for plan phase. Result is
|
||||
# captured as phase_signal_result and used at Agent-spawn sites below
|
||||
# to override the orchestrator model when a signal is present.
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/phase-signal-resolver.mjs --brief "{dir}/brief.md" --phase plan --json
|
||||
|
||||
# Research briefs (if any) — drift-warn only, none of these block the run
|
||||
[ -d "{dir}/research" ] && \
|
||||
|
|
@ -824,7 +824,7 @@ Never let tracking failures block the main workflow.
|
|||
|
||||
## Profile (v4.1)
|
||||
|
||||
Accepts `--profile <name>` where `<name>` is `economy`, `balanced`, `premium`, `fable`,
|
||||
Accepts `--profile <name>` where `<name>` is `economy`, `balanced`, `premium`,
|
||||
or a custom profile under `voyage-profiles/`. Default: `premium`.
|
||||
|
||||
Resolution order (per `lib/profiles/resolver.mjs`):
|
||||
|
|
@ -859,15 +859,13 @@ model_for_phase = brief.phase_signals[<phase>]?.model ?? profile.phase_models[
|
|||
```
|
||||
|
||||
The brief signal wins per-phase when present; the profile fills any
|
||||
gaps. Both fields are mechanically resolved by the single composed CLI
|
||||
`node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/resolver.mjs --resolve-phase-model`
|
||||
invoked in Phase 1; the resolved JSON `{effort, model, source}` is captured
|
||||
as `phase_signal_result` and passed to `Agent` tool calls explicitly. The
|
||||
resolver controls the `model` parameter at Agent-spawn sites only — the
|
||||
orchestrator's own model is fixed at invocation time (command frontmatter
|
||||
omits `model:`, so it follows the session model) and cannot be switched
|
||||
mid-turn. Sub-agents fall back to `model:` in their own `agents/*.md`
|
||||
frontmatter when no spawn-site injection happens.
|
||||
gaps. Composition is mechanically resolved via
|
||||
`node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/phase-signal-resolver.mjs`
|
||||
invoked in Phase 1; the resolved JSON is captured as `phase_signal_result`
|
||||
and passed to `Agent` tool calls explicitly. The resolver controls only
|
||||
the orchestrator and the model parameter at Agent-spawn sites — sub-agents
|
||||
otherwise read `model:` from their own `agents/*.md` frontmatter (still
|
||||
pinned to `opus`).
|
||||
|
||||
For `/trekplan` specifically: `effort == 'low'` activates the existing
|
||||
`--quick`-equivalent code-path (skip Phase 5 agent swarm — plan directly
|
||||
|
|
@ -912,11 +910,10 @@ Standard and low effort: do NOT run the additional pass.
|
|||
inadequate, stop and ask the user to run `/trekbrief` again.
|
||||
- **Scope**: Only explore the current working directory and its subdirectories.
|
||||
Never read files outside the repo (no ~/.env, no credentials, no other repos).
|
||||
- **Cost**: Model resolution at Agent-spawn sites is a three-layer fallback:
|
||||
brief `phase_signals[<phase>].model` > `profile.phase_models[<phase>]` >
|
||||
agent frontmatter `model:`. The composed resolver returns the first two
|
||||
layers as `phase_signal_result.model`; spawn sites inject it, and agent
|
||||
frontmatter is the fallback when no injection happens.
|
||||
- **Cost**: Sub-agents use their pinned `model:` frontmatter (currently `opus`).
|
||||
When `phase_signals[<phase>].model` is set, the orchestrator AND Agent-spawn
|
||||
sites use the resolved model (`phase_signal_result.model`) for that phase.
|
||||
Frontmatter is the default; brief signal is the per-phase override.
|
||||
- **Privacy**: Never log, store, or repeat file contents that look like
|
||||
secrets, tokens, or credentials. Never log prompt text.
|
||||
- **No premature execution**: Do not modify any project files until the user
|
||||
|
|
|
|||
|
|
@ -1,7 +1,8 @@
|
|||
---
|
||||
name: trekresearch
|
||||
description: Deep research combining local codebase analysis with external knowledge, producing structured research briefs with triangulation and confidence ratings
|
||||
argument-hint: "[--project <dir>] [--quick | --local | --external | --fg] [--engine swarm|deep-research] <research question>"
|
||||
argument-hint: "[--project <dir>] [--quick | --local | --external | --fg] <research question>"
|
||||
model: opus
|
||||
allowed-tools: Agent, Read, Glob, Grep, Write, Edit, Bash, AskUserQuestion, WebSearch, WebFetch, mcp__tavily__tavily_search, mcp__tavily__tavily_research
|
||||
---
|
||||
|
||||
|
|
@ -54,18 +55,15 @@ Supported flags:
|
|||
Create `{dir}/research/` if it does not already exist.
|
||||
|
||||
When `{dir}/brief.md` exists, ALWAYS run the brief-validator (soft mode)
|
||||
AND the composed phase-model resolver for this command's phase before
|
||||
continuing. The resolver's JSON output `{effort, model, source}`
|
||||
(brief signal > profile > default) is captured as `phase_signal_result`
|
||||
and used at Agent-spawn sites in Phase 4 to inject the resolved model:
|
||||
AND the phase-signal-resolver for this command's phase before continuing.
|
||||
The resolver's JSON output is captured as `phase_signal_result` and used
|
||||
at Agent-spawn sites in Phase 4 to inject the brief-resolved model:
|
||||
|
||||
```bash
|
||||
# When --min-brief-version was passed, append --min-version {min_brief_version}
|
||||
# so an older brief raises BRIEF_VERSION_BELOW_MINIMUM (warn, never block).
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/validators/brief-validator.mjs --soft --json [--min-version {min_brief_version}] "{dir}/brief.md"
|
||||
# v5.9 — composed resolver: ONE call returns {effort, model, source}.
|
||||
# Append --profile {profile} when the operator passed --profile.
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/resolver.mjs --resolve-phase-model --phase research --brief-path "{dir}/brief.md" [--profile {profile}] --json
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/phase-signal-resolver.mjs --brief "{dir}/brief.md" --phase research --json
|
||||
```
|
||||
|
||||
6. `--gates` — autonomy control. When present, set `gates_mode = true`. The
|
||||
|
|
@ -82,16 +80,6 @@ Supported flags:
|
|||
enforcement only fires at `≥ 2.2`. Absent → no version check. See
|
||||
`docs/HANDOVER-CONTRACTS.md` §Handover 1 for the pre-2.2 enforcement hole.
|
||||
|
||||
8. `--engine <name>` — opt-in external-research engine. Accepts `--engine <name>`
|
||||
where `<name>` is `swarm` or `deep-research`. **Default: `swarm`** (unchanged
|
||||
behavior). `swarm` runs Voyage's own external-research agent swarm;
|
||||
`deep-research` delegates the external phase to Claude Code's built-in
|
||||
`/deep-research` dynamic workflow and adapts its report into the research-brief
|
||||
schema (requires Claude Code 2.1.154+ and dynamic workflows enabled; falls back
|
||||
to `swarm` and notes the fallback if unavailable — never hard-fails). Orthogonal
|
||||
to `--profile`/`phase_signals`; only affects the external phase. Set
|
||||
**engine = {swarm|deep-research}** (the *requested* engine).
|
||||
|
||||
Flags can be combined:
|
||||
- `--local` — local-only research
|
||||
- `--external --quick` — external-only, lightweight
|
||||
|
|
@ -99,7 +87,7 @@ Flags can be combined:
|
|||
- `--quick` alone implies both local and external (lightweight)
|
||||
|
||||
Defaults: **scope = both**, **execution = foreground** (only mode as of
|
||||
v2.4.0), **project_dir = none**, **engine = swarm**.
|
||||
v2.4.0), **project_dir = none**.
|
||||
|
||||
After stripping flags, the remaining text is the **research question**.
|
||||
|
||||
|
|
@ -120,7 +108,6 @@ Modes:
|
|||
--external Only external research agents (skip codebase analysis)
|
||||
--fg No-op alias (foreground is the only mode as of v2.4.0)
|
||||
--project Write brief into an trekbrief project folder (auto-indexed)
|
||||
--engine Opt-in external-research engine: swarm (default) | deep-research
|
||||
|
||||
Flags can be combined: --local, --external --quick, --project <dir> --external
|
||||
|
||||
|
|
@ -131,7 +118,6 @@ Examples:
|
|||
/trekresearch --external What are the security implications of using Redis for sessions?
|
||||
/trekresearch --fg --local What patterns does this codebase use for database access?
|
||||
/trekresearch --project .claude/projects/2026-04-18-jwt-auth --external What JWT library is best for Node.js?
|
||||
/trekresearch --project <dir> --external --engine deep-research <research question>
|
||||
```
|
||||
|
||||
Do not continue past this step if no question was provided.
|
||||
|
|
@ -140,7 +126,6 @@ Report the detected mode:
|
|||
```
|
||||
Mode: {default | quick}, Scope: {both | local | external}, Execution: foreground
|
||||
Project: {project_dir or "-"}
|
||||
Engine (requested): {swarm | deep-research}
|
||||
Question: {research question}
|
||||
```
|
||||
|
||||
|
|
@ -307,63 +292,6 @@ For each local agent, prompt with the research question, NOT a task description:
|
|||
- convention-scanner: "Discover coding conventions relevant to evaluating {question}.
|
||||
What patterns would a solution need to follow?"
|
||||
|
||||
### Engine selection (scope = both or external)
|
||||
|
||||
`--engine` affects ONLY the external portion of research. The local agents
|
||||
(`### Local agents` above) and Phases 6–7 (triangulation, synthesis, brief
|
||||
writing) are **engine-agnostic** — they run identically regardless of engine.
|
||||
|
||||
`--engine` is **moot** (treated as `swarm`) whenever the external phase does not
|
||||
run at all: `--local`, `--quick`, `effort == 'low'`, or a profile with
|
||||
`external_research_enabled == false` (the `economy`/`balanced` auto-disable — see
|
||||
Profile below). The profile's on/off switch wins. Initialize
|
||||
`effective_engine = {requested engine}`.
|
||||
|
||||
**engine = swarm (default):** run the `### External agents` + `### Bridge agent`
|
||||
blocks below unchanged. This is byte-for-byte the current path, so `--engine swarm`
|
||||
changes nothing (SC1). Keep the native-swarm anchors intact ("in parallel",
|
||||
"single message", `model: "opus"`).
|
||||
|
||||
**engine = deep-research:**
|
||||
|
||||
1. **Coarse pre-gate (best-effort, NOT a trust signal).** `Bash: claude --version`;
|
||||
parse the leading `X.Y.Z` (e.g. from `2.1.196 (Claude Code)`) and compare
|
||||
numerically against `2.1.154` — split each on `.` and compare major, then minor,
|
||||
then patch as integers (do NOT string-compare; lexical comparison mis-orders
|
||||
multi-digit patch numbers). If the version is `< 2.1.154`, OR if
|
||||
`disableWorkflows: true` / `CLAUDE_CODE_DISABLE_WORKFLOWS=1` is set, skip to the
|
||||
fallback (step 4). **If `claude` is not on PATH inside the Bash tool (possible
|
||||
under `claude -p`) or the version cannot be parsed, treat the pre-gate as
|
||||
*indeterminate* and proceed to step 2 — do NOT hard-fail.** There is no positive
|
||||
availability probe (research Dim 4), so a passing pre-gate does not guarantee the
|
||||
workflow runs; the post-hoc check (step 3) is the authoritative guard.
|
||||
|
||||
2. **Run.** Instruct Claude (in prose, this turn) to run
|
||||
`/deep-research <research question>` and request per-claim citations. Note:
|
||||
interactive default/acceptEdits triggers a per-run approval prompt; `claude -p` /
|
||||
SDK / bypass runs immediately.
|
||||
|
||||
3. **Post-hoc presence + provenance check (the real guard).** Verify a real, cited
|
||||
`/deep-research` report actually landed in context — substantive findings with
|
||||
citations, not an empty/denied/errored turn and not bare error text. This check
|
||||
must be **robust to all failure manifestations** (workflow disabled, approval
|
||||
denied, runtime error, empty output), because the disabled-headless behavior is
|
||||
undocumented: no recognizable cited report in context → fall back, regardless of
|
||||
how the failure surfaces.
|
||||
|
||||
4. **On no real report (fallback):** set `effective_engine = swarm`, run the swarm
|
||||
blocks below, and **log the fallback at this decision point** — print
|
||||
`Engine: deep-research → swarm (fallback: <reason>)` and carry the reason into the
|
||||
Phase-8 Present summary and the brief's `## Executive Summary`. **NEVER fabricate
|
||||
or synthesize a substitute report** — a structurally-valid-but-invented brief
|
||||
passes the structure-only validator and silently poisons `/trekplan`; that is the
|
||||
worst outcome of this feature.
|
||||
|
||||
5. **On a real report:** keep `effective_engine = deep-research`, log
|
||||
`Engine: deep-research (active)`, and carry the report into Phase 6 triangulation
|
||||
as the external-findings input (adapted in Phase 7 — see the Deep-research engine
|
||||
adapter below).
|
||||
|
||||
### External agents (scope = both or external)
|
||||
|
||||
Launch the new research-specialized agents:
|
||||
|
|
@ -391,248 +319,15 @@ other agents — the value of Gemini is independence.
|
|||
small = halved, medium/large = default
|
||||
- convention-scanner: medium+ codebases only (50+ files)
|
||||
|
||||
## Phase 4.5 — Dimension discovery
|
||||
|
||||
**Skip this phase entirely unless `phase_signal_result.effort == 'high'` AND
|
||||
`VOYAGE_STORM_ENABLED=1`.** Both conditions, never either.
|
||||
|
||||
This phase never invokes `research-loop-cap.mjs`, so the cap's own flag check
|
||||
does not cover it — the flag has to be read here. Gating on effort alone would
|
||||
leave discovery mutating the dimension list at `effort: high` with the flag
|
||||
unset, making `dimensions` diverge from `dimensions_baseline` and putting the
|
||||
decline branch out of reach for half the mechanism. Doing nothing must leave
|
||||
**both** STORM phases inert.
|
||||
|
||||
Phase 4 retrieves more than the interview knew to ask for. This phase mines
|
||||
that surplus: findings that were **retrieved but unintegrated** — material an
|
||||
agent surfaced that no interview dimension claims.
|
||||
|
||||
1. **Mine.** Walk the Phase-4 agent results and collect findings that map to
|
||||
no existing dimension.
|
||||
2. **Rerank.** Order candidates by relevance to the research question **and**
|
||||
dissimilarity to the dimensions already on the list. A candidate that
|
||||
restates an existing dimension is not a discovery.
|
||||
3. **Augment under the existing ceiling.** Append candidates to the dimension
|
||||
list only while the **whole** list (interview + discovered) stays at or
|
||||
below `maxDimensions: 8` (`settings.json:16`). The ceiling is **not**
|
||||
raised here, so the documented 3–8 dimension range stays true and the
|
||||
README prose about it stays untouched. If the interview already produced 8
|
||||
dimensions, this phase discovers nothing and says so.
|
||||
|
||||
**The ceiling has a reader — use it.** Once the final list is settled, run
|
||||
the check below. Exit 1 means the list exceeded the ceiling: drop discovered
|
||||
dimensions until it passes. Do not proceed to Phase 5 on a rejected list —
|
||||
the turn budget is sized against this same ceiling, so a list over it spends
|
||||
a budget that was never approved for it.
|
||||
|
||||
```bash
|
||||
# Same VOYAGE_ROOT resolution as the per-turn protocol in Phase 5. Exit 0 =
|
||||
# within the ceiling, exit 1 = rejected. JSON on stdout: {ok, count, ceiling, reason?}
|
||||
node "$VOYAGE_ROOT/lib/util/research-loop-cap.mjs" --check-dimensions {final dimension count}
|
||||
```
|
||||
|
||||
The ceiling constant is `MAX_TOTAL_DIMENSIONS` in
|
||||
`lib/util/research-loop-cap.mjs` — deliberately the same constant that sizes
|
||||
the Phase 5 turn budget, so the two axes of the bounded-cost NFR cannot end
|
||||
up enforcing different numbers for one `settings.json:16` value.
|
||||
4. **Record the baseline.** Keep the interview-derived count as
|
||||
`dimensions_baseline` so the discovered delta is machine-readable against
|
||||
the final `dimensions` (Phase 8 stats).
|
||||
5. **Attest membership, not just the count.** Set
|
||||
`dimensions_baseline_preserved: true` only if EVERY interview-derived
|
||||
dimension is still on the final list — this phase appends, it never replaces.
|
||||
Set it `false` if any was dropped, merged away, or rewritten. A count delta
|
||||
cannot show this: dropping two interview dimensions and appending three
|
||||
discovered ones is `+1` and still not a superset, which is exactly what the
|
||||
Success Criterion forbids. `storm-measure.mjs --activation-check` reads the
|
||||
field and fails when it is absent, so omitting it is not the silent default.
|
||||
|
||||
Every outbound query generated from a discovered dimension passes
|
||||
`query-privacy-gate.mjs` before it leaves the machine — see the per-turn
|
||||
protocol in Phase 5. That gate controls the **egress** risk this phase's
|
||||
Independence crossing creates: local paths and identifiers travelling inside a
|
||||
query. It does **not** control the **bias** risk — it inspects query content and
|
||||
cannot stop a local finding from steering an external agent's question. The bias
|
||||
controls are structural (blind initial swarm, append-only crossing, unconditional
|
||||
`contrarian-researcher` at `effort: high`); see Hard rules → Independence.
|
||||
|
||||
## Phase 5 — Targeted follow-ups
|
||||
|
||||
Review all agent results. Identify knowledge gaps — dimensions where findings
|
||||
are thin, contradictory, or missing (**under-illuminated dimensions**).
|
||||
Review all agent results. Identify knowledge gaps — areas where findings are
|
||||
thin, contradictory, or missing.
|
||||
|
||||
**Standard and low effort — unchanged single pass.** For each significant gap,
|
||||
launch a targeted follow-up agent (model: "opus") with a narrow, specific
|
||||
brief. Maximum 2 follow-ups. If no gaps exist, skip: "Initial research
|
||||
sufficient — no follow-ups needed." Then go to Phase 6.
|
||||
For each significant gap, launch a targeted follow-up agent (model: "opus")
|
||||
with a narrow, specific brief. Maximum 2 follow-ups.
|
||||
|
||||
**The bounded loop below runs ONLY when `phase_signal_result.effort == 'high'`**
|
||||
(resolved in Phase 1; see `### High-effort behavior (v5.1.1)`). At any other
|
||||
effort this whole sub-section is inert — no loop, no cap ledger, no new
|
||||
counters beyond zero.
|
||||
|
||||
### Loop bound
|
||||
|
||||
**Maximum 3 turns per under-illuminated dimension.** The bound is per
|
||||
dimension, not per run: the worst case is 3 turns × the whole dimension list
|
||||
under the `maxDimensions: 8` ceiling (`settings.json:16`), which is what
|
||||
`research-loop-cap.mjs` sizes itself against. The cap counts itself from its
|
||||
own append-only ledger — it never asks this prose how many turns it has used.
|
||||
|
||||
The loop is **default-off**: `research-loop-cap.mjs` grants a budget of 0
|
||||
unless `VOYAGE_STORM_ENABLED=1`. Doing nothing leaves the mechanism off.
|
||||
|
||||
### Loop scope marker
|
||||
|
||||
`hooks/scripts/pre-agent-cap.mjs` (PreToolUse on `WebSearch|WebFetch|Task`)
|
||||
enforces the same bound in the harness rather than trusting this prose — but it
|
||||
enforces **only** for a session that carries a scope marker, and allows
|
||||
unconditionally for every session that does not. That is what keeps a globally
|
||||
wired PreToolUse hook from denying tool calls in unrelated sessions. Write the
|
||||
marker once, immediately before the first turn:
|
||||
|
||||
`CLAUDE_PLUGIN_DATA` is **empty in the Bash tool's process env** even in a
|
||||
plugin-enabled session, so the root is resolved with the same fallback
|
||||
`research-loop-cap.mjs` uses — `~/.claude/voyage`. Reader and writer must
|
||||
resolve identically; a marker written where the hook does not look leaves the
|
||||
hook allowing unconditionally while the docs call it enforcing.
|
||||
|
||||
```bash
|
||||
# Arms the PreToolUse cap for THIS session only.
|
||||
# CLAUDE_CODE_SESSION_ID is the same id the hook reads as `session_id`.
|
||||
DATA="${CLAUDE_PLUGIN_DATA:-$HOME/.claude/voyage}"
|
||||
case "$DATA" in /*) SCOPE_DIR="$DATA/trekresearch-loop-scope" ;; *) SCOPE_DIR="" ;; esac
|
||||
if [ -n "$SCOPE_DIR" ] && [ -n "${CLAUDE_CODE_SESSION_ID:-}" ] && mkdir -p "$SCOPE_DIR" 2>/dev/null; then
|
||||
printf '{"runId":"%s","startedAt":"%s"}\n' "{run_id}" "$(date -u +%Y-%m-%dT%H:%M:%SZ)" \
|
||||
> "$SCOPE_DIR/${CLAUDE_CODE_SESSION_ID}.json"
|
||||
else
|
||||
echo "[voyage] scope marker not written — harness cap stays inert for this run"
|
||||
fi
|
||||
```
|
||||
|
||||
An empty `CLAUDE_CODE_SESSION_ID` is checked before the path is composed, not
|
||||
after: unset, the marker becomes `.json`, which no hook lookup matches and no
|
||||
TTL sweep ever cleans up.
|
||||
|
||||
`runId` MUST be the same `{run_id}` passed to `research-loop-cap.mjs --run-id`.
|
||||
The hook counts ledger lines carrying that id, so a marker written with any
|
||||
other id counts zero turns and enforces nothing.
|
||||
|
||||
**A failed marker write is not a reason to stop.** The hook is defence in
|
||||
depth; `research-loop-cap.mjs` is the gate and stays correct on its own. Report
|
||||
the failure to the operator and run the loop. The reverse — skipping the budget
|
||||
gate because a marker exists — is never allowed.
|
||||
|
||||
**Removal belongs to every exit below, especially the exhausted one.** The hook
|
||||
denies `WebSearch`/`WebFetch`/`Task` once `research-loop-cap.mjs` has denied a
|
||||
turn — the gate records its own denials, so the LAST granted turn still runs its
|
||||
queries and exhaustion reaches you through exit 2 of the budget gate below, not
|
||||
through a blocked tool call. Once denied, the hook keeps denying for as long as
|
||||
the marker is there — including Phase 6, which spawns agents. A marker that outlives the loop turns a bound on this loop into a brick
|
||||
on the rest of the session. Cleanup covers the three exits and nothing else: a
|
||||
crashed session runs no cleanup at all, and is covered instead by the hook's
|
||||
TTL (default 2h, `VOYAGE_CAP_SCOPE_TTL_MS`), which auto-resets a stale marker.
|
||||
A crash mid-loop leaves no denial record, so a resumed session is not blocked by
|
||||
it; only a crash AFTER the cap denied a turn hands the resume a deny window, and
|
||||
every denial prints the marker path to delete.
|
||||
|
||||
```bash
|
||||
# Removal — idempotent, safe to repeat. Same root, same absolute-path guard as
|
||||
# the write: a remove that accepts a root the write rejected (or vice versa)
|
||||
# leaves markers the loop believes it cleaned up.
|
||||
DATA="${CLAUDE_PLUGIN_DATA:-$HOME/.claude/voyage}"
|
||||
case "$DATA" in /*) SCOPE_DIR="$DATA/trekresearch-loop-scope" ;; *) SCOPE_DIR="" ;; esac
|
||||
[ -n "$SCOPE_DIR" ] && [ -n "${CLAUDE_CODE_SESSION_ID:-}" ] && \
|
||||
rm -f "$SCOPE_DIR/${CLAUDE_CODE_SESSION_ID}.json"
|
||||
```
|
||||
|
||||
### Per-turn protocol
|
||||
|
||||
Each turn targets exactly one under-illuminated dimension, and runs two gates
|
||||
before it spends anything:
|
||||
|
||||
```bash
|
||||
# 0. Resolve the plugin root ONCE. ${CLAUDE_PLUGIN_ROOT} is substituted in this
|
||||
# command's text but is EMPTY in the Bash tool's process env, and a bare
|
||||
# `node ${CLAUDE_PLUGIN_ROOT}/lib/…` then runs `node /lib/…`, which exits 1 —
|
||||
# indistinguishable from a gate that said no.
|
||||
VOYAGE_ROOT="${CLAUDE_PLUGIN_ROOT:-}"
|
||||
case "$VOYAGE_ROOT" in
|
||||
/*) ;;
|
||||
*) VOYAGE_ROOT="$(ls -d "$HOME"/.claude/plugins/cache/*/voyage 2>/dev/null | head -1)" ;;
|
||||
esac
|
||||
if [ ! -f "$VOYAGE_ROOT/lib/util/research-loop-cap.mjs" ]; then
|
||||
echo "[voyage] gates could not run — plugin root unresolved (exit 2). NOT a denial:"
|
||||
echo " stop the loop and report to the operator. Never proceed ungated."
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# 1. Budget gate — per turn, per dimension. Exit 0 = granted, exit 1 = denied.
|
||||
# JSON on stdout: {ok, used, budget, reason?}
|
||||
node "$VOYAGE_ROOT/lib/util/research-loop-cap.mjs" \
|
||||
--run-id {run_id} --dimension {dimension} --effort {phase_signal_result.effort}
|
||||
|
||||
# 2. Privacy gate — EVERY outbound query, before it leaves the machine.
|
||||
# Exit 0 = send as-is; exit 1 = rewrite the query and re-gate. Never bypass.
|
||||
node "$VOYAGE_ROOT/lib/validators/query-privacy-gate.mjs" "{query text}"
|
||||
```
|
||||
|
||||
**Exit 2 is not exit 1.** A denied budget gate is an exit condition, not a
|
||||
retry. A failed privacy gate is a rewrite: the hard-block tier (secret-shaped
|
||||
strings) is never operator-overridable, so a query that trips it must be
|
||||
reformulated, not forced through. A gate that *could not run* is neither — no
|
||||
rewrite can clear it, so treat it as a hard stop and say so, rather than
|
||||
rewriting a query that was never the problem.
|
||||
|
||||
**Empty turns.** A turn that returns no findings, or findings without
|
||||
citations, is marked `empty`. An empty turn is counted in `empty_turns` and
|
||||
does NOT re-target the same dimension — re-asking the same question of the
|
||||
same silence is how a bounded loop turns into an unbounded one. Move to the
|
||||
next under-illuminated dimension, or exit.
|
||||
|
||||
### Exits (all three, always one of them)
|
||||
|
||||
1. **Converged** — the dimension carries findings with citations and no
|
||||
remaining contradiction. Stop turning on it. This is the normal exit. Once
|
||||
the last dimension has converged, remove the scope marker.
|
||||
2. **Cap exhausted** — `research-loop-cap.mjs` denies the turn. Print the
|
||||
exhaustion **visibly** to the operator, never silently:
|
||||
`Loop bound reached for dimension {dimension} after {N} turns — remaining
|
||||
gaps are carried into the brief as open questions.` A silent cap is
|
||||
indistinguishable from convergence, and that confusion is exactly what this
|
||||
phase exists to prevent. Then remove the scope marker — leaving it here is
|
||||
what would block Phase 6.
|
||||
3. **Operator stop** — the operator interrupts. Remove the scope marker first,
|
||||
then carry whatever has been gathered into Phase 6 and record the remaining
|
||||
gaps as open questions. Do not re-enter the loop after a stop.
|
||||
|
||||
### When the loop does not apply
|
||||
|
||||
**No-brief default.** Without `--project` (or with a project whose `brief.md`
|
||||
is absent), there are no `phase_signals` to resolve, so `effort = 'standard'`,
|
||||
the loop does not run, and all new counters (`conv_turns`, `empty_turns`) are
|
||||
emitted as `0`.
|
||||
|
||||
**Precedence matrix — each entry independently makes the loop moot**, the same
|
||||
way `--engine` is moot when the external phase does not run (see the moot gate
|
||||
in Phase 4):
|
||||
|
||||
| Condition | Effect on the loop |
|
||||
|-----------|--------------------|
|
||||
| `--quick` | Moot — Phase 3.5 skips to Phase 8; the swarm never runs |
|
||||
| `--local` | Moot — no outbound queries to bound |
|
||||
| `external_research_enabled: false` (profile) | Moot — the profile's on/off switch wins |
|
||||
|
||||
**Interaction rule.** A brief that carries `effort: high` **without** a
|
||||
`model`, under a cheap profile (`economy`/`balanced`): the effort signal
|
||||
governs orchestration shape, so the loop is armed, but the profile still
|
||||
supplies the model — and if that profile disables external research, the
|
||||
matrix above wins and the loop is moot regardless of effort.
|
||||
|
||||
**Honesty (hard rule, restated for this loop).** More turns do not make a
|
||||
finding more credible. Turn count is a cost, not evidence: report what the
|
||||
citations support, and let an exhausted cap show up as open questions rather
|
||||
than as confidence.
|
||||
If no gaps exist, skip: "Initial research sufficient — no follow-ups needed."
|
||||
|
||||
## Phase 6 — Triangulation
|
||||
|
||||
|
|
@ -678,45 +373,6 @@ Write the brief to the `brief_destination` computed in Phase 1:
|
|||
|
||||
Create the parent directory if it does not exist.
|
||||
|
||||
### Deep-research engine adapter (engine = deep-research only)
|
||||
|
||||
**Only when `effective_engine == deep-research`.** The swarm path skips this
|
||||
entirely — its findings already flow through Phases 6–7 unchanged (SC1).
|
||||
|
||||
Transform the in-context `/deep-research` report INTO
|
||||
`@${CLAUDE_PLUGIN_ROOT}/templates/research-brief-template.md` — do NOT paste the
|
||||
raw report. Specifically:
|
||||
|
||||
- Reduce the report to ≥ 1 `### {Dimension} -- Confidence: {high|medium|low}`
|
||||
entry, each carrying **External findings** bullets with per-claim source URLs.
|
||||
Local findings still come from the local agents (Phase 4) and are merged in per
|
||||
dimension as usual.
|
||||
- Emit a numeric `confidence ∈ [0,1]` in frontmatter and a 3-sentence
|
||||
`## Executive Summary` (answer, confidence, key caveat).
|
||||
- Populate `## Sources` from the report's citations.
|
||||
- **If the report lacks per-claim URLs, lower the confidence and note the gap in
|
||||
`## Open Questions` — do NOT fabricate URLs.** Provenance you cannot cite is not
|
||||
provenance.
|
||||
- If the report is large, bound the transform to the top dimensions to avoid
|
||||
context truncation.
|
||||
|
||||
### Output self-check (engine = deep-research only)
|
||||
|
||||
**Only when `effective_engine == deep-research`.** After writing to
|
||||
`brief_destination`, run the output validator and repair-or-fall-back. This mirrors
|
||||
the trekplan Phase-8 write→validate→repair self-check; the swarm path does NOT run
|
||||
it, so swarm behavior is unchanged (SC1):
|
||||
|
||||
```bash
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/validators/research-validator.mjs --json "{brief_destination}"
|
||||
```
|
||||
|
||||
On `valid: false`, repair the brief to satisfy the reported errors and re-run the
|
||||
validator. If it cannot be made valid (e.g. the report was too thin to yield even
|
||||
one dimension), set `effective_engine = swarm`, fall back to the swarm engine for
|
||||
this run (and log the fallback per the Engine selection step), rather than emit an
|
||||
invalid brief.
|
||||
|
||||
## Phase 8 — Present and track
|
||||
|
||||
Present a summary to the user:
|
||||
|
|
@ -728,7 +384,6 @@ Present a summary to the user:
|
|||
**Mode:** {default | quick}, Scope: {both | local | external}
|
||||
**Brief:** {brief_destination}
|
||||
**Project:** {project_dir or "-"}
|
||||
**Engine (effective):** {swarm | deep-research}{, with fallback reason if it fell back}
|
||||
**Confidence:** {overall confidence 0.0-1.0}
|
||||
**Dimensions:** {N} researched
|
||||
**Agents:** {N} local + {N} external + {gemini: used | unavailable | skipped}
|
||||
|
|
@ -763,17 +418,10 @@ Record format (one JSON line):
|
|||
"question": "{research question (first 100 chars)}",
|
||||
"mode": "{default|quick}",
|
||||
"scope": "{both|local|external}",
|
||||
"engine": "{effective engine: swarm|deep-research}",
|
||||
"slug": "{brief slug}",
|
||||
"project_dir": "{project_dir or null}",
|
||||
"brief_path": "{brief_destination}",
|
||||
"dimensions": {N},
|
||||
"dimensions_baseline": {N},
|
||||
"dimensions_baseline_preserved": {true|false},
|
||||
"effort": "{low|standard|high}",
|
||||
"conv_turns": {N},
|
||||
"empty_turns": {N},
|
||||
"unique_sources": {N},
|
||||
"agents_local": {N},
|
||||
"agents_external": {N},
|
||||
"gemini_used": {true|false},
|
||||
|
|
@ -783,26 +431,11 @@ Record format (one JSON line):
|
|||
}
|
||||
```
|
||||
|
||||
**The six measurement fields (v5.10).** `effort` is the grouping key — the
|
||||
resolved `phase_signal_result.effort` for the `research` phase, a
|
||||
low-cardinality label (`low|standard|high`), and the only axis on which a
|
||||
high-effort run can be compared against a standard one. Four are numeric:
|
||||
`unique_sources` (distinct sources cited across the brief),
|
||||
`dimensions_baseline` (the interview-derived dimension count, so the Phase 4.5
|
||||
delta against `dimensions` is machine-readable), `conv_turns` (Phase 5 loop
|
||||
turns actually spent), and `empty_turns` (loop turns that returned no findings
|
||||
or no citations). The sixth is boolean: `dimensions_baseline_preserved`, the
|
||||
Phase 4.5 attestation (step 5) that every interview dimension survived onto the
|
||||
final list — the count delta cannot show membership, and the dimension NAMES
|
||||
that could are prose the exporter allowlist denies. On a standard run the loop
|
||||
never arms, so `dimensions_baseline == dimensions`, both turn counters are `0`,
|
||||
and `dimensions_baseline_preserved` is `true` (nothing touched the list).
|
||||
|
||||
If `${CLAUDE_PLUGIN_DATA}` is not set or not writable, skip tracking silently.
|
||||
|
||||
## Profile (v4.1)
|
||||
|
||||
Accepts `--profile <name>` where `<name>` is `economy`, `balanced`, `premium`, `fable`,
|
||||
Accepts `--profile <name>` where `<name>` is `economy`, `balanced`, `premium`,
|
||||
or a custom profile under `voyage-profiles/`. Default: `premium`.
|
||||
|
||||
Resolution order (per `lib/profiles/resolver.mjs`):
|
||||
|
|
@ -822,8 +455,8 @@ VOYAGE_PROFILE=balanced /trekresearch
|
|||
```
|
||||
|
||||
Stats records emit `profile`, `phase_models`, `parallel_agents`,
|
||||
`external_research_enabled`, `profile_source`, and `engine` so operators can
|
||||
audit which profile and engine drove which session.
|
||||
`external_research_enabled`, and `profile_source` so operators can audit
|
||||
which profile drove which session.
|
||||
|
||||
## Composition rule (v5.1)
|
||||
|
||||
|
|
@ -837,15 +470,13 @@ model_for_phase = brief.phase_signals[<phase>]?.model ?? profile.phase_models[
|
|||
```
|
||||
|
||||
The brief signal wins per-phase when present; the profile fills any
|
||||
gaps. Both fields are mechanically resolved by the single composed CLI
|
||||
`node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/resolver.mjs --resolve-phase-model`
|
||||
invoked in Phase 1; the resolved JSON `{effort, model, source}` is captured
|
||||
as `phase_signal_result` and passed to `Agent` tool calls explicitly. The
|
||||
resolver controls the `model` parameter at Agent-spawn sites only — the
|
||||
orchestrator's own model is fixed at invocation time (command frontmatter
|
||||
omits `model:`, so it follows the session model) and cannot be switched
|
||||
mid-turn. Sub-agents fall back to `model:` in their own `agents/*.md`
|
||||
frontmatter when no spawn-site injection happens.
|
||||
gaps. Composition is mechanically resolved via
|
||||
`node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/phase-signal-resolver.mjs`
|
||||
invoked in Phase 1; the resolved JSON is captured as `phase_signal_result`
|
||||
and passed to `Agent` tool calls explicitly. The resolver controls only
|
||||
the orchestrator and the model parameter at Agent-spawn sites — sub-agents
|
||||
otherwise read `model:` from their own `agents/*.md` frontmatter (still
|
||||
pinned to `opus`).
|
||||
|
||||
For `/trekresearch` specifically: `effort == 'low'` activates the
|
||||
existing `--quick`-equivalent code-path (inline research, no agent swarm).
|
||||
|
|
@ -877,14 +508,6 @@ significant architectural questions or when triangulation value is
|
|||
high; in high-effort mode it runs unconditionally to provide an
|
||||
independent second opinion.
|
||||
|
||||
High effort additionally arms the Phase 5 bounded follow-up loop (max 3
|
||||
turns per under-illuminated dimension, budgeted by
|
||||
`research-loop-cap.mjs`, every outbound query gated by
|
||||
`query-privacy-gate.mjs`). The loop stays default-off until
|
||||
`VOYAGE_STORM_ENABLED=1`, and the moot matrix in Phase 5 (`--quick`,
|
||||
`--local`, `external_research_enabled: false`) overrides the effort
|
||||
signal whenever the external phase does not run at all.
|
||||
|
||||
Standard effort (or absent): use the existing conditional triggers.
|
||||
Low effort: inline research only, no agent swarm (existing
|
||||
`--quick`-equivalent code-path).
|
||||
|
|
@ -896,39 +519,12 @@ Low effort: inline research only, no agent swarm (existing
|
|||
- **Sources required:** Every claim must cite a source. No unsourced findings.
|
||||
- **Independence:** Do not pre-bias external agents with local findings or vice versa.
|
||||
Triangulate AFTER independent research.
|
||||
**Amended (v5.10) for Phase 4.5:** dimension discovery deliberately crosses this
|
||||
rule. Its candidate dimensions are mined from the Phase-4 result set, which
|
||||
contains output from the five local codebase agents, so a discovered dimension
|
||||
can carry local context into an external query.
|
||||
|
||||
The crossing creates **two distinct risks**, and they do not share a control:
|
||||
|
||||
- **Bias** — a local finding shapes what an external agent is asked. Its
|
||||
controls are structural, not a gate: the initial external swarm stays blind
|
||||
to local findings, so an **independent baseline already exists** before
|
||||
anything crosses; the crossing is confined to Phase 4.5 and the Phase 5 loop
|
||||
it feeds, which only ADD to that baseline and never revise it; and at
|
||||
`effort: high` — the only effort at which any of this runs —
|
||||
`contrarian-researcher` is forced always-on, so the brief always carries an
|
||||
adversarial counter-evidence pass over the result the crossed queries fed.
|
||||
Triangulation still happens AFTER independent research.
|
||||
- **Egress** — local paths, repo identifiers or secret-shaped strings leave the
|
||||
machine inside a query. That is what `query-privacy-gate.mjs` controls: every
|
||||
outbound query is inspected before it leaves, with a hard-block tier for
|
||||
secret-shaped strings that no operator flag can override.
|
||||
|
||||
The privacy gate was previously named as the compensating control for the
|
||||
crossing as a whole. It is not: it inspects query CONTENT and cannot stop a
|
||||
local finding from steering an external agent's question. Attributing the bias
|
||||
risk to it left that risk with no control while the text read as though it had
|
||||
one.
|
||||
- **Graceful degradation:** If MCP tools are unavailable (Tavily, Gemini, MS Learn),
|
||||
proceed with available tools and note limitations in brief metadata.
|
||||
- **Cost:** Model resolution at Agent-spawn sites is a three-layer fallback:
|
||||
brief `phase_signals[<phase>].model` > `profile.phase_models[<phase>]` >
|
||||
agent frontmatter `model:`. The composed resolver returns the first two
|
||||
layers as `phase_signal_result.model`; spawn sites inject it, and agent
|
||||
frontmatter is the fallback when no injection happens.
|
||||
- **Cost:** Sub-agents use their pinned `model:` frontmatter (currently `opus`).
|
||||
When `phase_signals[<phase>].model` is set, the orchestrator AND Agent-spawn
|
||||
sites use the resolved model (`phase_signal_result.model`) for that phase.
|
||||
Frontmatter is the default; brief signal is the per-phase override.
|
||||
- **Privacy:** Never log secrets, tokens, or credentials.
|
||||
- **Honesty:** If the question is trivially answerable, say so. Don't inflate research.
|
||||
- **Scope of codebase:** Only analyze the current working directory for local research.
|
||||
|
|
|
|||
|
|
@ -5,6 +5,7 @@ description: |
|
|||
review.md with severity-tagged findings (BLOCKER/MAJOR/MINOR/SUGGESTION)
|
||||
per Handover 6 (review → plan).
|
||||
argument-hint: "--project <dir> [--since <ref>] [--quick] [--validate] [--dry-run]"
|
||||
model: opus
|
||||
allowed-tools: Agent, Read, Glob, Grep, Write, Edit, Bash, AskUserQuestion
|
||||
---
|
||||
|
||||
|
|
@ -91,12 +92,10 @@ as the file is parseable:
|
|||
```bash
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/validators/brief-validator.mjs --soft --json "{brief_path}"
|
||||
|
||||
# v5.9 — composed phase-model resolution (brief > profile > default) for the
|
||||
# review phase. ONE call returns {effort, model, source}; captured as
|
||||
# v5.1.1 — resolve the review-phase brief signal. The JSON is captured as
|
||||
# phase_signal_result and used in Phase 7 at the reviewer-launch site to
|
||||
# inject the resolved model. Append --profile {profile} when the operator
|
||||
# passed --profile.
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/resolver.mjs --resolve-phase-model --phase review --brief-path "{brief_path}" [--profile {profile}] --json
|
||||
# inject the brief-resolved model.
|
||||
node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/phase-signal-resolver.mjs --brief "{brief_path}" --phase review --json
|
||||
```
|
||||
|
||||
Read the JSON output. If `valid: false` AND any error has code
|
||||
|
|
@ -421,7 +420,7 @@ the contract for that handover (see `docs/HANDOVER-CONTRACTS.md`).
|
|||
|
||||
## Profile (v4.1)
|
||||
|
||||
Accepts `--profile <name>` where `<name>` is `economy`, `balanced`, `premium`, `fable`,
|
||||
Accepts `--profile <name>` where `<name>` is `economy`, `balanced`, `premium`,
|
||||
or a custom profile under `voyage-profiles/`. Default: `premium`.
|
||||
|
||||
Resolution order (per `lib/profiles/resolver.mjs`):
|
||||
|
|
@ -453,15 +452,13 @@ model_for_phase = brief.phase_signals[<phase>]?.model ?? profile.phase_models[
|
|||
```
|
||||
|
||||
The brief signal wins per-phase when present; the profile fills any
|
||||
gaps. Both fields are mechanically resolved by the single composed CLI
|
||||
`node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/resolver.mjs --resolve-phase-model`
|
||||
invoked in Phase 2; the resolved JSON `{effort, model, source}` is captured
|
||||
as `phase_signal_result` and passed to `Agent` tool calls explicitly. The
|
||||
resolver controls the `model` parameter at Agent-spawn sites only — the
|
||||
orchestrator's own model is fixed at invocation time (command frontmatter
|
||||
omits `model:`, so it follows the session model) and cannot be switched
|
||||
mid-turn. Sub-agents fall back to `model:` in their own `agents/*.md`
|
||||
frontmatter when no spawn-site injection happens.
|
||||
gaps. Composition is mechanically resolved via
|
||||
`node ${CLAUDE_PLUGIN_ROOT}/lib/profiles/phase-signal-resolver.mjs`
|
||||
invoked in Phase 2; the resolved JSON is captured as `phase_signal_result`
|
||||
and passed to `Agent` tool calls explicitly. The resolver controls only
|
||||
the orchestrator and the model parameter at Agent-spawn sites — sub-agents
|
||||
otherwise read `model:` from their own `agents/*.md` frontmatter (still
|
||||
pinned to `opus`).
|
||||
|
||||
For `/trekreview` specifically: `effort == 'low'` activates the existing
|
||||
`--quick`-equivalent code-path (skip the brief-conformance reviewer; run
|
||||
|
|
@ -518,11 +515,10 @@ Low effort: skip the brief-conformance reviewer entirely (existing
|
|||
`findings:\n - a\n - b`.
|
||||
- **Refuse-with-suggestion above 100 files / 100K tokens.** Never run
|
||||
blind on a giant diff. Use AskUserQuestion to surface the gate.
|
||||
- **Cost.** Model resolution at Agent-spawn sites is a three-layer fallback:
|
||||
brief `phase_signals[<phase>].model` > `profile.phase_models[<phase>]` >
|
||||
agent frontmatter `model:`. The composed resolver returns the first two
|
||||
layers as `phase_signal_result.model`; spawn sites inject it, and agent
|
||||
frontmatter is the fallback when no injection happens.
|
||||
- **Cost.** Sub-agents use their pinned `model:` frontmatter (currently `opus`).
|
||||
When `phase_signals[<phase>].model` is set, the orchestrator AND Agent-spawn
|
||||
sites use the resolved model (`phase_signal_result.model`) for that phase.
|
||||
Frontmatter is the default; brief signal is the per-phase override.
|
||||
- **Privacy.** Never log secrets, tokens, or credentials in review.md.
|
||||
Findings citing files with secret-like content must redact the secret
|
||||
in the `detail` field.
|
||||
|
|
|
|||
|
|
@ -45,8 +45,6 @@ Every validator exposes a CLI: `node lib/validators/<name>.mjs --json <path>` re
|
|||
|
||||
**Stability tier: PUBLIC CONTRACT.** Handover 1 is the *only* public integration boundary of the pipeline: `brief.md` is what an upstream producer hands to Voyage, and Voyage consumes it without any knowledge of who produced it. This asymmetry is a hard invariant (see `CLAUDE.md` §Trinity context) — the interactive `/trekbrief` interview is just one producer; a `manual` brief, or an external per-app / per-portfolio producer, is equally valid as long as the artifact conforms to the schema below. No producer is privileged. The consequence: changing this schema — renaming or removing a field, narrowing an enum, or promoting an optional field to required — is a **breaking change for every downstream consumer** and MUST follow the [breaking-change protocol](#breaking-change-protocol) (version bump + N-1 compatibility window). Additive *optional* fields are non-breaking by design: the validator tolerates unknown frontmatter keys silently (forward-compat), so a newer brief still validates against an older consumer. The v5.5.0 contract formalization established `brief_version` **2.1** as the public-contract baseline and, in the same coordinated release, **evolved it to `2.2`** under the breaking-change protocol — adding the required `framing` field + `## TL;DR` section gated at ≥ 2.2 (existing 2.0/2.1 briefs stay valid). See the Versioning paragraph below for the 2.2 details.
|
||||
|
||||
**Trinity producer context (informational).** In the author's private three-tier design, Voyage is **Tier 1** (per-task). Its upstream producers are **Tier 2 `app-creator`** (per-app — "what does the app need, what's the next brief?") and **Tier 3 `app-factory`** (per-portfolio — "which app needs me now?"), both pre-implementation and bound for Forgejo when ready. None is privileged: any schema-conforming producer is equally valid, and Voyage stays unaware of Tier 2/3 — this Handover is the only coupling.
|
||||
|
||||
**Producer:** `/trekbrief` Phase 4g (after `brief-reviewer` stop-gate passes or iteration cap is hit).
|
||||
|
||||
**Consumer:** `/trekresearch` Phase 1 (mode parse + brief validation).
|
||||
|
|
@ -115,17 +113,9 @@ Optional but standard sections: `## Non-Goals`, `## Constraints`, `## Preference
|
|||
- `BRIEF_INVALID_PHASE_SIGNALS` → strict halt; phase_signals must be a list of `{phase, effort?, model?}` entries.
|
||||
- `BRIEF_INVALID_PHASE_SIGNAL_PHASE` → strict halt; phase ∉ `[research, plan, execute, review]`.
|
||||
- `BRIEF_INVALID_EFFORT` → strict halt; effort ∉ `[low, standard, high]`.
|
||||
- `BRIEF_INVALID_MODEL` → strict halt; model ∉ `BASE_ALLOWED_MODELS` (currently `[sonnet, opus, fable]`).
|
||||
- `BRIEF_INVALID_MODEL` → strict halt; model ∉ `BASE_ALLOWED_MODELS` (currently `[sonnet, opus]`).
|
||||
- `BRIEF_SIGNALS_MUTUALLY_EXCLUSIVE` → strict halt; cannot set both `phase_signals` and `phase_signals_partial: true`.
|
||||
|
||||
**Compatibility direction of the v5.9 allowlist widening (`fable`):** enum
|
||||
widening is safe for new readers of old data, not old readers of new data.
|
||||
Existing sonnet/opus briefs stay valid under the v5.9+ validator (non-breaking,
|
||||
no `brief_version` bump — value-space extension, not a schema change). The
|
||||
reverse does NOT hold: a fable-bearing brief REQUIRES a v5.9+ validator — an
|
||||
older cached brief-validator (e.g. a stale v5.8 marketplace clone) rejects it
|
||||
with `BRIEF_INVALID_MODEL`.
|
||||
|
||||
---
|
||||
|
||||
## Handover 2 — research/*.md → plan
|
||||
|
|
|
|||
BIN
docs/SLDC-AI.pdf
BIN
docs/SLDC-AI.pdf
Binary file not shown.
|
|
@ -1,153 +0,0 @@
|
|||
---
|
||||
type: trekbrief
|
||||
brief_version: "2.2"
|
||||
created: 2026-06-26
|
||||
task: "Relocate <example> blocks out of voyage agent description frontmatter to cut always-loaded tokens"
|
||||
slug: agent-description-token-trim
|
||||
project_dir: .claude/projects/2026-06-26-agent-description-token-trim/
|
||||
research_topics: 0
|
||||
research_status: skipped
|
||||
auto_research: false
|
||||
interview_turns: 0
|
||||
source: manual
|
||||
framing: refine
|
||||
phase_signals:
|
||||
- phase: research
|
||||
effort: low
|
||||
- phase: plan
|
||||
effort: standard
|
||||
- phase: execute
|
||||
effort: standard
|
||||
- phase: review
|
||||
effort: standard
|
||||
---
|
||||
|
||||
# Task: Trim voyage agent `description:` frontmatter — relocate `<example>` blocks into the agent body
|
||||
|
||||
> **Cross-session coordination brief.** Authored 2026-06-26 by the **config-audit machine-tuning session** (a parallel session that audits this machine's Claude Code always-loaded token footprint). It is dropped here per operator instruction so the **voyage session** owns the implementation — config-audit does **not** edit the voyage repo. The source measurement lives in config-audit's local worklist (`config-audit/docs/machine-tuning-worklist.local.md`, item **M4**, session m). This brief is the contract; `/trekplan` can consume it directly.
|
||||
|
||||
## TL;DR
|
||||
|
||||
voyage's 24 agent `description:` fields cost **~4,418 always-loaded tokens every turn** (injected into the system prompt of every session that has voyage enabled). **17** of those agents carry **two `<example>` blocks each (34 total, ~3,147 tok)** in their frontmatter `description`. Those example blocks exist to drive *autonomous* agent-selection — but voyage agents are launched **by explicit name** from the orchestrator commands, so the examples are cost without function. Move them into each agent's body (preserve, don't delete). Expected saving: **~2,500–3,100 always-loaded tokens.** framing: **refine** — this continues the token-trim line started in M1a (v5.6.1).
|
||||
|
||||
## Intent
|
||||
|
||||
The machine's #1 tuning lever is *always-loaded* tokens — context injected on every turn, paid on every request. config-audit's Fase-1 measurement (2026-06-26) found the machine's global layer is already lean (CLAUDE.md trimmed, rules/agents empty, output style off, MCP deferred), and the **single largest remaining tunable source is voyage's agent listing (~4,418 tok)**. The bulk of that is `<example>` blocks in the spawnable agents' `description:` frontmatter. These blocks follow the documented agent-authoring pattern whose purpose is to help the *main loop autonomously decide* when to delegate to an agent. voyage doesn't rely on that path: `/trekplan`, `/trekresearch`, and `/trekreview` launch their agents **deterministically by name** ("Launch the **architecture-mapper** agent", explicit per-codebase-size launch tables, explicit `voyage:review-coordinator` references). So the examples are paying a per-turn token tax for an auto-selection behavior voyage never uses.
|
||||
|
||||
## Goal
|
||||
|
||||
Each of the **17** example-bearing voyage agents has its frontmatter `description:` reduced to its **triggering lead sentence(s)** (the "Use this agent to/when…" summary — what any residual autonomous selection actually needs), with the `<example>` blocks **relocated verbatim into the agent's markdown body** under a clearly-marked section (e.g. `## When to use — examples`). No agent is deleted, no behavior changes, no example content is lost. The injected agent-listing cost drops by ~2,500–3,100 tokens, landing on every machine that reloads voyage after release.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- **Do NOT delete the examples** — relocate them into the body. They remain useful documentation and `/trekbrief`/onboarding reference.
|
||||
- **Do NOT change any agent's system prompt, body logic, tools, model, or `name`** (`name` is the agent's identity).
|
||||
- **Do NOT touch the 7 zero-example agents** (`planning-orchestrator`, `research-orchestrator`, `review-orchestrator`, `synthesis-agent`, `review-coordinator`, `code-correctness-reviewer`, `brief-conformance-reviewer`) — already lean (M1a handled the reference/dormant ones; the orchestrators keep their pinned "reference document, not a spawnable capability" phrasing).
|
||||
- **Do NOT modify the orchestrator commands** (`/trekplan` etc.) or change which agents exist.
|
||||
- This is **not** a behavioral or capability change — purely a token/packaging optimization.
|
||||
|
||||
## Constraints
|
||||
|
||||
- Frontmatter must remain valid YAML. The retained `description:` must keep a meaningful **triggering lead sentence**, because a few agents (notably the `*-researcher` agents) *can* legitimately be selected autonomously when a user asks a research question directly — keep their first 1–2 sentences descriptive enough to still trigger.
|
||||
- **The test suite must stay green.** `node --test 'tests/**/*.test.mjs'` is the gate. voyage has inventory/frontmatter-pin tests — M1a's note warned that "synthesis-schema + inventory tests may pin description content." If any test asserts on `<example>` presence/count or description length, update that test's expectation **deliberately** (it is asserting the old structure) and document why.
|
||||
- Follow voyage's release discipline: version-sync (plugin.json + package.json + README badge/What's-new + CHANGELOG), annotated tag, and catalog `ref` bump with `check-versions` green. This is a perf/token change → a **patch** release (e.g. v5.6.2) or fold into the work already in progress on the branch.
|
||||
- The saving only reaches a machine after the operator reloads voyage (`/plugin marketplace update` + update + `/exit`) — note this in the CHANGELOG so the delta isn't expected instantly.
|
||||
|
||||
## Preferences
|
||||
|
||||
- Relocate the two examples per agent into a single body section with a stable heading so they're easy to find and so any future inventory test can pin the body instead of the frontmatter.
|
||||
- Keep the diff mechanical and uniform across the 17 agents (same heading, same ordering) for reviewability.
|
||||
- If a lead sentence is currently entangled with the first `<example>`, lightly rewrite it into a clean 1–2 sentence trigger — minimal, not a rewrite.
|
||||
|
||||
## Non-Functional Requirements
|
||||
|
||||
- **Zero new dependencies.**
|
||||
- **Always-loaded `description` budget:** total frontmatter `description` chars across all 24 agents drops from **~17,672** to **≤ ~6,000** (~2,500–3,100 tok saved).
|
||||
- No regression in agent auto-selection for the agents that are legitimately autonomous (the researchers) — spot-check that a direct research-style prompt still routes sensibly.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- **Full suite green:** `node --test 'tests/**/*.test.mjs'` exits 0 (same pass count, or deliberately updated pins with rationale).
|
||||
- **No `<example>` in any frontmatter `description`:** the frontmatter-only scan below prints `0`:
|
||||
```bash
|
||||
python3 - <<'PY'
|
||||
import re,glob
|
||||
n=0
|
||||
for f in glob.glob('agents/*.md'):
|
||||
t=open(f).read(); m=re.match(r'^---\n(.*?)\n---', t, re.S)
|
||||
fm=m.group(1) if m else ''
|
||||
dm=re.search(r'description:\s*(.*?)(?=\n[a-zA-Z_][a-zA-Z_-]*:\s|\Z)', fm, re.S)
|
||||
n+=(dm.group(1).count('<example>') if dm else 0)
|
||||
print(n) # must be 0
|
||||
PY
|
||||
```
|
||||
- **Examples preserved in bodies:** `grep -l '<example>' agents/*.md` still lists the 17 agents (the blocks moved, not vanished).
|
||||
- **Agent count unchanged:** 24 agents, each with non-empty `name` + `description`. Plugin validates (plugin-validator / inventory tests green).
|
||||
- **Token budget met:** total frontmatter `description` chars ≤ ~6,000 (measure with the per-agent script in the worklist / the snippet above adapted to sum lengths).
|
||||
- **Behavior unchanged:** `/trekplan` (and trekresearch/trekreview) still launch the same agents by name — diff touches only `agents/*.md`, no command files.
|
||||
|
||||
## Research Plan
|
||||
|
||||
No external research needed — the codebase and this brief contain sufficient context for planning. (research_topics = 0.)
|
||||
|
||||
## Open Questions / Assumptions
|
||||
|
||||
- **[ASSUMPTION]** The 17 example-bearing agents are invoked **by name** from orchestrator commands — verified by the config-audit session against `voyage/commands/*.md` (explicit "Launch the X agent" prose + per-size launch tables + explicit `voyage:<name>` references). The `<example>` auto-trigger blocks are therefore non-load-bearing for voyage's actual invocation path.
|
||||
- **[Q]** Do any voyage tests assert on `<example>` presence/count or `description` length in `agents/*.md`? **Check first** (`grep -rn 'example\|description' tests/`), so a pin is updated intentionally, not discovered as a red test.
|
||||
- **[ASSUMPTION]** The `*-researcher` agents (community/contrarian/docs/security) and `gemini-bridge` may also be autonomously selected outside the trek pipeline → keep their lead sentences trigger-worthy (don't reduce to a bare label).
|
||||
|
||||
## Prior Attempts
|
||||
|
||||
- **M1a — voyage v5.6.1 (2026-06-24):** trimmed the 4 non-spawnable reference/dormant agents' (`planning/research/review-orchestrator` + `synthesis-agent`) verbose `description:` fields to one-liners; orchestrators kept their pinned phrase, synthesis kept its DORMANT flag. Suite stayed green (756). **M4 extends that same always-loaded-token-trim to the 17 spawnable agents' `<example>` blocks** — the larger remaining chunk M1a deliberately left.
|
||||
|
||||
## Reference data (config-audit Fase-1 measurement, 2026-06-26)
|
||||
|
||||
24 agents, **17,672 `description` chars (~4,418 tok)**, **34 `<example>` blocks**, ~12,591 chars (~3,147 tok) inside `<example>…</example>`. The 17 example-bearing agents (descending `description` size):
|
||||
|
||||
| agent | desc chars | examples |
|
||||
|---|---:|---:|
|
||||
| community-researcher | 1280 | 2 |
|
||||
| contrarian-researcher | 1270 | 2 |
|
||||
| gemini-bridge | 1226 | 2 |
|
||||
| security-researcher | 1207 | 2 |
|
||||
| docs-researcher | 1138 | 2 |
|
||||
| session-decomposer | 946 | 2 |
|
||||
| convention-scanner | 939 | 2 |
|
||||
| brief-reviewer | 870 | 2 |
|
||||
| research-scout | 847 | 2 |
|
||||
| test-strategist | 823 | 2 |
|
||||
| dependency-tracer | 818 | 2 |
|
||||
| task-finder | 817 | 2 |
|
||||
| git-historian | 795 | 2 |
|
||||
| risk-assessor | 794 | 2 |
|
||||
| architecture-mapper | 789 | 2 |
|
||||
| scope-guardian | 750 | 2 |
|
||||
| plan-critic | 692 | 2 |
|
||||
|
||||
Zero-example (leave as-is): review-coordinator (345), code-correctness-reviewer (292), brief-conformance-reviewer (266), synthesis-agent (228), planning-orchestrator (184), research-orchestrator (179), review-orchestrator (177).
|
||||
|
||||
## Metadata
|
||||
|
||||
- **Created:** 2026-06-26
|
||||
- **Interview turns:** 0
|
||||
- **Auto-research opted in:** no
|
||||
- **Source:** manual (cross-session coordination brief from the config-audit machine-tuning session)
|
||||
|
||||
---
|
||||
|
||||
## How to continue
|
||||
|
||||
```bash
|
||||
# (optional) confirm no test pins <example> in descriptions first:
|
||||
grep -rn 'example' tests/ | grep -i descr
|
||||
|
||||
# plan + execute via the voyage pipeline:
|
||||
/trekplan --project .claude/projects/2026-06-26-agent-description-token-trim/
|
||||
/trekexecute --project .claude/projects/2026-06-26-agent-description-token-trim/
|
||||
|
||||
# or implement directly (it's a uniform, mechanical 17-file edit) and release as a patch (v5.6.2):
|
||||
# - move each agent's 2 <example> blocks from frontmatter description into a body
|
||||
# "## When to use — examples" section
|
||||
# - node --test 'tests/**/*.test.mjs' (green)
|
||||
# - version-sync + CHANGELOG + tag + catalog ref-bump (check-versions green)
|
||||
```
|
||||
|
|
@ -12,8 +12,6 @@ Imported from `CLAUDE.md` via pointer.
|
|||
- `lib/stats/event-emit.mjs` — single-source stats event emitter for autonomy-gate transitions and main-merge-gate (v3.4.0)
|
||||
- `lib/validators/{brief,research,plan,progress,session-state}-validator.mjs` — schema validators with CLI shims (`node lib/validators/X.mjs --json <path>`)
|
||||
- `lib/validators/architecture-discovery.mjs` — drift-WARN external-contract discovery for `architecture/overview.md`
|
||||
- `lib/util/research-loop-cap.mjs` — stateful, **default-off** turn budget for the `/trekresearch` bounded conversation loop. `allowTurn()` derives the used-turn count from its own append-only JSONL ledger; it never asks the caller how many turns it has spent, because a cap that does is not a cap. Each grant first claims a turn **slot** with `O_EXCL` under `trekresearch-loop-claims/`, so the bound survives several callers deciding at once — counting the ledger and then appending is read-then-write, and Phase 4.5/5 can spawn several agents in one message. Budget = `TREKRESEARCH_MAX_CONV_TURNS` (default `3`, invalid values fall back to `3`) × `maxDimensions` (8, `settings.json:16`). **Default-off:** grants 0 unless `VOYAGE_STORM_ENABLED=1`, which is also the second condition on Phase 4.5's skip-guard (that phase does not call this module): unset, **both** STORM phases are inert. `resolveDataRoot()` is the single root for everything the loop writes — `CLAUDE_PLUGIN_DATA` when the harness sets it, `~/.claude/voyage` when it does not (it is empty in the Bash tool's process env, which is where the loop actually runs); the cap hook resolves through the same function, so writer and reader cannot disagree. A ledger that cannot be **written** denies the turn, and one that exists but cannot be **read** denies it too — only `ENOENT` counts as zero turns spent, that being the legitimate first-turn state (fail-closed — the opposite of `event-emit.mjs`, which is telemetry and must never block). The exported `readLedger()` is the single counting rule; the cap hook calls it rather than keeping a private copy. `checkDimensionCeiling()` is the reader for the OTHER axis of the bounded-cost NFR — the size of the whole dimension list after Phase 4.5 discovery — against the same `MAX_TOTAL_DIMENSIONS`, so the two axes cannot enforce different numbers for one `settings.json:16` value; an unreadable count is rejected, not waved through. CLI shim, two modes: `--run-id ID --dimension D --effort E` (budget gate) and `--check-dimensions N` (ceiling, no run id/effort/flag required since Phase 4.5 never calls the budget gate)
|
||||
- `lib/validators/query-privacy-gate.mjs` — gates **every** outbound research query before it leaves the machine; the hard-block tier (secret-shaped strings) is not operator-overridable, so a query that trips it must be reformulated rather than forced through. CLI shim: `node lib/validators/query-privacy-gate.mjs "<query>"`
|
||||
|
||||
Wiring points (replaces previous prose-grep instructions):
|
||||
- `/trekbrief` Phase 4g → `brief-validator` (post-write sanity check)
|
||||
|
|
@ -33,15 +31,13 @@ Doc-consistency test at `tests/lib/doc-consistency.test.mjs` pins agent-table co
|
|||
|
||||
`hooks/scripts/post-bash-stats.mjs` (PostToolUse, CC v2.1.97+) appends `duration_ms` for each Bash call into `${CLAUDE_PLUGIN_DATA}/trekexecute-stats.jsonl`. Useful for finding long-running verify or checkpoint commands.
|
||||
|
||||
`hooks/scripts/pre-agent-cap.mjs` (PreToolUse on `WebSearch|WebFetch|Task`) enforces the `/trekresearch` Phase 5 loop bound in the harness, so the cap is not merely prose the model is asked to obey. It counts spent turns read-only from the append-only ledger `research-loop-cap.mjs` writes, through that module's own exported `readLedger()`. It denies (exit 2) once the budget **gate has denied a turn** — the primitive records its denials as tombstones, and the tombstone is the boundary rather than the count, because `allowTurn()` appends before the turn runs and so the final granted turn already shows `budget` records. A ledger showing more granted turns than the budget denies too, as a backstop. Being in scope but unable to count the ledger also denies: a budget control that cannot count must not grant. Scope key = `session_id` + a marker file only the loop writes (`<resolveDataRoot()>/trekresearch-loop-scope/<session_id>.json`, the same root the ledger uses); without a marker the hook allows unconditionally, which is what keeps a globally-wired `PreToolUse` hook from over-blocking unrelated sessions. Stale markers auto-reset on a TTL and `VOYAGE_DISABLE_CAP_HOOK=1` is the kill switch. Defence in depth only — `lib/util/research-loop-cap.mjs` must stay correct if the hook stops firing (see `docs/spike-pretooluse-subagent-reach.md`).
|
||||
|
||||
`hooks/scripts/post-compact-flush.mjs` (PostCompact event, v3.4.0) re-injects `.session-state.local.json` after context compaction so multi-session work survives a compaction boundary. Companion to `pre-compact-flush.mjs` (which writes the state file before compaction); together they form the rehydrate cycle that keeps `/trekcontinue` reliable across long-running multi-session work.
|
||||
|
||||
## Architecture
|
||||
|
||||
**Brief:** 7-phase workflow: Parse mode → Create project dir → Phase 3 completeness loop (section-driven, no question cap) → Phase 3.5 per-phase effort dialog (v5.1) → Phase 4 draft/review/revise with `brief-reviewer` as stop-gate (max 3 iterations; gate = all dimensions ≥ 4 and research plan = 5) → Finalize (`brief.md` on pass, or `brief_quality: partial` on cap/force-stop) → Manual/auto opt-in → Stats. Always interactive. Auto mode runs research + plan inline in the main context (v2.4.0).
|
||||
|
||||
**Phase 3.5 (v5.1) — adaptive-depth signals:** Between Phase 3 completeness exit and Phase 4 draft, the operator commits an effort level (`low | standard | high`) and an optional `model` (`sonnet | opus | fable`) per downstream phase (`research`, `plan`, `execute`, `review`) via 4 tier-coupled `AskUserQuestion` calls. The choices land in `brief.md` frontmatter as `phase_signals:` (a list of `{phase, effort?, model?}` entries) when committed, or `phase_signals_partial: true` when the operator force-stops. `brief_version: 2.1` activates the **sequencing gate**: validator emits `BRIEF_V51_MISSING_SIGNALS` if a 2.1-versioned brief lacks both fields. Downstream commands surface a friendly hint pointing back to `/trekbrief` — enforcement is validator-only. Composition is documented prose in each downstream command's `## Composition rule (v5.1)` section: `brief.phase_signals[phase] > profile.phase_models[phase]`. The brief signal wins per-phase when present; the profile fills gaps. `effort == low` activates each command's existing `--quick`-equivalent code-path (`/trekexecute` low-effort = `--gates open` + sequential-only). High-effort behavior is deferred to v5.1.1 per brief Non-Goal.
|
||||
**Phase 3.5 (v5.1) — adaptive-depth signals:** Between Phase 3 completeness exit and Phase 4 draft, the operator commits an effort level (`low | standard | high`) and an optional `model` (`sonnet | opus`) per downstream phase (`research`, `plan`, `execute`, `review`) via 4 tier-coupled `AskUserQuestion` calls. The choices land in `brief.md` frontmatter as `phase_signals:` (a list of `{phase, effort?, model?}` entries) when committed, or `phase_signals_partial: true` when the operator force-stops. `brief_version: 2.1` activates the **sequencing gate**: validator emits `BRIEF_V51_MISSING_SIGNALS` if a 2.1-versioned brief lacks both fields. Downstream commands surface a friendly hint pointing back to `/trekbrief` — enforcement is validator-only. Composition is documented prose in each downstream command's `## Composition rule (v5.1)` section: `brief.phase_signals[phase] > profile.phase_models[phase]`. The brief signal wins per-phase when present; the profile fills gaps. `effort == low` activates each command's existing `--quick`-equivalent code-path (`/trekexecute` low-effort = `--gates open` + sequential-only). High-effort behavior is deferred to v5.1.1 per brief Non-Goal.
|
||||
|
||||
**Research:** Foreground workflow (v2.4.0): Parse mode → Interview → Parallel research swarm (5 local + 4 external + 1 bridge, spawned from main context) → Follow-ups → Triangulation → Synthesis + brief → Stats. With `--project`, writes to `{dir}/research/NN-slug.md`.
|
||||
|
||||
|
|
@ -96,7 +92,7 @@ Which native Claude Code primitive each pipeline step runs on today, and the alt
|
|||
| **execute** | Inline step loop; multi-session via `git worktree` + `claude -p` waves; deterministic manifest audit | CC `TaskCreate`/`TodoWrite` for progress/resume (insufficient — carries no step status / attempts / SHA / drift → own typed `progress.json` contract) |
|
||||
| **review** | Inline parallel reviewers (no cross-feed) → `review-coordinator` Judge | **Workflow** substrate for Phase 5–6 (bake-off POSITIVE: +4.4 % tokens / +54 % wall-time → shipped **opt-in `--workflow`**, not default; wholesale substrate swap declined) |
|
||||
| **continue** | Inline reads `.session-state.local.json` → zero-confirm resume | CC `--resume` (transcript replay, not typed work-state → insufficient) |
|
||||
| **cross-cutting** | 8 hook scripts: `pre-bash` + `pre-write` guards, `pre-agent-cap` loop-bound enforcement, `post-bash` stats, `session-title`, `pre-`/`post-compact` flush, **`Stop`→OTEL** export | — |
|
||||
| **cross-cutting** | 7 hook scripts: `pre-bash` + `pre-write` guards, `post-bash` stats, `session-title`, `pre-`/`post-compact` flush, **`Stop`→OTEL** export | — |
|
||||
|
||||
¹ MCP per research agent: `docs-researcher` → Microsoft Learn + Tavily · `community-`/`security-`/`contrarian-researcher` → Tavily (+ WebSearch/WebFetch) · `gemini-bridge` → Gemini Deep Research MCP. Graceful degradation when an MCP server is absent.
|
||||
|
||||
|
|
|
|||
|
|
@ -1,40 +0,0 @@
|
|||
# Brief — trim voyage CLAUDE.md to invariants (always-loaded token reduction)
|
||||
|
||||
**Author:** config-audit machine-tuning loop (KTG). **Date:** 2026-06-29.
|
||||
**Why config-audit didn't do this directly:** voyage is the operator's plugin — config-audit only writes briefs here, never code. Run this in a voyage session.
|
||||
|
||||
## Goal
|
||||
|
||||
`CLAUDE.md` is injected **every turn** while working in the voyage repo (the whole per-repo always-loaded delta). Measured with config-audit `manifest`: **2,261 always-loaded tokens**. The tables (commands, agents) are invariant; the cost sits in six design-history / context **block-quote notes** whose detail already lives in referenced docs. Moving them out (leaving a terse invariant + pointer) should cut **~900–1,200 tok (~40–50 %)** with zero capability loss.
|
||||
|
||||
## Hard constraints (do NOT break these)
|
||||
|
||||
- `tests/lib/doc-consistency.test.mjs` reads CLAUDE.md. Keep the **Commands table** and the **Agents table including the `Model` column** and any **count** the test cross-checks. After editing, `node --test 'tests/**/*.test.mjs'` and `bash verify.sh` must stay green.
|
||||
- Keep every **invariant fact** — just relocate the verbose *rationale/history* to the doc that already owns it, and leave a one-line pointer. Don't delete facts; move them.
|
||||
- Docs-only change → no version bump / no catalog ref unless the doc-consistency test demands a badge sync (it shouldn't for a CLAUDE.md prose trim).
|
||||
|
||||
## Trim targets (line numbers as of commit at brief time; match by content)
|
||||
|
||||
| Block | ~chars / ~tok | Action | Destination (already referenced) |
|
||||
|-------|---------------|--------|----------------------------------|
|
||||
| L5 design-principle parenthetical (synthesis-PoC Δ≈0 caveat) | 620 / ~155 | **Keep the honesty caveat in one short clause** ("main-context relief is asserted-by-design, not measured — see doc"), move the PoC explanation out | `docs/T1-synthesis-poc-results.md` |
|
||||
| L7 v3.0.0 architect-extraction note | 376 / ~94 | Replace with a one-liner: the plan command auto-discovers `architecture/overview.md` if present (migration history → CHANGELOG) | `CHANGELOG.md` |
|
||||
| L9 Trinity context (Tier 1/2/3) | 859 / ~215 | Keep the **asymmetry invariant** terse ("Voyage stays unaware of Tier 2/3; Handover 1 = the only integration point; brief-schema changes are breaking"), move the Tier 2/3 description out | `docs/HANDOVER-CONTRACTS.md` §Handover 1 |
|
||||
| **L11 brief-framing 3-layer gate** | **1,585 / ~396** | Biggest. Keep one invariant line ("brief framing must match operator intent — enforced as the `brief_version 2.2` gate: `framing` frontmatter + memory-alignment dim 6 + `## TL;DR`; all BLOCKER for ≥2.2"), move the implementation detail out | `docs/HANDOVER-CONTRACTS.md` §Handover 1 (PUBLIC CONTRACT) |
|
||||
| L56 S33 agent-inventory reconcile | 740 / ~185 | Keep the terse fact ("24 agent files = 21 spawnable + 3 orchestrator reference docs; `synthesis-agent` dormant; all `opus`"), move the sonnet-downgrade rationale out | `docs/voyage-vs-cc-balance-analysis.md` §10 |
|
||||
| L58 Model & effort axes | 588 / ~147 | Keep one line ("`opus`=Opus 4.8 / `high`; select agents carry native `effort:`; distinct from brief `phase_signals.effort`"), move the per-agent effort table + axes explanation out | `docs/profiles.md` §Model & effort axes |
|
||||
|
||||
**Net:** ~2,261 → ~1,300 tok. (Verify the after-figure with `node /path/to/config-audit/scanners/manifest.mjs <voyage-path>` → the `project` `claude-md` source.)
|
||||
|
||||
## Procedure
|
||||
|
||||
1. For each block above, confirm the destination doc already contains the detail (it's referenced, so it should) — if not, move the prose there first.
|
||||
2. Replace each block-quote with its terse invariant + pointer.
|
||||
3. Leave Commands/Agents tables byte-exact (counts + model column).
|
||||
4. `node --test 'tests/**/*.test.mjs'` + `bash verify.sh` green; confirm `doc-consistency` passes.
|
||||
5. Re-measure; commit `docs(claude-md): trim CLAUDE.md to invariants` to voyage's Forgejo. Δ lands on the machine after operator `/exit` + reload.
|
||||
|
||||
## Notes
|
||||
|
||||
- This mirrors the same trim already done directly in config-audit (−662 tok), linkedin-studio (−2,266), ms-ai-architect (−346), okr (−15), portfolio-optimiser (−93) — same principle: CLAUDE.md carries invariants for working on the plugin; history/rationale lives in CHANGELOG + docs/.
|
||||
- The voyage honesty culture (no overclaiming) is *preserved*, not trimmed: the synthesis-PoC Δ≈0 caveat stays as a one-liner; only the long-form explanation moves to its doc.
|
||||
|
|
@ -9,7 +9,7 @@ Per-command flag tables, imported from `CLAUDE.md` via pointer.
|
|||
| _(default)_ | Dynamic interview until quality gates pass → brief.md with research plan |
|
||||
| `--quick` | Compact start; still escalates if required sections are weak or the brief-review gate fails → brief.md with research plan |
|
||||
| `--gates {true\|false}` | (v3.4.0) Boolean autonomy-gate flag; present → gating on. Policy (`gates_mode`) detailed under `## Autonomy mode` in `docs/operations.md`. |
|
||||
| `--profile <name>` | (v4.1.0) Model profile: `economy` / `balanced` / `premium` / `fable` / `<custom>`. Sets `phase_models` for the brief phase. See `## Profile system` in `docs/operations.md`. |
|
||||
| `--profile <name>` | (v4.1.0) Model profile: `economy` / `balanced` / `premium` / `<custom>`. Sets `phase_models` for the brief phase. See `## Profile system` in `docs/operations.md`. |
|
||||
|
||||
Always interactive. Phase 3 is a section-driven completeness loop (no hard cap on question count); Phase 4 runs a `brief-reviewer` stop-gate with max 3 review iterations. After writing the brief, asks the user to choose manual (print commands) or auto (Claude runs research + plan in foreground).
|
||||
|
||||
|
|
@ -26,26 +26,9 @@ Always interactive. Phase 3 is a section-driven completeness loop (no hard cap o
|
|||
| `--gates {true\|false}` | (v3.4.0) Boolean autonomy-gate flag; present → gating on. Policy (`gates_mode`) detailed under `## Autonomy mode` in `docs/operations.md`. |
|
||||
| `--min-brief-version <ver>` | (S18) Warn — never block — if an attached `--project` brief declares a version below `<ver>` (e.g. `2.2`), i.e. sidesteps framing enforcement |
|
||||
| `--profile <name>` | (v4.1.0) Model profile for the research phase. |
|
||||
| `--engine {swarm\|deep-research}` | (deep-research-engine) Opt-in external-research engine; `deep-research` delegates the external phase to Claude Code's built-in `/deep-research` workflow (CC 2.1.154+), falls back to `swarm`. Default `swarm`. |
|
||||
|
||||
Flags combine: `--project <dir> --local`, `--external --quick`.
|
||||
|
||||
### Bounded conversation loop (Phase 4.5 + Phase 5) — env-vars
|
||||
|
||||
Dimension discovery and the multi-turn follow-up loop are **default-off** and
|
||||
have no flag; they are environment-gated, because turning them on costs turns.
|
||||
They run only at `effort: high` (resolved from the brief's `phase_signals`).
|
||||
|
||||
| Env-var | Default | Behavior |
|
||||
|---------|---------|----------|
|
||||
| `VOYAGE_STORM_ENABLED` | _(unset — default-off)_ | `=1` grants the Phase 5 loop a non-zero turn budget and is the second condition on Phase 4.5's skip-guard. Unset, `research-loop-cap.mjs` grants 0 turns and **both** phases are inert: doing nothing keeps the whole mechanism off. |
|
||||
| `TREKRESEARCH_MAX_CONV_TURNS` | `3` | Max turns per under-illuminated dimension. Budget = this × `maxDimensions` (8, `settings.json:16`). Empty, non-numeric, zero, negative, `Infinity`, or any fraction that floors below `1` (`0.5`, `0.9`) fall back to `3` — never to unbounded, and never to `0`. A fraction at or above `1` floors (`2.7` → `2`). |
|
||||
| `VOYAGE_DISABLE_CAP_HOOK` | _(unset)_ | `=1` disables `hooks/scripts/pre-agent-cap.mjs`, the `PreToolUse` enforcement of the turn budget. The cap primitive still applies; only the second gate is switched off. |
|
||||
|
||||
Adoption of the loop as a default is gated on a pre-registered measurement —
|
||||
protocol, thresholds, and the exact commands in
|
||||
[`docs/storm-measurement.md`](storm-measurement.md).
|
||||
|
||||
## /trekplan modes
|
||||
|
||||
| Flag | Behavior |
|
||||
|
|
|
|||
|
|
@ -1,62 +0,0 @@
|
|||
---
|
||||
type: trekbrief
|
||||
brief_version: "2.2"
|
||||
task: "Opt-in /deep-research-motor for /trekresearch ekstern-fase"
|
||||
slug: deep-research-engine
|
||||
framing: refine
|
||||
status: ready
|
||||
brief_quality: complete
|
||||
research_topics: 3
|
||||
research_status: complete
|
||||
phase_signals_partial: true
|
||||
---
|
||||
|
||||
# Brief — Opt-in `/deep-research`-motor for `/trekresearch`
|
||||
|
||||
## TL;DR
|
||||
- `--engine {swarm|deep-research}` på `/trekresearch` ekstern-fase; `swarm` default (uendret oppførsel).
|
||||
- `deep-research` delegerer ekstern research til Claude Codes innebygde workflow og adapterer inn i research-brief-skjemaet.
|
||||
- Lokal analyse, triangulering og H2-output (`research/NN-*.md`) er motor-uavhengig.
|
||||
- Feature-detekteres med auto-fallback til `swarm` — aldri hard feil.
|
||||
- Surface-only: `commands/` + `agents/`, ingen nye `lib/`-avhengigheter.
|
||||
|
||||
## Intent
|
||||
`/trekresearch` eier i dag hele den eksterne research-fasen selv (Tavily / MS-Learn / Gemini-sverm). Det er portabelt, men du vedlikeholder fan-out, kryssjekk og syntese selv. Anthropic sin innebygde `/deep-research` vedlikeholder fan-out, adversariell påstandsverifisering, sitatfiltrering og websøk for deg. Lar vi operatøren *velge* `/deep-research` for den eksterne fasen, får brukeren en vedlikeholdsfri "turbo" når den er tilgjengelig — uten at Voyage mister sin egen lokal-analyse, triangulering eller H2-kontrakt.
|
||||
|
||||
## Goal
|
||||
`/trekresearch` får et `--engine {swarm|deep-research}`-valg på den eksterne fasen. `swarm` er default (dagens oppførsel). `deep-research` delegerer den eksterne fasen til den innebygde workflowen og adapterer resultatet inn i research-brief-skjemaet. Lokal analyse, triangulering og H2-output (`research/NN-*.md`) er uendret uansett motor.
|
||||
|
||||
## Non-Goals
|
||||
- Bygge om `/trekresearch` på et eget dynamic-workflow-substrat (vurdert og forkastet, Δ≈0).
|
||||
- Multi-provider (Perplexity / Google) — hører til en senere egen `workflow`-motor, ikke denne.
|
||||
- Endre den lokale svermen, handover-kontraktene eller validator-skjemaene.
|
||||
- Gjøre `/deep-research` til default.
|
||||
|
||||
## Constraints / NFR
|
||||
- `/deep-research` krever Claude Code v2.1.154+ og må være på (Pro: via `/config`). Må feature-detekteres med auto-fallback til `swarm` — aldri hard feil.
|
||||
- Output må passere `research-validator.mjs` (confidence ∈ [0,1], dimensions ≥ 1, påkrevde body-seksjoner).
|
||||
- Ingen nye `lib/`-avhengigheter; overflate-endring i `commands/` + `agents/`.
|
||||
- Doc-consistency: oppdater command-tabellen i `CLAUDE.md` + `README.md`, kjør `npm test`.
|
||||
|
||||
## Success Criteria
|
||||
1. `/trekresearch --external --engine swarm "<q>"` gir identisk oppførsel som før (regresjon grønn).
|
||||
2. `/trekresearch --external --engine deep-research "<q>"` produserer en `research/NN-*.md` der `research-validator.mjs --json` rapporterer `valid: true`.
|
||||
3. Med dynamic workflows avskrudd faller `--engine deep-research` automatisk tilbake til `swarm` og logger valget — ingen exception.
|
||||
4. `npm test` (inkl. doc-consistency) er grønn.
|
||||
|
||||
## Research Plan
|
||||
1. **Programmatisk trigging av `/deep-research`** — Kan en plugin-kommando starte den innebygde workflowen og fange artefakten, eller må `commands/trekresearch.md` instruere Claude til å kjøre den? Scope: external · Konfidens: høy · Kost: lav.
|
||||
`/trekresearch --external "Kan en Claude Code plugin-kommando programmatisk starte /deep-research og hente rapporten?"`
|
||||
2. **Output-adapter** — Hvordan mappe `/deep-research` sin siterte rapport til research-brief-skjemaet (dimensjoner, confidence, sitater)? Scope: both · Konfidens: høy · Kost: lav.
|
||||
`/trekresearch --project <dir> --local "Hva krever research-validator.mjs av research/NN-*.md?"`
|
||||
3. **Feature-deteksjon + fallback** — Hvordan oppdage om workflows er på, og hvor i fasen fallback-grenen bør sitte? Scope: local · Konfidens: middels · Kost: lav.
|
||||
|
||||
> **Research-status (2026-06-30, operatør-beslutning «option A»):** Topic 1 (eneste ekte eksterne ukjente) er undersøkt → `docs/deep-research-engine-research.md` (validator-grønn). Funn: `/deep-research` er en innebygd **dynamic workflow** (ikke skill) → trigging KUN via prosa-instruksjon, output **inline i kontekst** (ingen on-disk-artefakt) → motoren må være instruksjons-basert + in-context transform. **SC3 er korrekt som skrevet** (`/deep-research` ER en dynamic workflow). Topic 2 (validator-skjema: `type/created/question` + `## Executive Summary`/`## Dimensions`, `dimensions ≥ 1`) og topic 3 (fallback-plassering) er lokale kode-spørsmål reklassifisert til `/trekplan`-utforskning. `research_status: complete` reflekterer denne beslutningen.
|
||||
|
||||
## Open Questions / Assumptions
|
||||
- Antar at `/deep-research`-rapporten kan reduseres til ≥ 1 dimensjon med per-påstand-sitater uten å bryte trianguleringen. Verifiseres i topic 2.
|
||||
- Uavklart om delegering skjer via instruksjon (trygt) eller programmatisk API (raskere) — topic 1 avgjør.
|
||||
|
||||
## Prior Attempts
|
||||
- Dvalende synthesis-agent bygget, målt (Δ≈0) og forkastet → samme lærdom gjelder substrat-bytte: mål før du adopterer.
|
||||
- Workflow-substratet er allerede dokumentert som vurdert-og-valgt-bort i `docs/architecture.md` §Primitives per step.
|
||||
|
|
@ -1,213 +0,0 @@
|
|||
---
|
||||
type: trekresearch-brief
|
||||
created: 2026-06-30
|
||||
question: "Can a Claude Code plugin command programmatically trigger the built-in /deep-research workflow and capture its report, or must it instruct Claude to run it?"
|
||||
confidence: 0.85
|
||||
dimensions: 4
|
||||
mcp_servers_used: []
|
||||
local_agents_used: [claude-code-guide]
|
||||
external_agents_used: []
|
||||
slug: deep-research-engine
|
||||
feeds_brief: docs/deep-research-engine-brief.md
|
||||
research_topic: 1
|
||||
---
|
||||
|
||||
# Research — Programmatic trigging of `/deep-research` from a plugin command
|
||||
|
||||
> Targeted single-topic research for the `deep-research-engine` brief (Topic 1 of
|
||||
> the brief's Research Plan). Topics 2 & 3 are local Voyage-code questions folded
|
||||
> into `/trekplan` exploration; only Topic 1 was a genuine external unknown.
|
||||
> Method: `claude-code-guide` agent (Anthropic docs + CHANGELOG, cited) +
|
||||
> direct inspection of this machine (CC 2.1.196 binary, a real local `/deep-research`
|
||||
> run). Not run through the `/trekresearch` swarm — see brief reconcile note.
|
||||
|
||||
## Research Question
|
||||
|
||||
Can a Claude Code **plugin slash-command** (markdown under `commands/`)
|
||||
**programmatically** start the built-in `/deep-research` workflow and **capture its
|
||||
report artifact** for adaptation into the research-brief schema — or must the command
|
||||
instead **instruct** Claude (in prose) to run `/deep-research` and then transform the
|
||||
in-context result?
|
||||
|
||||
## Executive Summary
|
||||
|
||||
Programmatic trigger + file-based capture is **not feasible**: `/deep-research` is a
|
||||
built-in **dynamic workflow** (not a skill), deliberately outside the Skill-tool
|
||||
allowlist, and it returns its report **inline into conversation context with no
|
||||
documented on-disk report artifact**. The **only reliable path is instruction-based
|
||||
delegation** — the command's prose tells Claude to run `/deep-research <q>`, then
|
||||
transforms the in-context report in the same turn. Confidence **high** (Anthropic docs +
|
||||
CHANGELOG + local run), with one residual gap: there is no positive "is it enabled?"
|
||||
probe, so feature-detection must lean on the documented *disable* switches + version
|
||||
floor + a `swarm` default.
|
||||
|
||||
## Dimensions
|
||||
|
||||
### 1. Origin + gating -- Confidence: high
|
||||
|
||||
**External findings:**
|
||||
- `/deep-research` is a **built-in dynamic workflow**, not a command and not a bundled
|
||||
skill. `commands.md` marks the `/deep-research <question>` row as **[Workflow]**;
|
||||
`workflows.md`: "Claude Code includes `/deep-research` as a built-in workflow."
|
||||
[VERIFIED — code.claude.com/docs/en/commands.md, code.claude.com/docs/en/workflows.md]
|
||||
- Gating: available on **all paid plans** (pro/max/team/enterprise) + API/Bedrock/Vertex/
|
||||
Foundry. **On Pro it must be turned on** in the *Dynamic workflows* row of `/config`.
|
||||
Also requires the **WebSearch tool** to be available.
|
||||
[VERIFIED — workflows.md, commands.md bundled-workflows row]
|
||||
- Version floor: dynamic workflows were **introduced in CC 2.1.154**; "Dynamic workflows
|
||||
require Claude Code v2.1.154 or later." The brief's `v2.1.154+` + `(Pro: via /config)`
|
||||
constraints are **both correct**. The exact version that first shipped the *named*
|
||||
`/deep-research` workflow is **[NOT DOCUMENTED]** (only a 2.1.196 bugfix mentions it by
|
||||
name) — treat 2.1.154 as the substrate floor, not a proven introduction point.
|
||||
[VERIFIED — workflows.md + CHANGELOG 2.1.154]
|
||||
|
||||
**Local findings:**
|
||||
- This machine runs **CC 2.1.196** (`claude --version`) — substrate floor satisfied.
|
||||
- The exact skill-description string lives **compiled into the binary**
|
||||
(`/Users/ktg/.local/share/claude/versions/2.1.196`, Mach-O 235 MB); there is **no
|
||||
`SKILL.md`** for it anywhere under `~/.claude` (system-wide `find`/`grep` — only hits
|
||||
are this brief + unrelated harness notes). Confirms "Anthropic-bundled, not user skill."
|
||||
|
||||
### 2. Invocation mechanism (programmatic vs instruction) -- Confidence: high
|
||||
|
||||
**External findings:**
|
||||
- **Not** via the `Skill` tool. The Skill-tool built-in allowlist is closed: only
|
||||
`/init`, `/review`, `/security-review` are reachable; "Other built-in commands such as
|
||||
`/compact` are not." `/deep-research` is a Workflow and is not on that list.
|
||||
[VERIFIED — code.claude.com/docs/en/skills.md]
|
||||
- **Instruction-based delegation is the documented mechanism.** Workflows launch when the
|
||||
user types the command, or **when Claude is asked in natural language** ("use a
|
||||
workflow" / "run a workflow") or via the dedicated opt-in keyword CC ships for
|
||||
multi-agent orchestration (CC 2.1.160; spelled out in
|
||||
`docs/cc-upgrade-2.1.181-decision-matrix.md`, which `verify.sh`'s SC1 excludes
|
||||
precisely because it legitimately cites CC keywords). A plugin command whose
|
||||
markdown instructs Claude to run `/deep-research <q>` is therefore the supported path.
|
||||
[VERIFIED — workflows.md "Have Claude write a workflow"]
|
||||
- **Approval gate caveat:** launching a workflow triggers a per-run approval prompt —
|
||||
*every run* in default/acceptEdits; *first launch only* in auto; **never in `claude -p`
|
||||
/ Agent SDK / bypass-permissions** ("the run starts immediately"). Voyage's headless
|
||||
surface (`claude -p`) thus delegates without an interactive gate; interactive sessions
|
||||
hit a prompt. [VERIFIED — workflows.md "Behavior and limits"]
|
||||
- No documented blanket "commands/skills cannot nest" prohibition beyond the Skill-tool
|
||||
allowlist + workflow runtime limits (no mid-run user input; 16 concurrent agents;
|
||||
1000 agents/run). [VERIFIED — workflows.md; NOT DOCUMENTED for a general nesting ban]
|
||||
|
||||
### 3. Output capture -- Confidence: high
|
||||
|
||||
**External findings:**
|
||||
- "When the run finishes, **the report lands in your session**"; "Claude's context holds
|
||||
only the final answer." The report is **in-context**, not a file.
|
||||
[VERIFIED — workflows.md]
|
||||
- What *is* written to disk is the orchestration **script**, not the report: "Every run
|
||||
writes its script to a file under your session's directory in `~/.claude/projects/`."
|
||||
[VERIFIED — workflows.md "How a workflow runs"]
|
||||
|
||||
**Local findings:**
|
||||
- A real `/deep-research` run on this machine left exactly one file —
|
||||
`~/.claude/projects/<session>/workflows/scripts/deep-research-wf_<id>.js` — and **no
|
||||
`.md` report** beside it. [VERIFIED — local filesystem inspection by claude-code-guide]
|
||||
- **Consequence:** a calling command cannot `grep` a results file off disk (none is
|
||||
documented to exist). It can only consume the report **as it sits in conversation
|
||||
context in the same turn** — Claude reads its own prior output and transforms it into
|
||||
research-brief schema. [INFERENCE — file-based capture not feasible; in-context
|
||||
transform is the only avenue]
|
||||
|
||||
### 4. Feature detection + fallback -- Confidence: medium
|
||||
|
||||
**External findings:**
|
||||
- **No positive enumeration / "is-enabled" API or flag is documented.** `claude --help`
|
||||
exposes `--disable-slash-commands` but no "list skills/workflows" or "is-feature-on"
|
||||
flag. [VERIFIED — local `claude --help`; NOT DOCUMENTED for any positive probe]
|
||||
- The documented signals are the **off-switches**, read defensively: `/config` Dynamic-
|
||||
workflows off; `disableWorkflows: true` / `disableBundledSkills: true` in settings.json;
|
||||
`CLAUDE_CODE_DISABLE_WORKFLOWS=1` / `CLAUDE_CODE_DISABLE_BUNDLED_SKILLS=1`.
|
||||
[VERIFIED — workflows.md "Turn workflows off"; CHANGELOG 2.1.x]
|
||||
- **The gap:** the Pro `/config` *on*-state is the very thing you most need to detect, and
|
||||
only the *disable* keys are documented; the persisted key/value for the Pro enable-state
|
||||
is **[NOT DOCUMENTED]**, so it cannot be reliably grepped. A failed/disabled
|
||||
`/deep-research` is also **not documented** to raise a signal a sibling command can
|
||||
catch. [VERIFIED gap]
|
||||
|
||||
**Local findings:**
|
||||
- On this machine both disable keys are unset/absent — workflows are not disabled.
|
||||
|
||||
## External Knowledge
|
||||
|
||||
### Best Practice
|
||||
The "Have Claude write a workflow" + "Behavior and limits" sections of `workflows.md`
|
||||
establish that workflows are operator/Claude-launched, run isolated, and return one
|
||||
in-context report. The Skill-tool allowlist (`skills.md`) is the authoritative statement
|
||||
that only three built-ins are tool-invocable.
|
||||
|
||||
### Known Issues
|
||||
Per-run approval prompts outside `-p`/SDK/bypass mean an interactive `/trekresearch
|
||||
--engine deep-research` will pause for operator approval on each launch — acceptable, but
|
||||
worth documenting in the command UX. The absence of a positive availability probe is the
|
||||
single biggest design constraint (see Dimension 4).
|
||||
|
||||
## Synthesis
|
||||
|
||||
Three cross-cutting insights that only emerge from combining the docs with Voyage's brief:
|
||||
|
||||
1. **The brief's SC3 is correct as written — an earlier review note was wrong.** Because
|
||||
`/deep-research` *is itself* a dynamic workflow, "med dynamic workflows avskrudd faller
|
||||
`--engine deep-research` tilbake til swarm" is the right feature-detection axis. A
|
||||
prior brief-review remark that SC3 "conflated `/deep-research` (skill) with dynamic
|
||||
workflows" was based on a wrong premise (that `/deep-research` was a skill) and is
|
||||
retracted. **No SC3 brief edit is needed.**
|
||||
|
||||
2. **The engine must be instruction-based + in-context, never file-based.** Topic 1's
|
||||
open question ("instruction (safe) vs. programmatic API (faster) — topic 1 decides")
|
||||
resolves decisively to **instruction-based**: there is no programmatic API and no
|
||||
on-disk report. The `deep-research` engine path in `commands/trekresearch.md` must
|
||||
(a) instruct Claude to run `/deep-research <q>`, then (b) transform the in-context
|
||||
report into the `research/NN-*.md` schema in the same turn. This is surface-only
|
||||
(`commands/` prose), matching the brief's "ingen nye `lib/`-avhengigheter."
|
||||
|
||||
3. **SC3's "ingen exception" cannot rest on runtime detection — pin it to a default.**
|
||||
Since no positive availability probe exists, robust fallback = version-floor check
|
||||
(≥ 2.1.154) + disable-key heuristic (`disableWorkflows` / env) + **`--engine` default
|
||||
of `swarm`** (explicit opt-in). The fallback is "graceful degradation by design,"
|
||||
not "catch an exception at runtime." This refines, but does not contradict, SC3.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- **Adapter fidelity (brief Topic 2, partly answered locally):** `research-validator.mjs`
|
||||
requires `type: trekresearch-brief` + `created` + `question`; validates `confidence ∈
|
||||
[0,1]` and `dimensions ≥ 1` if present; body must carry `## Executive Summary` +
|
||||
`## Dimensions`. So the `/deep-research` report **can** be reduced to ≥ 1 dimension with
|
||||
per-claim citations and a confidence number — the brief's assumption holds. The
|
||||
remaining open part (how cleanly the in-context report maps to per-dimension
|
||||
local/external splits) is a `/trekplan` exploration concern, not an external unknown.
|
||||
- **Exact intro version of the *named* `/deep-research` workflow** — not documented; the
|
||||
2.1.154 dynamic-workflows floor is the safe pin.
|
||||
|
||||
## Recommendation
|
||||
|
||||
Build the `deep-research` engine as **instruction-based delegation with in-context
|
||||
adaptation**, not a programmatic trigger:
|
||||
1. `--engine deep-research` makes `commands/trekresearch.md`'s external phase instruct
|
||||
Claude to run `/deep-research <q>` and transform the returned in-context report into
|
||||
`research/NN-*.md` (validator-conformant: `## Executive Summary` + `## Dimensions`,
|
||||
`confidence`, `dimensions ≥ 1`).
|
||||
2. Feature-detect by **graceful degradation**: version floor + disable-key heuristic +
|
||||
`--engine` default `swarm`. Do not depend on a positive availability probe (none
|
||||
exists). Log the chosen engine. This satisfies SC3 without a runtime exception.
|
||||
3. Keep `swarm` the default (brief's Non-Goal: "ikke default-bytte"). Document the
|
||||
per-run approval prompt for interactive (non-`-p`) sessions.
|
||||
|
||||
Confidence in the recommendation: **high** for the mechanism (instruction-based +
|
||||
in-context), **medium** for the exact fallback-detection ergonomics (the one documented
|
||||
gap). This is sufficient to green-light `/trekplan` with Topic 1 resolved.
|
||||
|
||||
## Sources
|
||||
|
||||
| # | Source | Type | Quality | Used in |
|
||||
|---|--------|------|---------|---------|
|
||||
| 1 | code.claude.com/docs/en/workflows.md | official | high | Dim 1,2,3,4 + Synthesis |
|
||||
| 2 | code.claude.com/docs/en/commands.md | official | high | Dim 1 (Workflow classification, WebSearch req) |
|
||||
| 3 | code.claude.com/docs/en/skills.md | official | high | Dim 2 (Skill-tool allowlist) |
|
||||
| 4 | github.com/anthropics/claude-code CHANGELOG (2.1.154, 2.1.x, 2.1.196) | official | high | Dim 1,4 (version floor, disable keys) |
|
||||
| 5 | Local: CC 2.1.196 binary inspection (`find`/`grep`, `claude --version`) | codebase | high | Dim 1 (bundled, no SKILL.md) |
|
||||
| 6 | Local: real `/deep-research` run — only `.js` script written, no `.md` report | codebase | high | Dim 3 (no on-disk artifact) |
|
||||
| 7 | Local: `lib/validators/research-validator.mjs` schema | codebase | high | Open Questions (adapter feasibility / Topic 2) |
|
||||
|
|
@ -7,11 +7,9 @@ review/coordinator agent **misfires** on a real task, the case is distilled into
|
|||
a machine-readable record and added here, so the failure can never silently
|
||||
regress.
|
||||
|
||||
This corpus is the gold that the **offline gold-scored output eval (SKAL-1·4b)**
|
||||
scores against — that eval is now implemented and wired into `node --test`
|
||||
(see [§Gold-scored output eval](#gold-scored-output-eval-skal-14b) below). The
|
||||
corpus records here are committed fixtures + a schema; the scoring run is the
|
||||
separate piece that consumes them.
|
||||
This tier is the **staging ground** for the offline gold-scored output eval
|
||||
(SKAL-1·4b). It is deliberately **not wired into CI** yet — these are committed
|
||||
fixtures + a schema, not a scoring run.
|
||||
|
||||
## Seed example
|
||||
|
||||
|
|
@ -67,38 +65,10 @@ A corpus file is one JSON object:
|
|||
3. Add (or extend) a loader test in the `tests/lib/gold-corpus.test.mjs` shape so
|
||||
the record's shape + catalogue membership are pinned under `node --test`.
|
||||
|
||||
## Gold-scored output eval (SKAL-1·4b)
|
||||
|
||||
The offline scoring run that grades a recorded agent run against this corpus.
|
||||
**Offline** = committed reviewer payloads, no live agent spawn, no LLM, no
|
||||
network (the LLM-in-the-loop grading is the separate 4c tier).
|
||||
|
||||
- **Committed runs** live under `tests/fixtures/bakeoff-rich/runs/`. Each file is
|
||||
the JSON **reviewer payloads** (one object per reviewer, `{ reviewer, findings }`)
|
||||
that a recorded run produced. `run-perfect.json` is the regression guard: fed
|
||||
through the coordinator contract it must reproduce every seeded gold finding.
|
||||
- **The contract** `lib/review/coordinator-contract.mjs::runContract(payloads)`
|
||||
turns those payloads into the deterministic coordinator output (4a).
|
||||
- **The scorer** `lib/review/gold-scorer.mjs`:
|
||||
- `scoreFindings(runFindings, goldFindings)` matches at **`(file, rule_key)`
|
||||
granularity** (line + severity ignored) → `{ tp, fp, fn, precision, recall,
|
||||
f1, matched, missed, spurious }`. Vacuous-set conventions (empty run →
|
||||
recall 0; empty gold → precision 0; f1 collapses to 0) are documented in the
|
||||
module header.
|
||||
- `scoreVerdict(runVerdict, goldVerdict)` → exact verdict match.
|
||||
- **The scoring run** `tests/lib/gold-eval.test.mjs` asserts `run-perfect`
|
||||
reproduces gold at precision/recall/f1 = 1.0 and `verdict === expected_verdict`.
|
||||
- **Third test-census category.** `lib/util/test-census.mjs` now reports a
|
||||
`goldEval` bucket (matched by `GOLD_EVAL_FILE_RE`) separately from `behavior`
|
||||
and `docPins` — a scoring run is neither behavior coverage nor a prose pin, so
|
||||
the honest-count invariant is now a 3-way sum.
|
||||
|
||||
## Future hardening (not in this tier)
|
||||
|
||||
- SKAL-1·4b: an offline gold-scored output eval that scores recorded agent runs
|
||||
against these records at `(file, rule_key)` granularity, with committed runs
|
||||
and a third test-census category.
|
||||
- A prose↔JSON cross-assertion (parse the fixture README table, diff against
|
||||
`gold.json`) to mechanically bound the two-source-of-truth drift.
|
||||
- Degraded-run fixtures (a run that misses or invents findings) to exercise the
|
||||
scorer's discriminating path end-to-end; today that path is covered by
|
||||
`tests/lib/gold-scorer.test.mjs` with inline synthetic findings.
|
||||
- SKAL-1·4c: the LLM-in-the-loop eval that grades live agent runs (needs a
|
||||
filesystem + model judgement, deliberately excluded from this deterministic tier).
|
||||
|
|
|
|||
|
|
@ -28,14 +28,13 @@ A revived Path C (post-v2.2.xxx) would require: (1) re-architecting tool-list to
|
|||
|
||||
## Profile system (`--profile`, v4.1.0)
|
||||
|
||||
Four built-in model profiles plus operator-defined `<custom>.yaml`. Each profile pins `phase_models` for the six pipeline phases (`brief`, `research`, `plan`, `execute`, `review`, `continue`). Profile is recorded in plan.md frontmatter as `profile: <name>` and emitted to `${CLAUDE_PLUGIN_DATA}/trek*-stats.jsonl` for cost-attribution.
|
||||
Three built-in model profiles plus operator-defined `<custom>.yaml`. Each profile pins `phase_models` for the six pipeline phases (`brief`, `research`, `plan`, `execute`, `review`, `continue`). Profile is recorded in plan.md frontmatter as `profile: <name>` and emitted to `${CLAUDE_PLUGIN_DATA}/trek*-stats.jsonl` for cost-attribution.
|
||||
|
||||
| Profile | Brief | Research | Plan | Execute | Review | Continue | Use case |
|
||||
|---------|-------|----------|------|---------|--------|----------|----------|
|
||||
| `economy` | sonnet | sonnet | sonnet | sonnet | sonnet | sonnet | ⚠ **Experimental** (uncalibrated Jaccard floor) — lowest cost; high-confidence small-scope tasks (operator-opt-in via `--profile economy`) |
|
||||
| `balanced` | sonnet | sonnet | opus | sonnet | opus | sonnet | Mixed — opus where reasoning depth pays off (operator-opt-in via `--profile balanced`) |
|
||||
| `premium` (default) | opus | opus | opus | opus | opus | opus | Maximum quality — Opus on every phase. Default since 2026-05-13 operator request; also the hardcoded resolver default returned by `resolveProfile()` in `lib/profiles/resolver.mjs` |
|
||||
| `fable` | fable | fable | fable | fable | fable | fable | Max quality — Fable 5 (Mythos-class, above Opus) on every phase (operator-opt-in via `--profile fable`); reasoning effort inherits from the session — see `docs/profiles.md` §Model & effort axes |
|
||||
|
||||
### Lookup order
|
||||
|
||||
|
|
@ -46,7 +45,7 @@ Four built-in model profiles plus operator-defined `<custom>.yaml`. Each profile
|
|||
|
||||
### Custom profiles
|
||||
|
||||
Create `voyage-profiles/<custom>.yaml` in the repo root (or `~/.claude/voyage-profiles/<custom>.yaml`) to define a **new** tier — the name must not be a built-in. The validator (`lib/validators/profile-validator.mjs`) enforces: every `phase_models[].phase` must be a known phase enum; every `phase_models[].model` must exactly match an entry in `BASE_ALLOWED_MODELS` (`['sonnet', 'opus', 'fable']`; `haiku` only with `VOYAGE_ALLOW_HAIKU=1`). `findProfilePath` (`lib/profiles/resolver.mjs`) resolves **built-in first** (`lib/profiles/<name>.yaml` for `economy`/`balanced`/`premium`/`fable`), then repo-root `voyage-profiles/`, then `~/.claude/voyage-profiles/`. A custom file named after a built-in therefore **cannot** shadow it (custom profiles must use new names); for the same custom name, repo-root takes precedence over home.
|
||||
Create `voyage-profiles/<custom>.yaml` in the repo root (or `~/.claude/voyage-profiles/<custom>.yaml`) to define a **new** tier — the name must not be a built-in. The validator (`lib/validators/profile-validator.mjs`) enforces: every `phase_models[].phase` must be a known phase enum; every `phase_models[].model` must match `^(opus|sonnet)(\b|-).*` or one of the canonical short names. `findProfilePath` (`lib/profiles/resolver.mjs`) resolves **built-in first** (`lib/profiles/<name>.yaml` for `economy`/`balanced`/`premium`), then repo-root `voyage-profiles/`, then `~/.claude/voyage-profiles/`. A custom file named after a built-in therefore **cannot** shadow it (custom profiles must use new names); for the same custom name, repo-root takes precedence over home.
|
||||
|
||||
Drift between plan-frontmatter `profile:` and step-manifest `profile_used:` emits a `MANIFEST_PROFILE_DRIFT` warning from `plan-validator --strict` (Step 20). Plan remains valid; the warning surfaces accidental tier-mismatch.
|
||||
|
||||
|
|
|
|||
|
|
@ -6,15 +6,14 @@ cost estimation (with disclaimer).
|
|||
|
||||
## Built-in profiles
|
||||
|
||||
Four pre-defined tiers ship with the plugin (fable added in v5.9), located at
|
||||
`lib/profiles/{economy,balanced,premium,fable}.yaml`.
|
||||
Three pre-defined tiers ship with v4.1, located at
|
||||
`lib/profiles/{economy,balanced,premium}.yaml`.
|
||||
|
||||
| Profile | Brief | Research | Plan | Execute | Review | Continue | Use case |
|
||||
|---------|-------|----------|------|---------|--------|----------|----------|
|
||||
| `economy` | sonnet | sonnet | sonnet | sonnet | sonnet | sonnet | ⚠ **Experimental** (uncalibrated Jaccard floor) — lowest cost; small-scope tasks where you have high confidence the brief is right |
|
||||
| `balanced` | sonnet | sonnet | opus | sonnet | opus | sonnet | Mixed — opus where reasoning depth pays off (plan synthesis + adversarial review); opt-in via `--profile balanced` |
|
||||
| `premium` (default) | opus | opus | opus | opus | opus | opus | Maximum quality — Opus on every phase + external research on (default since the 2026-05-13 operator decision) |
|
||||
| `fable` | fable | fable | fable | fable | fable | fable | Max quality — Fable 5 (Mythos-class, above Opus) on every phase; opt-in via `--profile fable`; reasoning effort inherits from the session (see Model & effort axes) |
|
||||
|
||||
`premium` is the default tier — set by the 2026-05-13 operator decision and
|
||||
matched by the hardcoded resolver default in `lib/profiles/resolver.mjs`. It
|
||||
|
|
@ -23,8 +22,7 @@ roughly 5× the sub-agent cost of an all-sonnet run, accepted as a deliberate
|
|||
trade-off. Drop to `--profile balanced` (opus only on the two phases where
|
||||
quality matters most — Plan synthesis + Review — and sonnet everywhere else)
|
||||
or `--profile economy` (sonnet everywhere) when cost or latency matters more
|
||||
than depth. Step up to `--profile fable` (Fable 5 on every phase) when
|
||||
maximum quality is wanted end-to-end and cost is not a constraint.
|
||||
than depth.
|
||||
|
||||
`economy` is *strictly experimental* in v4.1, and says so in the profile
|
||||
data itself: `lib/profiles/economy.yaml` carries `experimental: true`. The
|
||||
|
|
@ -38,19 +36,10 @@ back to `balanced`.
|
|||
|
||||
## Model & effort axes
|
||||
|
||||
`opus`, `sonnet`, and `fable` are model **aliases**, not pinned ids. As of
|
||||
Claude Code 2.1.154 the `opus` alias resolves to **Opus 4.8**, whose default
|
||||
reasoning effort is **`high`**; `sonnet` resolves to Sonnet 4.6; `fable`
|
||||
resolves to **Fable 5** (Mythos-class, positioned above Opus), whose default
|
||||
reasoning effort is also `high`. The profile table above selects *which
|
||||
alias* runs each phase — it does not touch reasoning effort.
|
||||
|
||||
**Reasoning effort inherits from the session.** Voyage effort (orchestration
|
||||
shape — which agents/passes run) and model reasoning effort are different
|
||||
axes. Fable 5's default reasoning effort is `high`, NOT xhigh, and switching
|
||||
model resets effort to the model default — xhigh does not follow the model.
|
||||
To run the fable tier at xhigh, set it at session level: `/effort xhigh`, the
|
||||
`effortLevel` setting, or `CLAUDE_CODE_EFFORT_LEVEL`.
|
||||
`opus` and `sonnet` are model **aliases**, not pinned ids. As of Claude Code
|
||||
2.1.154 the `opus` alias resolves to **Opus 4.8**, whose default reasoning
|
||||
effort is **`high`**; `sonnet` resolves to Sonnet 4.6. The profile table above
|
||||
selects *which alias* runs each phase — it does not touch reasoning effort.
|
||||
|
||||
Two different things share the word "effort" in Voyage. They are **orthogonal
|
||||
axes** — same name, different mechanism:
|
||||
|
|
@ -65,8 +54,8 @@ axes** — same name, different mechanism:
|
|||
|
||||
The `phase-signal-resolver.mjs` helper only reads the **orchestration** axis
|
||||
(`phase_signals.effort`, gated against `low/standard/high`) plus the optional
|
||||
per-phase `model` (gated against `['sonnet','opus','fable']`). It never emits
|
||||
native `effort:`.
|
||||
per-phase `model` (gated against `['sonnet','opus']`). It never emits native
|
||||
`effort:`.
|
||||
|
||||
**Native `effort:` on agents.** Voyage sets the reasoning axis statically on
|
||||
selected agents, additively over the Opus-4.8 default:
|
||||
|
|
@ -131,13 +120,11 @@ The validator (`lib/validators/profile-validator.mjs`) enforces:
|
|||
|
||||
- Every `phase_models[].phase` must be a known phase enum:
|
||||
`brief` / `research` / `plan` / `execute` / `review` / `continue`
|
||||
- Every `phase_models[].model` must exactly match an entry in
|
||||
`BASE_ALLOWED_MODELS` (`['sonnet', 'opus', 'fable']` in
|
||||
`lib/validators/profile-validator.mjs`; `haiku` only with
|
||||
`VOYAGE_ALLOW_HAIKU=1`)
|
||||
- Every `phase_models[].model` must match `^(opus|sonnet)(\b|-).*` or
|
||||
one of the canonical short names
|
||||
- All six phases must be present (no partial profiles)
|
||||
|
||||
The four built-in names (`economy`, `balanced`, `premium`, `fable`) resolve to their
|
||||
The three built-in names (`economy`, `balanced`, `premium`) resolve to their
|
||||
bundled yaml first — `findProfilePath()` returns the built-in before consulting
|
||||
`voyage-profiles/`, so a same-named custom file is ignored and cannot shadow a
|
||||
built-in. To customize, give your profile a new name and reference it via
|
||||
|
|
|
|||
|
|
@ -1,160 +0,0 @@
|
|||
# Spike: does a plugin `PreToolUse` hook reach sub-agent tool calls?
|
||||
|
||||
**Date:** 2026-08-09
|
||||
**Claude Code version:** 2.1.226
|
||||
**Plugin:** voyage 5.9.1 (installed from `ktg-plugin-marketplace`)
|
||||
**Gates:** Step 10 of `plan.md` (`2026-06-30-trekresearch-storm-upgrade`)
|
||||
|
||||
## Question
|
||||
|
||||
Step 10 wants to enforce the conversation-turn cap in a `PreToolUse` hook. That
|
||||
is only viable if a **plugin** `PreToolUse` hook fires on tool calls made
|
||||
*inside a sub-agent*. If it does not, the cap must be enforced somewhere else
|
||||
and Step 10 becomes a documented downgrade instead.
|
||||
|
||||
The question is genuinely open, not answerable from docs alone: the current
|
||||
[hooks reference](https://code.claude.com/docs/en/hooks) states that sub-agent
|
||||
tool calls fire the same hooks and carry `agent_id` / `agent_type`, while
|
||||
[issue #34692](https://github.com/anthropics/claude-code/issues/34692) reported
|
||||
the exact opposite behaviour. The answer is therefore version-dependent and had
|
||||
to be measured on the version actually in use.
|
||||
|
||||
## Method
|
||||
|
||||
A one-shot probe hook was registered for the `WebSearch` matcher, and a
|
||||
**headless child session** was launched to exercise it. The child was used
|
||||
because hooks are resolved when a session starts — a matcher added mid-session
|
||||
cannot be observed by the session that added it.
|
||||
|
||||
Two deviations from the step as originally written, both forced and both
|
||||
verified not to affect the result:
|
||||
|
||||
1. **The matcher was injected into the installed plugin's `hooks.json`, not the
|
||||
repository's.** The plan assumed the repo working tree *is* the active plugin
|
||||
root. It is not: `~/.claude/plugins/cache/ktg-plugin-marketplace/voyage/5.9.1/`
|
||||
is a plain directory holding its own copy, and that copy is what loads.
|
||||
Editing `hooks/hooks.json` in the repo would have measured nothing. The cache
|
||||
file was backed up, modified, and restored — verified byte-identical to the
|
||||
repo file afterwards.
|
||||
2. **The probe script lives under the session scratchpad, not `${TMPDIR}`.** A
|
||||
pathguard hook refuses writes to `${TMPDIR}`. The load-bearing property was
|
||||
only that the script sit **outside `hooks/scripts/`**, which
|
||||
`tests/lib/doc-consistency.test.mjs:75-84` counts via `readdirSync`; the
|
||||
scratchpad satisfies that just as well. The directory still holds 7 scripts.
|
||||
|
||||
### Probe hook
|
||||
|
||||
Logged every invocation as one JSON line (`tool_name`, `agent_id`,
|
||||
`agent_type`, plus the untouched stdin) and always exited `0`, so it could not
|
||||
alter the child's behaviour.
|
||||
|
||||
### Commands
|
||||
|
||||
```bash
|
||||
# 1. inject the temporary matcher into the INSTALLED plugin
|
||||
node -e '...push {matcher:"WebSearch", ...voyage-spike-hook.mjs} into hooks.PreToolUse...'
|
||||
|
||||
# 2. exercise it from a fresh child session
|
||||
claude -p "Spawn exactly one sub-agent via the Agent tool (subagent_type: general-purpose).
|
||||
Instruct that sub-agent to perform exactly ONE WebSearch for the query
|
||||
'claude code hooks reference' and report back the first result title.
|
||||
You MUST NOT call WebSearch yourself in the main context - only the
|
||||
sub-agent may call it. When the sub-agent returns, reply with the word DONE." \
|
||||
--allowedTools "Agent,Task,WebSearch" \
|
||||
--max-turns 15
|
||||
|
||||
# 3. restore
|
||||
cp "${TMPDIR}voyage-hooks-backup.json" <cache>/hooks/hooks.json
|
||||
```
|
||||
|
||||
The child returned `DONE`.
|
||||
|
||||
> An earlier attempt additionally passed `--permission-mode bypassPermissions`
|
||||
> and was refused by the auto-mode classifier. The flag was dropped;
|
||||
> `--allowedTools` alone was sufficient.
|
||||
|
||||
## Raw observation
|
||||
|
||||
The log contains **exactly one** record — so the main context did not call
|
||||
`WebSearch` itself, and the single entry is unambiguously the sub-agent's call:
|
||||
|
||||
```json
|
||||
{"at":"2026-08-09T12:07:47.796Z","tool_name":"WebSearch",
|
||||
"agent_id":"aa6d19525a4680fe0","agent_type":"general-purpose",
|
||||
"raw_stdin":"{\"session_id\":\"b126fd6a-...\",\"cwd\":\"/Users/ktg/repos/ktg-plugin-marketplace/voyage\",
|
||||
\"permission_mode\":\"auto\",\"agent_id\":\"aa6d19525a4680fe0\",\"agent_type\":\"general-purpose\",
|
||||
\"effort\":{\"level\":\"xhigh\"},\"hook_event_name\":\"PreToolUse\",\"tool_name\":\"WebSearch\",
|
||||
\"tool_input\":{\"query\":\"claude code hooks reference\"},\"tool_use_id\":\"toolu_01Ka4...\"}"}
|
||||
```
|
||||
|
||||
Both `agent_id` and `agent_type` are populated, matching the documented
|
||||
common-input fields for sub-agent-originated tool events. A main-context call
|
||||
would have carried neither.
|
||||
|
||||
## Consequence for Step 10
|
||||
|
||||
A plugin `PreToolUse` hook **does** observe sub-agent tool calls on CC 2.1.226,
|
||||
and can attribute them via `agent_id` / `agent_type`. Step 10 may therefore take
|
||||
the enforcement branch rather than the documented-downgrade branch.
|
||||
|
||||
Two limits worth carrying forward, neither of which changes the verdict:
|
||||
|
||||
- This measures `WebSearch` on one CC version. The behaviour regressed once
|
||||
before (#34692), so the hook must fail **open**, never assume it is the only
|
||||
gate, and the cap must remain correct if the hook silently stops firing.
|
||||
- The probe only establishes *reach*. Whether a **blocking** (exit 2) decision
|
||||
from inside a sub-agent propagates usefully was not measured — the probe
|
||||
always exited 0 by design.
|
||||
|
||||
RESULT: FIRES
|
||||
|
||||
## Enforcement outcome
|
||||
|
||||
Step 10 took the **enforcement branch**: `hooks/scripts/pre-agent-cap.mjs`
|
||||
(PreToolUse, matcher `WebSearch|WebFetch|Task`), pinned by
|
||||
`tests/hooks/agent-cap.test.mjs`.
|
||||
|
||||
What it does: counts turns spent by a run — read-only, from the append-only
|
||||
ledger that `lib/util/research-loop-cap.mjs` writes — and exits 2 once
|
||||
`turns_used >= max_conv_turns × maxDimensions`. It never appends to the
|
||||
ledger; a cap that recorded its own enforcement would count itself.
|
||||
|
||||
Scope key, the part that makes a globally-wired `PreToolUse` hook safe:
|
||||
`session_id` **+** a marker file only the Phase 5 loop writes, at
|
||||
`${CLAUDE_PLUGIN_DATA}/trekresearch-loop-scope/<session_id>.json`:
|
||||
|
||||
```json
|
||||
{ "runId": "<run id>", "startedAt": "<ISO-8601>" }
|
||||
```
|
||||
|
||||
No marker for the calling session ⇒ out of scope ⇒ allow, unconditionally.
|
||||
An unrelated session is never denied because some other run spent its budget.
|
||||
|
||||
Fail-open and fail-closed are split deliberately:
|
||||
|
||||
| Condition | Outcome | Why |
|
||||
|---|---|---|
|
||||
| No marker / no `session_id` / unparsable stdin | allow | Not evidence of a loop turn |
|
||||
| Marker older than TTL (default 2h, `VOYAGE_CAP_SCOPE_TTL_MS`) | allow + auto-reset | A crashed run must not deny tool calls forever, and `--resume` keeps the same `session_id` |
|
||||
| `VOYAGE_DISABLE_CAP_HOOK=1` | allow | Kill switch |
|
||||
| `VOYAGE_STORM_ENABLED` ≠ `1` | allow | Default-off: no loop runs, nothing to enforce |
|
||||
| In scope, `CLAUDE_PLUGIN_DATA` absent | **deny** | A budget control that cannot count must not grant — same stance as `research-loop-cap.mjs` |
|
||||
| In scope, budget spent | **deny (exit 2)** | The bound |
|
||||
|
||||
Both limits recorded above still hold and are not closed by this step. Reach
|
||||
was measured on one CC version for one tool, and blocking-propagation from
|
||||
inside a sub-agent was never measured — so this hook is **defence in depth**,
|
||||
and `research-loop-cap.mjs` must remain correct on its own if the hook
|
||||
silently stops firing.
|
||||
|
||||
**Follow-up, closed (S79).** Step 10 shipped the hook correct but **latent** —
|
||||
nothing wrote the marker, and it enforces exactly when a marker exists.
|
||||
`commands/trekresearch.md` Phase 5 now writes it (`### Loop scope marker`,
|
||||
keyed by `CLAUDE_CODE_SESSION_ID`, carrying the same `run_id` the ledger is
|
||||
counted under) and removes it on all three exits. Cleanup covers exits only; a
|
||||
crashed session is covered by the TTL above. Verified end-to-end: the snippet
|
||||
as shipped arms the hook, the hook denies at `8/8` turns, and the removal
|
||||
snippet returns it to allow so Phase 6 can still spawn agents. Pinned by
|
||||
`tests/lib/doc-consistency.test.mjs` (STORM marker), which derives the
|
||||
directory name from `SCOPE_DIRNAME` in the hook, so renaming either side
|
||||
fails.
|
||||
|
|
@ -1,131 +0,0 @@
|
|||
# STORM adoption gate — pre-registered measurement protocol
|
||||
|
||||
**Status:** thresholds registered, **no measurement run yet.** This document is
|
||||
committed *before* the first measurement by design: a threshold chosen after
|
||||
seeing the numbers is not a threshold. **Harness:** `scripts/storm-measure.mjs`
|
||||
(pinned by `tests/scripts/storm-measure.test.mjs`). **Decides:** whether the
|
||||
bounded Phase 5 conversation loop (`commands/trekresearch.md` §Loop bound) is
|
||||
worth turning on by default.
|
||||
|
||||
---
|
||||
|
||||
## 1. What this gate measures — and what it does not
|
||||
|
||||
The gate measures **source and coverage breadth**:
|
||||
|
||||
| Metric | Definition | Arm |
|
||||
|---|---|---|
|
||||
| `unique_sources` | distinct external sources cited by the run | between-arm median gain |
|
||||
| dimensions over baseline | `(dimensions − dimensions_baseline) / dimensions_baseline` | within-run median across the treatment arm |
|
||||
|
||||
It does **not** measure outline quality, answer correctness, synthesis
|
||||
usefulness, or operator satisfaction. A breadth win is a *necessary* condition
|
||||
for adoption, never a sufficient one. If the loop widens coverage by 40% and
|
||||
the resulting briefs read worse, the correct action is still decline — the
|
||||
number does not overrule a reading of the artifacts.
|
||||
|
||||
The second metric is a within-run delta by construction (`dimensions_baseline`
|
||||
is the interview dimension count, `dimensions` the post-Phase-4.5 list), so it
|
||||
needs no control arm; the control arm's value is 0 because the loop is inert
|
||||
below `effort: high`.
|
||||
|
||||
## 2. Pre-registered thresholds
|
||||
|
||||
Either metric clearing the bar is enough. The brief pre-registers "median
|
||||
forbedring ≥ 30 % på (a) eller (b) → adopt. < 15 % → decline" — a breadth win
|
||||
on one axis counts, because either axis widening is the effect the loop claims.
|
||||
Adopt is evaluated first, so a run that clears the adopt bar on one metric is
|
||||
an adopt even when the other metric sits under the decline bar.
|
||||
|
||||
| Median gain | Verdict | Action |
|
||||
|---|---|---|
|
||||
| ≥ 30% on **either** metric | **adopt** | Flip the `VOYAGE_STORM_ENABLED` default (see §5) |
|
||||
| < 15% on **either** metric (and no adopt) | **decline** | Leave the mechanism default-off. This is a **no-op**: nothing is rolled back |
|
||||
| both metrics in 15% – 30% | **inconclusive** | Keep default-off, gather more runs, re-measure |
|
||||
| either arm empty | **insufficient-data** | Not a decline — measure more |
|
||||
|
||||
`ADOPT_THRESHOLD = 0.30` and `DECLINE_THRESHOLD = 0.15` are exported constants
|
||||
in `scripts/storm-measure.mjs` and pinned by the test suite. Changing them is a
|
||||
deliberate, reviewable act, not a tuning knob to be nudged toward a result.
|
||||
|
||||
## 3. Excluded runs (the honest denominator)
|
||||
|
||||
Runs with `empty_turns > 0` are **excluded from the gain and reported as a
|
||||
count**. An empty turn is one that spent budget and returned no findings, or
|
||||
findings without citations. A run whose `empty_turns` cannot be **read** as a
|
||||
number is excluded on the same footing: a field that says `"many"` is not
|
||||
evidence of zero empty turns, and treating it as one would let the least
|
||||
trustworthy run back into the denominator in the direction that flatters
|
||||
adoption. Including those runs decides adoption on a broken
|
||||
denominator — the loop looks cheap because its failures are averaged into its
|
||||
successes. The harness prints the excluded count on every invocation; if that
|
||||
count is a large fraction of the treatment arm, the finding is about the loop's
|
||||
reliability, and it should be read before the gain figure is read at all.
|
||||
|
||||
Rows predating the Step 9 measurement fields carry no `effort` and are dropped
|
||||
as `legacy` with a count. A stats file where *no* row carries `effort` is a hard
|
||||
error, not an empty treatment group: a schema gap must never present itself as
|
||||
"no gain".
|
||||
|
||||
## 4. The measurement runs (operator-run, outside this plan)
|
||||
|
||||
The measurement itself is **not** part of the implementation plan that built
|
||||
this harness. It is an operator-run gate between that plan and any adopt commit.
|
||||
|
||||
Protocol:
|
||||
|
||||
- **n ≥ 5 runs per arm.** Fewer, and the median is an anecdote.
|
||||
- **The same question set in both arms.** Two briefs, run as `--project` runs.
|
||||
- The arms differ in exactly one thing: whether the loop is enabled.
|
||||
- Both arms append to the same `trekresearch-stats.jsonl`; `effort` is the
|
||||
grouping key that separates them.
|
||||
|
||||
```bash
|
||||
# Control arm (loop inert) — n >= 5
|
||||
claude -p "/trekresearch --project .claude/projects/<brief-standard-effort>"
|
||||
|
||||
# Treatment arm (loop live, effort: high in the brief's phase_signals) — n >= 5
|
||||
VOYAGE_STORM_ENABLED=1 \
|
||||
claude -p "/trekresearch --project .claude/projects/<brief-high-effort>"
|
||||
|
||||
# The gate
|
||||
node scripts/storm-measure.mjs --stats "${CLAUDE_PLUGIN_DATA}/trekresearch-stats.jsonl"
|
||||
|
||||
# Machine-readable, for a decision record
|
||||
node scripts/storm-measure.mjs --json --stats "${CLAUDE_PLUGIN_DATA}/trekresearch-stats.jsonl"
|
||||
|
||||
# SC activation check — BOTH halves: did the latest high-effort run discover at
|
||||
# least one dimension, AND is its final list a true SUPERSET of the interview
|
||||
# ones? The second half is read from the run's own `dimensions_baseline_preserved`
|
||||
# attestation, because the dimension NAMES that would show it directly are prose
|
||||
# the exporter allowlist denies. A run that does not attest it FAILS — dropping
|
||||
# two interview dimensions and appending three discovered ones is a +1 count
|
||||
# delta and not a superset. Exit 0 = both halves hold.
|
||||
node scripts/storm-measure.mjs --activation-check --stats "${CLAUDE_PLUGIN_DATA}/trekresearch-stats.jsonl"
|
||||
```
|
||||
|
||||
## 5. What "adopt" concretely means
|
||||
|
||||
Adoption is **one constant**: `isStormEnabled()` in
|
||||
`lib/util/research-loop-cap.mjs` currently requires `VOYAGE_STORM_ENABLED === '1'`.
|
||||
Adopt = make the loop's budget non-zero without that opt-in, in a commit that
|
||||
cites the measurement output.
|
||||
|
||||
This asymmetry is deliberate and was designed in before any code was written:
|
||||
|
||||
- **decline costs nothing** — the mechanism ships default-off, so declining is
|
||||
doing nothing. No revert, no removal from a command file two other steps
|
||||
already rewrote.
|
||||
- **adopt costs one constant** — plus the enforcement hook
|
||||
(`hooks/scripts/pre-agent-cap.mjs`) already in place to bound what gets turned
|
||||
on, and the operator-visible cap-exhaustion message already required by
|
||||
Phase 5's exit conditions.
|
||||
|
||||
## 6. Reading the result honestly
|
||||
|
||||
- The gate measures breadth. Say "breadth" in the decision record, not "quality".
|
||||
- Report the excluded count alongside the gain, always. A 35% gain computed
|
||||
after excluding 6 of 10 treatment runs is a finding about instability.
|
||||
- `insufficient-data` is not a decline. Do not resolve it by lowering n.
|
||||
- A verdict computed from a stats file mixing several question sets measures the
|
||||
question sets, not the loop.
|
||||
|
|
@ -20,16 +20,6 @@
|
|||
"args": ["${CLAUDE_PLUGIN_ROOT}/hooks/scripts/pre-write-executor.mjs"]
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"matcher": "WebSearch|WebFetch|Task",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "node",
|
||||
"args": ["${CLAUDE_PLUGIN_ROOT}/hooks/scripts/pre-agent-cap.mjs"]
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"UserPromptSubmit": [
|
||||
|
|
|
|||
|
|
@ -1,211 +0,0 @@
|
|||
#!/usr/bin/env node
|
||||
// Hook: pre-agent-cap.mjs
|
||||
// Event: PreToolUse (WebSearch | WebFetch | Task)
|
||||
// Purpose: Enforce the /trekresearch Phase 5 loop bound at the harness level,
|
||||
// so the cap is a reader that fells rather than prose the model obeys.
|
||||
//
|
||||
// Why this exists: the Phase 5 budget gate (lib/util/research-loop-cap.mjs) is
|
||||
// invoked BY the loop. A gate the caller chooses to consult is advice. The
|
||||
// spike in docs/spike-pretooluse-subagent-reach.md (RESULT: FIRES) established
|
||||
// that a plugin PreToolUse hook does observe tool calls made INSIDE sub-agents
|
||||
// on CC 2.1.226, which is what makes a second, non-optional gate possible.
|
||||
//
|
||||
// Two limits carried over from that spike, neither of which changes the design:
|
||||
// - Reach was measured on one CC version and regressed once before (#34692),
|
||||
// so this hook is defence in depth, never the only gate. research-loop-cap
|
||||
// must stay correct if this hook silently stops firing.
|
||||
// - Whether a blocking (exit 2) decision from inside a sub-agent propagates
|
||||
// usefully was NOT measured — the probe always exited 0 by design.
|
||||
//
|
||||
// Scope key — the property that makes this safe to wire globally:
|
||||
// session_id + a scope marker file that only the Phase 5 loop writes, at
|
||||
// <data root>/trekresearch-loop-scope/<session_id>.json, where the data root
|
||||
// comes from research-loop-cap.mjs's resolveDataRoot() — the same function
|
||||
// the writer resolves through, because a writer and a reader that resolve
|
||||
// the root separately are a hook that enforces nothing while reporting that
|
||||
// it does:
|
||||
// { "runId": "<run id>", "startedAt": "<ISO-8601>" }
|
||||
// No marker for this session => out of scope => allow, unconditionally. An
|
||||
// unrelated session must never be denied because some other run spent its
|
||||
// budget; a PreToolUse hook that over-blocks breaks every session on the box.
|
||||
//
|
||||
// Stated limit, because the guarantee above is about OTHER sessions and reads
|
||||
// as broader than it is: `claude --resume` keeps the same session_id, so a
|
||||
// resumed session is the same session by this key. If a run reached its cap
|
||||
// and then died before removing the marker, the resume inherits the remainder
|
||||
// of the TTL, for any WebSearch/WebFetch/Task — research or not. Three things
|
||||
// bound it rather than close it: only an EXHAUSTED run denies at all (a
|
||||
// part-spent crash leaves no tombstone and is allowed), the TTL is 2h rather
|
||||
// than a working day, and every denial prints the marker path to delete. A
|
||||
// liveness check would close it properly, but a PreToolUse hook has nothing
|
||||
// trustworthy to check liveness against — the marker's writer is a shell
|
||||
// snippet whose $$ is a subshell, not the session.
|
||||
//
|
||||
// Fail-open vs fail-closed, deliberately split:
|
||||
// - Out of scope (no marker, no session_id, unparsable stdin, stale marker,
|
||||
// kill switch, STORM off) => exit 0. Fail OPEN.
|
||||
// - In scope and over budget => exit 2. Fail CLOSED, mirroring
|
||||
// research-loop-cap.mjs's own stance: a budget control that cannot count
|
||||
// must not grant. (The former "CLAUDE_PLUGIN_DATA absent" deny is gone —
|
||||
// the root now always resolves, so that branch could no longer fire.)
|
||||
// - In scope and the ledger cannot be counted (EISDIR, EACCES, EIO — anything
|
||||
// but ENOENT) => exit 2, same reason. This branch used to ALLOW: the hook
|
||||
// carried a private countTurns() whose catch returned 0, so an unreadable
|
||||
// ledger read as "no turns spent". Counting now goes through the
|
||||
// primitive's exported readLedger(), so reader and writer cannot hold
|
||||
// different rules about what an unreadable ledger means.
|
||||
//
|
||||
// Counting is read-only. The ledger is append-only and written solely by
|
||||
// research-loop-cap.mjs's allowTurn(); if this hook appended, the cap would
|
||||
// count its own enforcement.
|
||||
//
|
||||
// Kill switch: VOYAGE_DISABLE_CAP_HOOK=1 disables enforcement entirely.
|
||||
|
||||
import { readFileSync, existsSync, rmSync } from 'node:fs';
|
||||
import { join, dirname } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
|
||||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
const { resolveLedgerPath, resolveDataRoot, resolveMaxConvTurns, isStormEnabled, readLedger, MAX_TOTAL_DIMENSIONS } =
|
||||
await import(join(HERE, '..', '..', 'lib', 'util', 'research-loop-cap.mjs'));
|
||||
|
||||
const SCOPE_DIRNAME = 'trekresearch-loop-scope';
|
||||
// 2h — comfortably longer than any real research run (a 24-turn loop at a couple
|
||||
// of minutes a turn is under an hour), and short enough that debris does not own
|
||||
// the rest of the working day. The TTL is measured from marker.startedAt rather
|
||||
// than from last activity, and `claude --resume` keeps the same session_id, so
|
||||
// this window is what a resumed session can inherit from a run that died holding
|
||||
// the marker. It was 6h; nothing needed six.
|
||||
const DEFAULT_TTL_MS = 2 * 60 * 60 * 1000;
|
||||
|
||||
const env = process.env;
|
||||
|
||||
function allow() {
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
function deny(message) {
|
||||
process.stderr.write(`[voyage] BLOCKED: trekresearch loop cap\n${message}\n`);
|
||||
process.exit(2);
|
||||
}
|
||||
|
||||
// 1. Kill switch.
|
||||
if (env.VOYAGE_DISABLE_CAP_HOOK === '1') allow();
|
||||
|
||||
// 2. Default-off: no loop runs unless STORM is enabled, so nothing to enforce.
|
||||
if (!isStormEnabled(env)) allow();
|
||||
|
||||
// 3. Parse stdin. Unparsable input is not evidence of a loop turn.
|
||||
let input;
|
||||
try {
|
||||
input = JSON.parse(readFileSync(0, 'utf-8'));
|
||||
} catch {
|
||||
allow();
|
||||
}
|
||||
|
||||
const sessionId = input?.session_id;
|
||||
if (!sessionId || typeof sessionId !== 'string') allow();
|
||||
|
||||
// 4. Resolve the scope marker through the writer's own root resolution.
|
||||
// VOYAGE_CAP_SCOPE_DIR stays as a test/override seam; unset, this lands on
|
||||
// exactly the directory the Phase 5 snippet writes into.
|
||||
const scopeDir = env.VOYAGE_CAP_SCOPE_DIR || resolveDataRoot(env);
|
||||
|
||||
const markerPath = join(scopeDir, SCOPE_DIRNAME, `${sessionId}.json`);
|
||||
if (!existsSync(markerPath)) allow();
|
||||
|
||||
let marker;
|
||||
try {
|
||||
marker = JSON.parse(readFileSync(markerPath, 'utf-8'));
|
||||
} catch {
|
||||
allow(); // A marker we cannot read cannot tell us which run we are in.
|
||||
}
|
||||
|
||||
if (!marker?.runId) allow();
|
||||
|
||||
// 5. TTL / auto-reset. A marker left behind by a crashed run must not deny
|
||||
// tool calls for the rest of the machine's life.
|
||||
const ttlRaw = Number(env.VOYAGE_CAP_SCOPE_TTL_MS);
|
||||
const ttlMs = Number.isFinite(ttlRaw) && ttlRaw > 0 ? ttlRaw : DEFAULT_TTL_MS;
|
||||
const startedAt = Date.parse(marker.startedAt ?? '');
|
||||
if (!Number.isFinite(startedAt) || Date.now() - startedAt > ttlMs) {
|
||||
try { rmSync(markerPath, { force: true }); } catch { /* best effort */ }
|
||||
allow();
|
||||
}
|
||||
|
||||
// --- In scope from here on. ---
|
||||
|
||||
// 6. The ledger is the only source of truth for turns spent, and it is counted
|
||||
// through the primitive's OWN readLedger(). This hook used to carry a
|
||||
// private copy of the counting rule whose read error returned 0 — so an
|
||||
// unreadable ledger read as "no turns spent" and ALLOWED, in the one branch
|
||||
// where this hook is supposed to fail closed.
|
||||
const ledgerPath = resolveLedgerPath(env);
|
||||
|
||||
// 7. Same bound the primitive uses: turns-per-dimension × the whole dimension
|
||||
// list under settings.json:16's maxDimensions ceiling.
|
||||
const budget = resolveMaxConvTurns(env) * MAX_TOTAL_DIMENSIONS;
|
||||
|
||||
let ledger;
|
||||
try {
|
||||
ledger = readLedger(ledgerPath, marker.runId);
|
||||
} catch (e) {
|
||||
deny(
|
||||
` Run ${marker.runId} is in scope, but its turn ledger could not be read:\n` +
|
||||
` ${e.message}\n` +
|
||||
` A budget control that cannot count must not grant. Fix or remove the\n` +
|
||||
` ledger, or set VOYAGE_DISABLE_CAP_HOOK=1 to disable enforcement.`,
|
||||
);
|
||||
}
|
||||
|
||||
const toolLine =
|
||||
` Tool: ${input?.tool_name ?? 'unknown'}${input?.agent_type ? ` (agent: ${input.agent_type})` : ''}\n`;
|
||||
|
||||
// Every denial names the marker. If this run is over and the marker outlived it
|
||||
// — the loop's own cleanup covers its three exits, but a crash between the
|
||||
// exhaustion record and the removal runs no cleanup at all — deleting this file
|
||||
// is the remedy, and a resumed session (same session_id) would otherwise sit out
|
||||
// the remaining TTL for work that has nothing to do with research.
|
||||
const remedyLines =
|
||||
` If this loop is not running, the marker is debris — delete it:\n` +
|
||||
` ${markerPath}\n` +
|
||||
` It also auto-resets ${Math.round(ttlMs / 3600000)}h after the run started (VOYAGE_CAP_SCOPE_TTL_MS).\n`;
|
||||
|
||||
// 8. The boundary is the TOMBSTONE, not the count.
|
||||
//
|
||||
// allowTurn() appends before the turn runs, so during the final granted turn the
|
||||
// ledger already holds `budget` records. Denying at `granted >= budget` blocked
|
||||
// that turn's own tool calls — the primitive granted B turns and this hook
|
||||
// permitted B-1 — and it forced every exhausted run out through an exit-2 tool
|
||||
// denial rather than the graceful "cap exhausted" exit, the only exit the prose
|
||||
// at commands/trekresearch.md teaches the model to handle.
|
||||
//
|
||||
// Moving the boundary to `granted > budget` alone would have made this hook
|
||||
// unable to fire at all once the claim mechanism made a breached ledger
|
||||
// impossible — a deny branch that cannot be reached is a dead security claim,
|
||||
// not a backstop. So the primitive records its own denials, and the case this
|
||||
// hook exists for is the one it now catches: the gate said no and a tool call
|
||||
// arrived anyway.
|
||||
if (ledger.exhausted > 0) {
|
||||
deny(
|
||||
` Run ${marker.runId} was already denied a turn by the budget gate\n` +
|
||||
` (${ledger.granted}/${budget} loop turns spent), and this call came after it.\n` +
|
||||
toolLine +
|
||||
` Remaining gaps belong in the brief as open questions, not in another turn.\n` +
|
||||
remedyLines +
|
||||
` Raise TREKRESEARCH_MAX_CONV_TURNS deliberately, or set VOYAGE_DISABLE_CAP_HOOK=1.`,
|
||||
);
|
||||
}
|
||||
|
||||
// 9. Backstop for a ledger that exceeded the bound however it managed to.
|
||||
if (ledger.granted > budget) {
|
||||
deny(
|
||||
` Run ${marker.runId} shows ${ledger.granted} granted turns against a budget of ${budget}.\n` +
|
||||
toolLine +
|
||||
` The ledger has been breached; the loop is over regardless of cause.\n` +
|
||||
remedyLines +
|
||||
` Raise TREKRESEARCH_MAX_CONV_TURNS deliberately, or set VOYAGE_DISABLE_CAP_HOOK=1.`,
|
||||
);
|
||||
}
|
||||
|
||||
allow();
|
||||
|
|
@ -105,16 +105,7 @@ const BLOCK_RULES = [
|
|||
// --- Executor-specific additions ---
|
||||
{
|
||||
name: 'System shutdown/reboot',
|
||||
// Anchored to command position — start of string/line, or after a
|
||||
// separator (`;`, `|`, `&`, `&&`), with optional `sudo` and an optional
|
||||
// absolute path. An unanchored \b match blocked the bare word anywhere,
|
||||
// including quoted grep patterns, heredoc data, and commit messages.
|
||||
//
|
||||
// Runs against commandView, not the whitespace-collapsed string: collapsing
|
||||
// \s+ to ' ' would erase the newline separator before the pattern ever saw
|
||||
// it, and quoted spans must read as data, not as command position.
|
||||
commandView: true,
|
||||
pattern: /(?:^|[\n;|&])\s*(?:sudo\s+(?:-[a-zA-Z]+\s+)*)?(?:[\w./-]*\/)?(?:shutdown|reboot|halt|poweroff)\b/,
|
||||
pattern: /\b(?:shutdown|reboot|halt|poweroff)\b/,
|
||||
description: 'System shutdown/reboot commands are blocked during execution.',
|
||||
},
|
||||
{
|
||||
|
|
@ -207,59 +198,6 @@ function normalizeCommand(cmd) {
|
|||
.trim();
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Command-position view — for rules that must distinguish a command from data
|
||||
// that merely names one. Whitespace is NOT collapsed, so newline stays a
|
||||
// separator. Quoted spans and heredoc bodies become data; the argument of a
|
||||
// shell wrapper stays a command; backslash-escaped names are seen through.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
// `out` ends with the `-c` of a shell invocation — the next quoted span is a
|
||||
// command string, not data. Optional leading `sudo` and absolute path.
|
||||
const SHELL_C_TAIL =
|
||||
/(?:^|[\s;|&])(?:sudo\s+(?:-[a-zA-Z]+\s+)*)?(?:[\w./-]*\/)?(?:ba|z|k|da|a)?sh\s+(?:-[a-zA-Z]+\s+)*-c\s*$/;
|
||||
|
||||
// Drop heredoc bodies, keeping the operator line. Their newlines are not
|
||||
// command separators, and without this every heredoc line that happens to
|
||||
// start with a matched word reads as command position. Runs before the quote
|
||||
// scan, since a body may contain quotes that would desync it.
|
||||
function stripHeredocBodies(cmd) {
|
||||
return cmd.replace(
|
||||
/(<<-?\s*(['"]?)(\w+)\2[^\n]*\n)[\s\S]*?(?:\n[ \t]*\3[ \t]*(?=\n|$)|$)/g,
|
||||
(_match, head) => head,
|
||||
);
|
||||
}
|
||||
|
||||
function commandPositionView(cmd) {
|
||||
const src = stripHeredocBodies(cmd);
|
||||
let out = '';
|
||||
let i = 0;
|
||||
while (i < src.length) {
|
||||
const ch = src[i];
|
||||
if (ch === "'" || ch === '"') {
|
||||
const close = src.indexOf(ch, i + 1);
|
||||
const inner = close === -1 ? src.slice(i + 1) : src.slice(i + 1, close);
|
||||
// Unterminated quote — treat the remainder as one span and stop.
|
||||
out += SHELL_C_TAIL.test(out) ? `;${inner};` : ' ';
|
||||
i = close === -1 ? src.length : close + 1;
|
||||
} else {
|
||||
out += ch;
|
||||
i += 1;
|
||||
}
|
||||
}
|
||||
return (
|
||||
out
|
||||
// `xargs [flags] <cmd>` puts <cmd> at command position with no separator
|
||||
// in front of it. Flags taking a separate argument (`-I {}`) are not
|
||||
// parsed — the separator lands before the argument, not the command.
|
||||
.replace(/\bxargs((?:\s+-[a-zA-Z0-9-]+)*)/g, 'xargs$1 ;')
|
||||
// `\name` runs name — the backslash only suppresses alias expansion.
|
||||
// normalizeBashExpansion covers the between-word-chars case; this covers
|
||||
// a backslash at command position.
|
||||
.replace(/\\(\w)/g, '$1')
|
||||
);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Main
|
||||
// ---------------------------------------------------------------------------
|
||||
|
|
@ -281,11 +219,10 @@ if (!command || typeof command !== 'string') {
|
|||
// Strip bash evasion, then normalize whitespace
|
||||
const deobfuscated = normalizeBashExpansion(command);
|
||||
const normalized = normalizeCommand(deobfuscated);
|
||||
const commandView = commandPositionView(deobfuscated.replace(/\x1B\[[0-9;]*m/g, ''));
|
||||
|
||||
// Check BLOCK rules first
|
||||
for (const rule of BLOCK_RULES) {
|
||||
if (rule.pattern.test(rule.commandView ? commandView : normalized)) {
|
||||
if (rule.pattern.test(normalized)) {
|
||||
process.stderr.write(
|
||||
`[voyage] BLOCKED: ${rule.name}\n` +
|
||||
` Command: ${normalized.slice(0, 200)}${normalized.length > 200 ? '...' : ''}\n` +
|
||||
|
|
|
|||
|
|
@ -25,27 +25,11 @@ const TREKBRIEF_ALLOWED = Object.freeze(new Set([
|
|||
]));
|
||||
|
||||
// Source: tests/fixtures/jsonl-schemas.md row 2 (trekresearch)
|
||||
// `engine` is a low-cardinality label (swarm|deep-research) emitted at
|
||||
// commands/trekresearch.md:533 and promised in prose (:570-572).
|
||||
// DENY BY OMISSION: question (free prose), project_dir + brief_path
|
||||
// (filesystem paths) are written into the jsonl but MUST NOT reach the
|
||||
// exporter.
|
||||
// The five v5.10 measurement fields are allowlisted too: `effort` is a
|
||||
// low-cardinality label (low|standard|high) and the grouping key the
|
||||
// measurement gate is computed on; `unique_sources`, `dimensions_baseline`,
|
||||
// `conv_turns` and `empty_turns` are plain counters. None of them carry prose
|
||||
// or paths. `dimensions_baseline_preserved` is a boolean, and it exists BECAUSE
|
||||
// prose is denied here: the activation check needs to know that the final
|
||||
// dimension list is a superset of the interview-derived one, and the dimension
|
||||
// NAMES that would show it directly are free prose that must not reach the
|
||||
// exporter. A boolean attestation carries the fact without the payload.
|
||||
const TREKRESEARCH_ALLOWED = Object.freeze(new Set([
|
||||
'ts', 'slug', 'mode', 'scope', 'engine', 'dimensions', 'agents_local',
|
||||
'ts', 'slug', 'mode', 'scope', 'dimensions', 'agents_local',
|
||||
'agents_external', 'gemini_used', 'confidence', 'contradictions',
|
||||
'open_questions', 'profile', 'parallel_agents',
|
||||
'external_research_enabled', 'profile_source',
|
||||
'effort', 'unique_sources', 'dimensions_baseline', 'conv_turns',
|
||||
'empty_turns', 'dimensions_baseline_preserved',
|
||||
]));
|
||||
|
||||
// Source: tests/fixtures/jsonl-schemas.md row 3 (trekplan)
|
||||
|
|
|
|||
|
|
@ -32,7 +32,7 @@ const OPTIONAL_KEYS = [
|
|||
const OPTIONAL_BOOLEAN_KEYS = new Set(OPTIONAL_KEYS);
|
||||
|
||||
// Optional string-typed manifest keys (v4.1 Step 3 — additive forward-compat).
|
||||
// `profile_used`: name of the model profile (economy|balanced|premium|fable|<custom>) the
|
||||
// `profile_used`: name of the model profile (economy|balanced|premium|<custom>) the
|
||||
// step was executed under. Absence is fine (v4.0 manifests have no
|
||||
// profile concept); presence MUST be a string.
|
||||
// Unlike OPTIONAL_BOOLEAN_KEYS, absence is NOT defaulted — the field is simply
|
||||
|
|
|
|||
|
|
@ -1,21 +0,0 @@
|
|||
---
|
||||
profile_version: "1.0"
|
||||
name: fable
|
||||
phase_models:
|
||||
- phase: brief
|
||||
model: fable
|
||||
- phase: research
|
||||
model: fable
|
||||
- phase: plan
|
||||
model: fable
|
||||
- phase: execute
|
||||
model: fable
|
||||
- phase: review
|
||||
model: fable
|
||||
- phase: continue
|
||||
model: fable
|
||||
parallel_agents_min: 6
|
||||
parallel_agents_max: 8
|
||||
external_research_enabled: true
|
||||
brief_reviewer_iter_cap: 3
|
||||
---
|
||||
|
|
@ -67,11 +67,6 @@ export function resolvePhaseSignalFromFile(briefPath, phase) {
|
|||
}
|
||||
|
||||
// CLI shim — mirrors lib/validators/brief-validator.mjs:168 pattern.
|
||||
// Footgun guard (v5.9): this shim's `model` output is brief-signal-only — it
|
||||
// never consults the profile layer. For command wiring, the composed resolver
|
||||
// CLI (`resolver.mjs --resolve-phase-model`, brief > profile > default) is the
|
||||
// single resolution source for {effort, model}. Do not re-wire commands/*.md
|
||||
// Bash blocks back to this shim.
|
||||
if (import.meta.url === `file://${process.argv[1]}`) {
|
||||
const args = process.argv.slice(2);
|
||||
const getArg = (name) => {
|
||||
|
|
|
|||
|
|
@ -45,7 +45,7 @@ import { resolvePhaseSignal } from './phase-signal-resolver.mjs';
|
|||
|
||||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||
const BUILTIN_PROFILES_DIR = __dirname; // lib/profiles/
|
||||
const BUILTIN_NAMES = new Set(['economy', 'balanced', 'premium', 'fable']);
|
||||
const BUILTIN_NAMES = new Set(['economy', 'balanced', 'premium']);
|
||||
|
||||
/**
|
||||
* Resolve the path to a profile file.
|
||||
|
|
@ -221,13 +221,7 @@ export function validateProfileFile(path, opts = {}) {
|
|||
* @param {string|null} briefPath Absolute or repo-relative path to brief.md, or null
|
||||
* @param {string[]|object} argv Full process.argv array OR parsed flags object
|
||||
* @param {object} [env] Environment-variable record (defaults to process.env)
|
||||
* @returns {{effort?: string, model: string, source: 'brief-signal'|'flag'|'env'|'default'}}
|
||||
*
|
||||
* `effort` (v5.9 ADDITIVE) is the brief signal's effort passed through when
|
||||
* present — so commands consume ONE coherent {effort, model, source} result
|
||||
* instead of two split CLI calls. Absent when the brief carries no valid
|
||||
* effort signal for the phase (commands default to 'standard' per the
|
||||
* composition rule).
|
||||
* @returns {{model: string, source: 'brief-signal'|'flag'|'env'|'default'}}
|
||||
*
|
||||
* Error handling contract:
|
||||
* - Never throws. Any failure (ENOENT on briefPath, malformed YAML, missing
|
||||
|
|
@ -240,10 +234,7 @@ export function validateProfileFile(path, opts = {}) {
|
|||
* directly; commands must inject {resolved model} at Agent-tool spawn sites.
|
||||
*/
|
||||
export function resolvePhaseModel(phase, briefPath, argv, env = process.env) {
|
||||
// Step 1: brief-signal lookup. `effort` is captured independently of `model`
|
||||
// so a signal like {effort: high} (no model) still passes effort through
|
||||
// while the model falls to the profile layer.
|
||||
let effort;
|
||||
// Step 1: brief-signal lookup
|
||||
if (typeof briefPath === 'string' && briefPath.length > 0 && existsSync(briefPath)) {
|
||||
let fm = null;
|
||||
try {
|
||||
|
|
@ -255,11 +246,8 @@ export function resolvePhaseModel(phase, briefPath, argv, env = process.env) {
|
|||
}
|
||||
if (fm) {
|
||||
const signal = resolvePhaseSignal(fm, phase);
|
||||
if (signal && typeof signal.effort === 'string') effort = signal.effort;
|
||||
if (signal && typeof signal.model === 'string' && signal.model.length > 0) {
|
||||
return effort !== undefined
|
||||
? { effort, model: signal.model, source: 'brief-signal' }
|
||||
: { model: signal.model, source: 'brief-signal' };
|
||||
return { model: signal.model, source: 'brief-signal' };
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -290,9 +278,7 @@ export function resolvePhaseModel(phase, briefPath, argv, env = process.env) {
|
|||
}
|
||||
}
|
||||
const model = phaseModels[phase] || 'opus';
|
||||
return effort !== undefined
|
||||
? { effort, model, source: profile_source }
|
||||
: { model, source: profile_source };
|
||||
return { model, source: profile_source };
|
||||
}
|
||||
|
||||
// CLI shim — invoked by commands/trek*.md via Bash.
|
||||
|
|
@ -314,8 +300,7 @@ if (import.meta.url === `file://${process.argv[1]}`) {
|
|||
if (args.includes('--json')) {
|
||||
process.stdout.write(JSON.stringify(r) + '\n');
|
||||
} else {
|
||||
const effort = 'effort' in r ? ` effort=${r.effort}` : '';
|
||||
process.stdout.write(`model=${r.model} source=${r.source}${effort}\n`);
|
||||
process.stdout.write(`model=${r.model} source=${r.source}\n`);
|
||||
}
|
||||
process.exit(0);
|
||||
}
|
||||
|
|
|
|||
Binary file not shown.
|
|
@ -27,7 +27,6 @@ import { join, dirname } from 'node:path';
|
|||
// Per-Mtok USD prices, resolved 2026-06-26 via the claude-api skill reference:
|
||||
// base input/output from the model table; cache rates from the prompt-caching
|
||||
// doc multipliers (cache_read 0.1x, write_5m 1.25x, write_1h 2.0x of input).
|
||||
// claude-fable-5 resolved 2026-07-02 from the official platform pricing docs.
|
||||
export const PRICE_TABLE = Object.freeze({
|
||||
'claude-opus-4-8': Object.freeze({
|
||||
input: 5.0,
|
||||
|
|
@ -36,17 +35,10 @@ export const PRICE_TABLE = Object.freeze({
|
|||
cache_write_5m: 6.25,
|
||||
cache_write_1h: 10.0,
|
||||
}),
|
||||
'claude-fable-5': Object.freeze({
|
||||
input: 10.0,
|
||||
output: 50.0,
|
||||
cache_read: 1.0,
|
||||
cache_write_5m: 12.5,
|
||||
cache_write_1h: 20.0,
|
||||
}),
|
||||
});
|
||||
|
||||
// Date the PRICE_TABLE values were resolved/verified. Bump when prices change.
|
||||
export const PRICE_TABLE_VERSION = '2026-07-02';
|
||||
export const PRICE_TABLE_VERSION = '2026-06-26';
|
||||
|
||||
function num(v) {
|
||||
return typeof v === 'number' && Number.isFinite(v) ? v : 0;
|
||||
|
|
|
|||
|
|
@ -1,350 +0,0 @@
|
|||
// lib/util/research-loop-cap.mjs
|
||||
// Stateful, default-off cost cap for the /trekresearch bounded conversation
|
||||
// loop (Phase 4.5 dimension discovery + Phase 5 loop turns).
|
||||
//
|
||||
// Three properties the plan review required:
|
||||
// (a) Default-off — VOYAGE_STORM_ENABLED must be '1'; otherwise the budget
|
||||
// is 0 regardless of effort. This IS the decline branch: doing nothing
|
||||
// leaves the mechanism off, and adopt is flipping this one constant.
|
||||
// (b) The cap counts itself — allowTurn() derives used-turn count from an
|
||||
// append-only JSONL ledger, never from a caller-supplied number. A cap
|
||||
// that asks the caller how many turns it has used is not a cap. Each
|
||||
// grant additionally claims a turn SLOT with O_EXCL, so the bound holds
|
||||
// when several callers decide at once instead of only when they queue.
|
||||
// (c) Correct size bound — worst case is max_conv_turns × max_total_dimensions,
|
||||
// where max_total_dimensions is the WHOLE list (interview + discovered)
|
||||
// under settings.json:16's cap of 8 — not × discovered-only.
|
||||
//
|
||||
// CLAUDE_PLUGIN_DATA absent => fall back to ~/.claude/voyage. The variable is
|
||||
// EMPTY in the Bash tool's process env (measured in a live plugin-enabled
|
||||
// session), and the Phase 5 bash snippet is this module's only caller — so
|
||||
// denying on its absence denied turn 1 of every real run. The root is resolved
|
||||
// in code rather than demanded of the environment, and hooks/scripts/
|
||||
// pre-agent-cap.mjs resolves it through the SAME function, so the writer and
|
||||
// the reader can never disagree about where the ledger lives.
|
||||
//
|
||||
// The fail-closed stance covers both directions of ledger IO: a ledger that
|
||||
// cannot be WRITTEN denies the turn, and a ledger that exists but cannot be
|
||||
// READ denies it too. Only ENOENT counts as zero turns spent, because that is
|
||||
// the legitimate first-turn state. This module is a budget control, not
|
||||
// telemetry — the opposite of lib/stats/event-emit.mjs's fail-open.
|
||||
//
|
||||
// CLI shim, two modes:
|
||||
// node lib/util/research-loop-cap.mjs --run-id ID --dimension D --effort E
|
||||
// → JSON: { ok, used, budget, reason? } (exit 0 = granted, exit 1 = denied)
|
||||
//
|
||||
// node lib/util/research-loop-cap.mjs --check-dimensions N
|
||||
// → JSON: { ok, count, ceiling, reason? } (exit 0 = within, exit 1 = rejected)
|
||||
// The second axis of the bounded-cost NFR. Phase 4.5 never calls the budget
|
||||
// gate, so this mode requires no run id, effort or STORM flag — but it reads
|
||||
// the SAME MAX_TOTAL_DIMENSIONS the budget is sized against.
|
||||
|
||||
import { existsSync, mkdirSync, appendFileSync, readFileSync, writeFileSync } from 'node:fs';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
|
||||
export const MAX_CONV_TURNS = 3;
|
||||
export const MAX_TOTAL_DIMENSIONS = 8; // settings.json:16 maxDimensions — whole list, not discovered-only
|
||||
|
||||
const LEDGER_FILENAME = 'trekresearch-loop-ledger.jsonl';
|
||||
const CLAIM_DIRNAME = 'trekresearch-loop-claims';
|
||||
|
||||
export function isStormEnabled(env = process.env) {
|
||||
return env.VOYAGE_STORM_ENABLED === '1';
|
||||
}
|
||||
|
||||
/**
|
||||
* The OTHER cost ceiling: how large the whole dimension list may get after
|
||||
* Phase 4.5 discovery has appended to it.
|
||||
*
|
||||
* The bounded-cost NFR asks for explicit ceilings on both axes. The turn axis
|
||||
* had a constant, a ledger-backed reader and a PreToolUse enforcer; the
|
||||
* discovery axis had only a sentence in Phase 4.5 prose — a cap nothing reads,
|
||||
* which is the failure mode the operator decision on brief_reviewer_iter_cap
|
||||
* warned about. This is the reader.
|
||||
*
|
||||
* The ceiling is MAX_TOTAL_DIMENSIONS on purpose: the value that sizes the turn
|
||||
* budget IS settings.json:16's maxDimensions, and a second constant for the same
|
||||
* number is how two readers end up enforcing different bounds.
|
||||
*
|
||||
* Accepts the dimension list or its count, because Phase 4.5 has the list and
|
||||
* the CLI has a number. A count that cannot be read is REJECTED — a cost ceiling
|
||||
* that waves through what it cannot measure is not a ceiling.
|
||||
*
|
||||
* @param {string[]|number|string} dimensions
|
||||
* @param {{ceiling?: number}} [opts]
|
||||
* @returns {{ok: boolean, count: number|null, ceiling: number, reason?: string}}
|
||||
*/
|
||||
export function checkDimensionCeiling(dimensions, opts = {}) {
|
||||
const ceiling = Number.isFinite(opts.ceiling) ? opts.ceiling : MAX_TOTAL_DIMENSIONS;
|
||||
const count = Array.isArray(dimensions) ? dimensions.length : Number(dimensions);
|
||||
if (dimensions === null || dimensions === undefined || !Number.isInteger(count) || count < 0) {
|
||||
return { ok: false, count: null, ceiling, reason: 'unreadable_dimension_count' };
|
||||
}
|
||||
if (count > ceiling) {
|
||||
return { ok: false, count, ceiling, reason: 'ceiling_exceeded' };
|
||||
}
|
||||
return { ok: true, count, ceiling };
|
||||
}
|
||||
|
||||
/**
|
||||
* Coerce TREKRESEARCH_MAX_CONV_TURNS. NaN, empty, negative, zero, Infinity, or
|
||||
* any fraction that floors below 1 all fall back to MAX_CONV_TURNS — never to
|
||||
* unbounded, and never to 0.
|
||||
*
|
||||
* The floor is applied BEFORE the `<= 0` guard, not after. Flooring afterwards
|
||||
* let '0.5' and '0.9' clear a guard written against the raw value and then
|
||||
* become 0, making the budget 0 × MAX_TOTAL_DIMENSIONS = 0: every turn denied
|
||||
* and the loop silently dead rather than bounded. A cap of 0 is not a narrower
|
||||
* cap, it is an off switch that the documented fallback promises not to be.
|
||||
*/
|
||||
export function resolveMaxConvTurns(env = process.env) {
|
||||
const raw = env.TREKRESEARCH_MAX_CONV_TURNS;
|
||||
if (raw === undefined || raw === null || raw === '') return MAX_CONV_TURNS;
|
||||
const n = Math.floor(Number(raw));
|
||||
if (!Number.isFinite(n) || n <= 0) return MAX_CONV_TURNS;
|
||||
return n;
|
||||
}
|
||||
|
||||
/**
|
||||
* The one data root for everything this loop writes: the turn ledger and the
|
||||
* PreToolUse scope marker. CLAUDE_PLUGIN_DATA when the harness provides it,
|
||||
* ~/.claude/voyage when it does not — which is the case in every Bash tool
|
||||
* invocation today.
|
||||
*/
|
||||
export function resolveDataRoot(env = process.env) {
|
||||
const dir = env.CLAUDE_PLUGIN_DATA;
|
||||
if (dir && typeof dir === 'string' && dir.length > 0) return dir;
|
||||
const home = env.HOME && env.HOME.length > 0 ? env.HOME : homedir();
|
||||
return join(home, '.claude', 'voyage');
|
||||
}
|
||||
|
||||
export function resolveLedgerPath(env = process.env) {
|
||||
return join(resolveDataRoot(env), LEDGER_FILENAME);
|
||||
}
|
||||
|
||||
/**
|
||||
* Where the per-turn claim files live. A claim is the ATOMIC record that a turn
|
||||
* slot is taken; the ledger is the readable record of what that turn was for.
|
||||
*
|
||||
* The claim exists because the ledger alone cannot bound the loop. Counting the
|
||||
* ledger and then appending is read-then-write: N callers that all observe
|
||||
* `used == budget - 1` all decide to grant, and the bound is exceeded by N-1 —
|
||||
* precisely the concurrent case (several agents spawned in one message) that
|
||||
* allowTurn's own comment named as the reason it had to be append-only.
|
||||
*
|
||||
* Claim files are empty, at most `budget` per run, and never cleaned up — the
|
||||
* same standing as the ledger itself, which also grows for the life of the data
|
||||
* root. Two consequences worth stating rather than discovering: reusing a
|
||||
* runId across runs finds its slots already taken and denies, and two runIds
|
||||
* that collide after filename sanitisation block each other. Both err toward
|
||||
* denying a turn, which is the safe direction for a budget control.
|
||||
*/
|
||||
export function resolveClaimDir(env = process.env) {
|
||||
return join(resolveDataRoot(env), CLAIM_DIRNAME);
|
||||
}
|
||||
|
||||
function claimFileName(runId, slot) {
|
||||
return `${String(runId).replace(/[^A-Za-z0-9._-]/g, '_')}-${slot}.claim`;
|
||||
}
|
||||
|
||||
/**
|
||||
* Try to take turn slot `slot` for `runId`. `wx` is O_CREAT|O_EXCL: the kernel
|
||||
* decides the winner, so exactly one caller can ever create a given slot file.
|
||||
*
|
||||
* @returns {boolean} true when this caller took the slot, false when it was already taken
|
||||
* @throws on any IO error other than EEXIST — the caller turns that into a denial
|
||||
*/
|
||||
function claimSlot(claimDir, runId, slot) {
|
||||
try {
|
||||
writeFileSync(join(claimDir, claimFileName(runId, slot)), '', { flag: 'wx' });
|
||||
return true;
|
||||
} catch (e) {
|
||||
if (e && e.code === 'EEXIST') return false;
|
||||
throw e;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Read one run's turn count off the append-only ledger.
|
||||
*
|
||||
* ENOENT is 0 turns spent — the legitimate first-turn state, and the reason
|
||||
* this cannot simply throw on every read failure. Every OTHER read error
|
||||
* (EISDIR, EACCES, EIO) THROWS, because returning 0 from an unreadable ledger
|
||||
* re-granted the full budget on every call: unbounded, and the exact
|
||||
* silently-grant-unlimited failure this module's header argues against three
|
||||
* lines above the code that did it. The missing-directory case already failed
|
||||
* closed; this makes the unreadable-file case agree with it.
|
||||
*
|
||||
* The `existsSync` pre-check is deliberately gone: readFileSync's own ENOENT
|
||||
* carries the same information without a second syscall that can disagree with
|
||||
* the read that follows it.
|
||||
*
|
||||
* Exported so hooks/scripts/pre-agent-cap.mjs counts through this exact
|
||||
* function. A reader and a writer with private copies of the counting rule are
|
||||
* how a hook ends up enforcing a different bound than the gate it backs.
|
||||
*
|
||||
* Two counts, deliberately separate. `granted` is turns handed out. `exhausted`
|
||||
* is tombstones — records this gate wrote when it DENIED a turn. A tombstone is
|
||||
* not a turn and must never consume budget; it exists so the PreToolUse hook can
|
||||
* tell "turn B is in flight" (granted == budget, no tombstone) apart from "the
|
||||
* gate already said no and something kept going" (tombstone present).
|
||||
*
|
||||
* @param {string} ledgerPath
|
||||
* @param {string} runId
|
||||
* @returns {{granted: number, exhausted: number}}
|
||||
* @throws when the ledger exists but cannot be read
|
||||
*/
|
||||
export function readLedger(ledgerPath, runId) {
|
||||
let text;
|
||||
try {
|
||||
text = readFileSync(ledgerPath, 'utf-8');
|
||||
} catch (e) {
|
||||
if (e && e.code === 'ENOENT') return { granted: 0, exhausted: 0 };
|
||||
const err = new Error(`ledger unreadable at ${ledgerPath}: ${e.message}`);
|
||||
err.code = 'VOYAGE_LEDGER_UNREADABLE';
|
||||
throw err;
|
||||
}
|
||||
let granted = 0;
|
||||
let exhausted = 0;
|
||||
for (const line of text.split('\n')) {
|
||||
if (!line) continue;
|
||||
try {
|
||||
const rec = JSON.parse(line);
|
||||
if (rec.runId !== runId) continue;
|
||||
if (rec.exhausted === true) exhausted++;
|
||||
else granted++;
|
||||
} catch { /* skip malformed lines */ }
|
||||
}
|
||||
return { granted, exhausted };
|
||||
}
|
||||
|
||||
/**
|
||||
* Record that this run has been denied a turn for budget.
|
||||
*
|
||||
* Best effort on purpose: the denial itself is already the correct answer, so a
|
||||
* ledger that cannot take the tombstone must not turn a denial into a grant. The
|
||||
* tombstone only strengthens the harness-level backstop.
|
||||
*/
|
||||
function markExhausted(ledgerPath, runId, now) {
|
||||
try {
|
||||
appendFileSync(ledgerPath, JSON.stringify({ ts: now.toISOString(), runId, exhausted: true }) + '\n');
|
||||
} catch { /* best effort — see above */ }
|
||||
}
|
||||
|
||||
/**
|
||||
* Decide whether one more research-loop turn may run.
|
||||
*
|
||||
* Phase 4.5/5 may spawn several agents in a single message, so the decision has
|
||||
* to survive concurrent callers. It does that by CLAIMING a turn slot with
|
||||
* O_EXCL (see claimSlot) and only then appending to the ledger. The comment
|
||||
* that used to sit here asserted "append-only: never read-modify-write" as if
|
||||
* appending were itself the concurrency guarantee — but the decision path was
|
||||
* count-then-append, which is read-then-write, so the claim was unsupported by
|
||||
* the code beneath it. The kernel now picks the winner for each slot.
|
||||
*
|
||||
* @param {{runId: string, dimension: string, effort: string}} args
|
||||
* @param {{env?: object, now?: Date}} [opts]
|
||||
* @returns {{ok: boolean, used: number, budget: number, reason?: string}}
|
||||
*/
|
||||
export function allowTurn({ runId, dimension, effort } = {}, opts = {}) {
|
||||
const env = opts.env || process.env;
|
||||
const now = opts.now || new Date();
|
||||
|
||||
if (!isStormEnabled(env)) {
|
||||
return { ok: false, used: 0, budget: 0, reason: 'storm_disabled' };
|
||||
}
|
||||
if (effort !== 'high') {
|
||||
return { ok: false, used: 0, budget: 0, reason: 'effort_not_high' };
|
||||
}
|
||||
if (!runId || !dimension) {
|
||||
return { ok: false, used: 0, budget: 0, reason: 'missing_args' };
|
||||
}
|
||||
|
||||
const maxConvTurns = resolveMaxConvTurns(env);
|
||||
const budget = maxConvTurns * MAX_TOTAL_DIMENSIONS;
|
||||
|
||||
const ledgerPath = resolveLedgerPath(env);
|
||||
let ledger;
|
||||
try {
|
||||
ledger = readLedger(ledgerPath, runId);
|
||||
} catch (e) {
|
||||
return { ok: false, used: 0, budget, reason: `ledger-read-failed: ${e.message}` };
|
||||
}
|
||||
const used = ledger.granted;
|
||||
|
||||
// Already tombstoned: this run is over. Short-circuit so a hammered gate
|
||||
// neither walks every slot again nor appends a second tombstone.
|
||||
if (ledger.exhausted > 0) {
|
||||
return { ok: false, used, budget, reason: 'budget_exhausted' };
|
||||
}
|
||||
// Claim a turn SLOT before spending anything. The ledger count only says
|
||||
// where to start looking; the claim is what makes the grant exclusive. Slot
|
||||
// numbers are bounded by `budget`, and each can be created exactly once, so
|
||||
// the total number of grants for a run can never exceed the budget however
|
||||
// many callers arrive at once.
|
||||
const claimDir = resolveClaimDir(env);
|
||||
let slot = used + 1;
|
||||
let claimed = false;
|
||||
try {
|
||||
mkdirSync(claimDir, { recursive: true });
|
||||
while (slot <= budget) {
|
||||
if (claimSlot(claimDir, runId, slot)) { claimed = true; break; }
|
||||
slot++;
|
||||
}
|
||||
} catch (e) {
|
||||
return { ok: false, used, budget, reason: `claim-failed: ${e.message}` };
|
||||
}
|
||||
if (!claimed) {
|
||||
markExhausted(ledgerPath, runId, now);
|
||||
return { ok: false, used, budget, reason: 'budget_exhausted' };
|
||||
}
|
||||
|
||||
try {
|
||||
const dir = dirname(ledgerPath);
|
||||
if (!existsSync(dir)) mkdirSync(dir, { recursive: true });
|
||||
appendFileSync(ledgerPath, JSON.stringify({ ts: now.toISOString(), runId, dimension, effort, slot }) + '\n');
|
||||
} catch (e) {
|
||||
return { ok: false, used, budget, reason: `ledger-write-failed: ${e.message}` };
|
||||
}
|
||||
|
||||
return { ok: true, used: slot, budget };
|
||||
}
|
||||
|
||||
// ---- CLI shim ----------------------------------------------------------------
|
||||
|
||||
function parseArgs(argv) {
|
||||
const out = {};
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
const a = argv[i];
|
||||
if (a === '--run-id') out.runId = argv[++i];
|
||||
else if (a === '--dimension') out.dimension = argv[++i];
|
||||
else if (a === '--effort') out.effort = argv[++i];
|
||||
else if (a === '--check-dimensions') out.checkDimensions = argv[++i];
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
if (import.meta.url === `file://${process.argv[1]}`) {
|
||||
const args = parseArgs(process.argv.slice(2));
|
||||
|
||||
// The dimension ceiling is a Phase 4.5 concern, and Phase 4.5 never calls the
|
||||
// budget gate — so this branch must not inherit the gate's preconditions
|
||||
// (run id, effort, STORM flag). It is a pure bound on list size.
|
||||
if (args.checkDimensions !== undefined) {
|
||||
const result = checkDimensionCeiling(args.checkDimensions);
|
||||
process.stdout.write(JSON.stringify(result) + '\n');
|
||||
process.exit(result.ok ? 0 : 1);
|
||||
}
|
||||
|
||||
if (!args.runId || !args.dimension || !args.effort) {
|
||||
process.stdout.write(JSON.stringify({
|
||||
ok: false,
|
||||
reason: 'usage: research-loop-cap.mjs --run-id ID --dimension D --effort standard|high|low',
|
||||
}) + '\n');
|
||||
process.exit(1);
|
||||
}
|
||||
const result = allowTurn(args);
|
||||
process.stdout.write(JSON.stringify(result) + '\n');
|
||||
process.exit(result.ok ? 0 : 1);
|
||||
}
|
||||
|
|
@ -1,16 +1,10 @@
|
|||
// lib/util/test-census.mjs
|
||||
// Census of the test suite: split top-level test() declarations into three
|
||||
// honest categories:
|
||||
// - "behavior" — ordinary behavior coverage (the default bucket).
|
||||
// - "docPins" — doc-consistency pins (string/existence assertions that pin
|
||||
// documentation against a source-of-truth).
|
||||
// - "goldEval" — the offline gold-scored output eval scoring RUN (SKAL-1·4b):
|
||||
// neither behavior nor prose-pin, but a run that scores a
|
||||
// committed agent-run fixture against the golden corpus.
|
||||
// Makes the cited test count honest — a prose-pin is not the same coverage as a
|
||||
// behavior test, and a scoring run is a third thing again, so a single
|
||||
// Census of the test suite: split top-level test() declarations into
|
||||
// "behavior" tests vs "doc-consistency pins" (string/existence assertions that
|
||||
// pin documentation against a source-of-truth). Makes the cited test count
|
||||
// honest — a prose-pin is not the same coverage as a behavior test, so a single
|
||||
// conflated total oversells behavior coverage.
|
||||
// Devil's-advocate audit §Top changes #8 (S19); third bucket added in SKAL-1·4b.
|
||||
// Devil's-advocate audit §Top changes #8 (S19).
|
||||
//
|
||||
// Metric: top-level `test(` declarations (static, deterministic, in-process).
|
||||
// This is distinct from node:test's runtime total, which additionally counts
|
||||
|
|
@ -24,11 +18,6 @@ import { join } from 'node:path';
|
|||
// is bucketed correctly without editing this module.
|
||||
export const PIN_FILE_RE = /(doc-consistency|prose-pins)\.test\.mjs$/;
|
||||
|
||||
// Files whose tests are the offline gold-scored output eval scoring run
|
||||
// (SKAL-1·4b). Kept as a regex (not a single filename) so a future split-out
|
||||
// is bucketed correctly without editing this module.
|
||||
export const GOLD_EVAL_FILE_RE = /gold-eval\.test\.mjs$/;
|
||||
|
||||
function walk(dir) {
|
||||
const out = [];
|
||||
for (const e of readdirSync(dir, { withFileTypes: true })) {
|
||||
|
|
@ -43,20 +32,18 @@ function countTests(file) {
|
|||
return (readFileSync(file, 'utf-8').match(/^\s*test\(/gm) || []).length;
|
||||
}
|
||||
|
||||
// Returns { behavior, docPins, goldEval, total, byFile } for all *.test.mjs
|
||||
// under testsRoot. behavior + docPins + goldEval === total by construction.
|
||||
// Returns { behavior, docPins, total, byFile } for all *.test.mjs under
|
||||
// testsRoot. behavior + docPins === total by construction.
|
||||
export function censusTests(testsRoot) {
|
||||
const files = walk(testsRoot).sort();
|
||||
const byFile = {};
|
||||
let behavior = 0;
|
||||
let docPins = 0;
|
||||
let goldEval = 0;
|
||||
for (const f of files) {
|
||||
const n = countTests(f);
|
||||
byFile[f] = n;
|
||||
if (PIN_FILE_RE.test(f)) docPins += n;
|
||||
else if (GOLD_EVAL_FILE_RE.test(f)) goldEval += n;
|
||||
else behavior += n;
|
||||
}
|
||||
return { behavior, docPins, goldEval, total: behavior + docPins + goldEval, byFile };
|
||||
return { behavior, docPins, total: behavior + docPins, byFile };
|
||||
}
|
||||
|
|
|
|||
|
|
@ -252,11 +252,7 @@ if (import.meta.url === `file://${process.argv[1]}`) {
|
|||
const minIdx = args.indexOf('--min-version');
|
||||
const minBriefVersion = minIdx >= 0 ? args[minIdx + 1] : undefined;
|
||||
// filePath is the first positional, skipping the --min-version value token.
|
||||
// Guard: when --min-version is absent (minIdx === -1) the skip index must be -1,
|
||||
// not 0 — otherwise the no-flag invocation `brief-validator.mjs <brief.md>` drops
|
||||
// the file (which sits at index 0) and bails to Usage.
|
||||
const skipIdx = minIdx >= 0 ? minIdx + 1 : -1;
|
||||
const filePath = args.find((a, i) => !a.startsWith('--') && i !== skipIdx);
|
||||
const filePath = args.find((a, i) => !a.startsWith('--') && i !== minIdx + 1);
|
||||
if (!filePath) {
|
||||
process.stderr.write('Usage: brief-validator.mjs [--soft] [--min-version <x.y>] <brief.md>\n');
|
||||
process.exit(2);
|
||||
|
|
|
|||
|
|
@ -21,7 +21,7 @@
|
|||
// PROFILE_READ_ERROR — file unreadable or parse-error
|
||||
// PROFILE_NOT_FOUND — file does not exist
|
||||
//
|
||||
// Allowed model values: ['sonnet', 'opus', 'fable']. Haiku is allowed only when
|
||||
// Allowed model values: ['sonnet', 'opus']. Haiku is allowed only when
|
||||
// VOYAGE_ALLOW_HAIKU=1 (per global CLAUDE.md modellvalg-prinsipp: Haiku skal
|
||||
// ikke brukes som default; eksplisitt opt-in for spesielle bruksmønstre).
|
||||
|
||||
|
|
@ -42,7 +42,7 @@ export const PROFILE_REQUIRED_PHASES = Object.freeze([
|
|||
'brief', 'research', 'plan', 'execute', 'review', 'continue',
|
||||
]);
|
||||
|
||||
export const BASE_ALLOWED_MODELS = Object.freeze(['sonnet', 'opus', 'fable']);
|
||||
export const BASE_ALLOWED_MODELS = Object.freeze(['sonnet', 'opus']);
|
||||
|
||||
function getAllowedModels(env = process.env) {
|
||||
if (env.VOYAGE_ALLOW_HAIKU === '1') {
|
||||
|
|
|
|||
|
|
@ -1,119 +0,0 @@
|
|||
// lib/validators/query-privacy-gate.mjs
|
||||
// Inspect an outbound research query before it leaves the machine. Called
|
||||
// only from the new high-effort steps (Phase 4.5 dimension discovery + the
|
||||
// bounded Phase 5 loop turns) — the existing single-pass Phase 5 path is
|
||||
// unchanged (Step 6, plan-v2).
|
||||
//
|
||||
// Two-tier, same shape as lib/exporters/endpoint-validator.mjs's SSRF gate:
|
||||
// - WARN tier — absolute filesystem paths, repo-internal identifiers.
|
||||
// Operator-overridable via `strict: false` / `--soft` (matches
|
||||
// lib/validators/research-validator.mjs's strict/soft convention), and
|
||||
// fully bypassable via the VOYAGE_QUERY_PRIVACY_ALLOW=1 opt-in.
|
||||
// - HARD-BLOCK tier — secret-shaped tokens. NEVER overridable by strict,
|
||||
// --soft, or the opt-in env var — mirrors endpoint-validator.mjs's
|
||||
// HARD_BLOCKED_HOSTS, where an opt-in widens the warn tier but never
|
||||
// unlocks the permanently-blocked one.
|
||||
//
|
||||
// CLI shim:
|
||||
// node lib/validators/query-privacy-gate.mjs [--soft] "<query text>"
|
||||
// → JSON {valid, errors, warnings}; exit 0 valid, 1 invalid.
|
||||
|
||||
import { issue } from '../util/result.mjs';
|
||||
|
||||
// WARN tier — absolute filesystem paths (leaks local directory layout).
|
||||
export const ABSOLUTE_PATH_PATTERNS = Object.freeze([
|
||||
/\/Users\/[^\s"'`]+/,
|
||||
/\/home\/[^\s"'`]+/,
|
||||
/[A-Za-z]:\\[^\s"'`]+/,
|
||||
/\$\{?HOME\}?\/[^\s"'`]+/,
|
||||
]);
|
||||
|
||||
// WARN tier — repo-internal identifiers that don't need to leave the
|
||||
// machine in a generic research query.
|
||||
export const REPO_IDENTIFIER_PATTERNS = Object.freeze([
|
||||
/git\.fromaitochitta\.com[^\s"'`]*/,
|
||||
/\bktg-plugin-marketplace\b/,
|
||||
/\bplugins\/cache\/[^\s"'`]+/,
|
||||
]);
|
||||
|
||||
// HARD-BLOCK tier — secret-shaped strings. Never operator-overridable.
|
||||
//
|
||||
// Token bodies that contain `-` or `_` break a plain `[A-Za-z0-9]{n,}` run, so
|
||||
// each such format needs its own pattern rather than relying on run length:
|
||||
// an Anthropic Console key runs out after `api03` (3 alphanumerics), and a
|
||||
// fine-grained GitHub PAT after its 22-character segment. Patterns whose body
|
||||
// class includes `-`/`_` carry no trailing `\b`, which would not fire on a
|
||||
// non-word final character.
|
||||
export const SECRET_SHAPED_PATTERNS = Object.freeze([
|
||||
/\bsk-[A-Za-z0-9]{20,}\b/, // OpenAI-style API key (sk-<48>)
|
||||
/\bsk-ant-[a-z0-9]+-[A-Za-z0-9_-]{20,}/, // Anthropic Console key (sk-ant-api03-/-oat01- + ~95 base64url)
|
||||
/\bAKIA[0-9A-Z]{16}\b/, // AWS access key ID
|
||||
/\bgh[pousr]_[A-Za-z0-9]{36,}\b/, // GitHub classic PAT / OAuth / user / server / refresh token
|
||||
/\bgithub_pat_[A-Za-z0-9_]{20,}/, // GitHub fine-grained PAT (github_pat_<22>_<59>)
|
||||
/\bxox[baprs]-[A-Za-z0-9-]{10,}\b/, // Slack token
|
||||
/-----BEGIN [A-Z ]*PRIVATE KEY-----/, // PEM private key block
|
||||
]);
|
||||
|
||||
function findMatch(patterns, text) {
|
||||
for (const re of patterns) {
|
||||
const m = re.exec(text);
|
||||
if (m) return m[0];
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* @param {string} text
|
||||
* @param {{strict?: boolean, env?: object}} [opts]
|
||||
* @returns {{valid: boolean, errors: import('../util/result.mjs').Issue[], warnings: import('../util/result.mjs').Issue[]}}
|
||||
*/
|
||||
export function validateOutboundQuery(text, opts = {}) {
|
||||
const strict = opts.strict !== false;
|
||||
const env = opts.env || process.env;
|
||||
// Bypasses the WARN tier entirely — never affects the hard-block tier below.
|
||||
const allowWarnTier = env.VOYAGE_QUERY_PRIVACY_ALLOW === '1';
|
||||
|
||||
if (typeof text !== 'string' || text.length === 0) {
|
||||
return { valid: false, errors: [issue('PRIVACY_EMPTY_QUERY', 'Outbound query must be a non-empty string')], warnings: [] };
|
||||
}
|
||||
|
||||
const errors = [];
|
||||
const warnings = [];
|
||||
|
||||
// Hard-block tier — checked unconditionally; no opt-in reaches this branch.
|
||||
const secretMatch = findMatch(SECRET_SHAPED_PATTERNS, text);
|
||||
if (secretMatch) {
|
||||
errors.push(issue('PRIVACY_SECRET_SHAPED', `Outbound query contains a secret-shaped token: ${secretMatch}`));
|
||||
}
|
||||
|
||||
if (!allowWarnTier) {
|
||||
const pathMatch = findMatch(ABSOLUTE_PATH_PATTERNS, text);
|
||||
if (pathMatch) {
|
||||
const issueObj = issue('PRIVACY_ABSOLUTE_PATH', `Outbound query contains an absolute filesystem path: ${pathMatch}`);
|
||||
if (strict) errors.push(issueObj); else warnings.push(issueObj);
|
||||
}
|
||||
|
||||
const repoMatch = findMatch(REPO_IDENTIFIER_PATTERNS, text);
|
||||
if (repoMatch) {
|
||||
const issueObj = issue('PRIVACY_REPO_IDENTIFIER', `Outbound query contains a repo-internal identifier: ${repoMatch}`);
|
||||
if (strict) errors.push(issueObj); else warnings.push(issueObj);
|
||||
}
|
||||
}
|
||||
|
||||
return { valid: errors.length === 0, errors, warnings };
|
||||
}
|
||||
|
||||
// ---- CLI shim ----------------------------------------------------------------
|
||||
|
||||
if (import.meta.url === `file://${process.argv[1]}`) {
|
||||
const args = process.argv.slice(2);
|
||||
const strict = !args.includes('--soft');
|
||||
const text = args.find(a => !a.startsWith('--'));
|
||||
if (text === undefined) {
|
||||
process.stderr.write('Usage: query-privacy-gate.mjs [--soft] "<query text>"\n');
|
||||
process.exit(2);
|
||||
}
|
||||
const r = validateOutboundQuery(text, { strict });
|
||||
process.stdout.write(JSON.stringify(r) + '\n');
|
||||
process.exit(r.valid ? 0 : 1);
|
||||
}
|
||||
4
package-lock.json
generated
4
package-lock.json
generated
|
|
@ -1,12 +1,12 @@
|
|||
{
|
||||
"name": "voyage",
|
||||
"version": "5.9.1",
|
||||
"version": "5.7.0",
|
||||
"lockfileVersion": 3,
|
||||
"requires": true,
|
||||
"packages": {
|
||||
"": {
|
||||
"name": "voyage",
|
||||
"version": "5.9.1",
|
||||
"version": "5.7.0",
|
||||
"license": "MIT",
|
||||
"engines": {
|
||||
"node": ">=18"
|
||||
|
|
|
|||
|
|
@ -1,6 +1,6 @@
|
|||
{
|
||||
"name": "voyage",
|
||||
"version": "5.9.1",
|
||||
"version": "5.7.0",
|
||||
"description": "Voyage — brief, research, plan, execute, review, continue. Contract-driven Claude Code pipeline. /trekbrief, /trekplan, and /trekreview each end by building a self-contained operator-annotation HTML (scripts/annotate.mjs, modelled on claude-code-100x): select text or click any heading/paragraph/list-item, pick intent (Fiks/Endre/Spørsmål), write comment, copy structured prompt, paste back, Claude revises the .md.",
|
||||
"type": "module",
|
||||
"engines": {
|
||||
|
|
|
|||
|
|
@ -1,342 +0,0 @@
|
|||
#!/usr/bin/env node
|
||||
// scripts/storm-measure.mjs
|
||||
// Step 11 — the STORM adoption gate: deterministic Δ accounting over
|
||||
// ${CLAUDE_PLUGIN_DATA}/trekresearch-stats.jsonl.
|
||||
//
|
||||
// What this decides: whether the bounded Phase 5 conversation loop buys enough
|
||||
// extra source/coverage breadth to justify flipping VOYAGE_STORM_ENABLED on by
|
||||
// default in lib/util/research-loop-cap.mjs. Nothing else. The thresholds live
|
||||
// in docs/storm-measurement.md and were committed BEFORE any measurement run —
|
||||
// that pre-registration is the whole point, so this script never invents them.
|
||||
//
|
||||
// What this does NOT measure: outline quality, answer correctness, or operator
|
||||
// satisfaction. It measures breadth (distinct sources, dimensions covered).
|
||||
// A breadth win is necessary for adoption, not sufficient on its own.
|
||||
//
|
||||
// Honesty properties, both load-bearing:
|
||||
// - Runs with empty_turns > 0 are EXCLUDED from the gain and REPORTED. An
|
||||
// empty turn means the loop spent budget and returned nothing; leaving those
|
||||
// in decides adoption on a broken denominator.
|
||||
// - A stats file carrying no `effort` field is a loud error. The silent
|
||||
// failure this prevents is an empty treatment group reading as "no gain",
|
||||
// which would decline the mechanism for a schema reason.
|
||||
//
|
||||
// Zero deps. Node stdlib only.
|
||||
|
||||
import { readFileSync, existsSync } from 'node:fs';
|
||||
import { join } from 'node:path';
|
||||
|
||||
// Pre-registered in docs/storm-measurement.md. Do not tune to fit a result.
|
||||
export const ADOPT_THRESHOLD = 0.30;
|
||||
export const DECLINE_THRESHOLD = 0.15;
|
||||
|
||||
const STATS_FILENAME = 'trekresearch-stats.jsonl';
|
||||
|
||||
// ---- pure core (unit-tested) -------------------------------------------------
|
||||
|
||||
/** @param {number[]} xs @returns {number|null} null for an empty list — 0 would read as a measurement. */
|
||||
export function median(xs) {
|
||||
if (!Array.isArray(xs) || xs.length === 0) return null;
|
||||
const s = [...xs].sort((a, b) => a - b);
|
||||
const mid = s.length >> 1;
|
||||
return s.length % 2 ? s[mid] : (s[mid - 1] + s[mid]) / 2;
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse a trekresearch-stats.jsonl body into effort-carrying records.
|
||||
*
|
||||
* Rows predating Step 9 have no `effort` field; they are dropped and counted as
|
||||
* `legacy` rather than silently pooled into the control arm. A file where NO row
|
||||
* carries `effort` throws — see the header note on silent failure.
|
||||
*
|
||||
* @param {string} text
|
||||
* @returns {{records: object[], malformed: number, legacy: number}}
|
||||
*/
|
||||
export function parseStats(text) {
|
||||
const lines = String(text ?? '').split('\n');
|
||||
const records = [];
|
||||
let malformed = 0;
|
||||
let legacy = 0;
|
||||
let parsed = 0;
|
||||
|
||||
for (const line of lines) {
|
||||
if (!line.trim()) continue;
|
||||
let rec;
|
||||
try {
|
||||
rec = JSON.parse(line);
|
||||
} catch {
|
||||
malformed++;
|
||||
continue;
|
||||
}
|
||||
parsed++;
|
||||
if (typeof rec?.effort !== 'string' || rec.effort.length === 0) {
|
||||
legacy++;
|
||||
continue;
|
||||
}
|
||||
records.push(rec);
|
||||
}
|
||||
|
||||
if (parsed === 0) {
|
||||
throw new Error(`storm-measure: no records in stats file (${malformed} malformed line(s)).`);
|
||||
}
|
||||
if (records.length === 0) {
|
||||
throw new Error(
|
||||
`storm-measure: no records carry an \`effort\` field (${legacy} legacy row(s)). ` +
|
||||
`The gate groups on \`effort\`; without it there is no treatment arm to measure. ` +
|
||||
`Re-run the measurement set on a build that emits the Step 9 fields.`,
|
||||
);
|
||||
}
|
||||
return { records, malformed, legacy };
|
||||
}
|
||||
|
||||
/**
|
||||
* Split off the runs that must not count toward a gain.
|
||||
*
|
||||
* A value that cannot be READ as a turn count is excluded, not treated as zero.
|
||||
* `Number('many')` is NaN, and testing `Number.isFinite(empty) && empty > 0`
|
||||
* sent NaN down the eligible branch — so a garbage field silently re-entered the
|
||||
* denominator, in the direction that flatters adoption: the run whose bookkeeping
|
||||
* broke is the run whose numbers deserve the least trust. Absent and null stay
|
||||
* eligible via `?? 0`, because a field that was never written is a genuine zero
|
||||
* on any run where the loop did not arm.
|
||||
*
|
||||
* @param {object[]} records
|
||||
* @returns {{eligible: object[], excluded: number}}
|
||||
*/
|
||||
export function partitionEligible(records) {
|
||||
const eligible = [];
|
||||
let excluded = 0;
|
||||
for (const r of records) {
|
||||
const empty = Number(r.empty_turns ?? 0);
|
||||
if (!Number.isFinite(empty) || empty > 0) excluded++;
|
||||
else eligible.push(r);
|
||||
}
|
||||
return { eligible, excluded };
|
||||
}
|
||||
|
||||
function numbers(records, pick) {
|
||||
return records.map(pick).filter((n) => Number.isFinite(n));
|
||||
}
|
||||
|
||||
/** Relative gain (treatment − control)/control. null when control is absent or zero. */
|
||||
function relGain(control, treatment) {
|
||||
if (control === null || treatment === null || control === 0) return null;
|
||||
return (treatment - control) / control;
|
||||
}
|
||||
|
||||
/**
|
||||
* Full measurement over a parsed record set.
|
||||
*
|
||||
* Treatment arm = `effort: high` (the only effort at which the loop runs).
|
||||
* Control arm = every other effort.
|
||||
*
|
||||
* - sources: between-arm median gain in `unique_sources`.
|
||||
* - dimensions: within-run median gain of (dimensions − dimensions_baseline)
|
||||
* / dimensions_baseline across the treatment arm. It is a
|
||||
* within-run delta by construction, so it needs no control arm —
|
||||
* the control arm's is 0, the loop being inert there.
|
||||
*
|
||||
* @param {object[]} records
|
||||
*/
|
||||
export function measure(records) {
|
||||
const { eligible, excluded } = partitionEligible(records);
|
||||
const treatment = eligible.filter((r) => r.effort === 'high');
|
||||
const control = eligible.filter((r) => r.effort !== 'high');
|
||||
|
||||
const srcControl = median(numbers(control, (r) => Number(r.unique_sources)));
|
||||
const srcTreatment = median(numbers(treatment, (r) => Number(r.unique_sources)));
|
||||
const sourcesGain = relGain(srcControl, srcTreatment);
|
||||
|
||||
const dimDeltas = treatment
|
||||
.map((r) => ({ d: Number(r.dimensions), b: Number(r.dimensions_baseline) }))
|
||||
.filter(({ d, b }) => Number.isFinite(d) && Number.isFinite(b) && b > 0)
|
||||
.map(({ d, b }) => (d - b) / b);
|
||||
const dimensionsGain = median(dimDeltas);
|
||||
|
||||
return {
|
||||
control: { n: control.length, sources: srcControl },
|
||||
treatment: { n: treatment.length, sources: srcTreatment },
|
||||
sources: { control: srcControl, treatment: srcTreatment, gain: sourcesGain },
|
||||
dimensions: { gain: dimensionsGain, n: dimDeltas.length },
|
||||
excluded,
|
||||
verdict: decideVerdict(sourcesGain, dimensionsGain),
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Pre-registered mapping, verbatim from the brief: "median forbedring >= 30 %
|
||||
* på (a) eller (b) → adopt. < 15 % → decline." OR on both sides, adopt
|
||||
* evaluated first — so a strong win on one axis is an adopt even when the
|
||||
* other axis sits under the decline bar. A stricter AND rule may well be the
|
||||
* better decision procedure, but changing it here is changing the
|
||||
* pre-registration after the fact, which is the one thing the constraint
|
||||
* exists to prevent.
|
||||
*
|
||||
* @param {number|null} sourcesGain
|
||||
* @param {number|null} dimensionsGain
|
||||
* @returns {'adopt'|'decline'|'inconclusive'|'insufficient-data'}
|
||||
*/
|
||||
export function decideVerdict(sourcesGain, dimensionsGain) {
|
||||
if (sourcesGain === null || sourcesGain === undefined) return 'insufficient-data';
|
||||
if (dimensionsGain === null || dimensionsGain === undefined) return 'insufficient-data';
|
||||
if (sourcesGain >= ADOPT_THRESHOLD || dimensionsGain >= ADOPT_THRESHOLD) return 'adopt';
|
||||
if (sourcesGain < DECLINE_THRESHOLD || dimensionsGain < DECLINE_THRESHOLD) return 'decline';
|
||||
return 'inconclusive';
|
||||
}
|
||||
|
||||
/**
|
||||
* SC activation check, BOTH halves.
|
||||
*
|
||||
* The SC asks two things of an `effort: high` run: that it discovered at least
|
||||
* one dimension, AND that the dimension list in the output brief is a TRUE
|
||||
* SUPERSET of the interview-derived ones. This function used to check only
|
||||
* `dimensions - dimensions_baseline >= 1`, which is a count delta and says
|
||||
* nothing about membership — a run that dropped two interview dimensions and
|
||||
* added three discovered ones passed while violating the second half.
|
||||
* Supersetness was asserted only by Phase 4.5's prose contract that discovery
|
||||
* APPENDS; nothing read it.
|
||||
*
|
||||
* The record cannot carry the dimension names: names are free prose, and
|
||||
* lib/exporters/field-allowlist.mjs denies prose by omission. So the run attests
|
||||
* membership with `dimensions_baseline_preserved`, a low-cardinality boolean set
|
||||
* in Phase 4.5, and this gate refuses to call activation OK without it. An
|
||||
* ABSENT attestation is not an attestation — legacy rows fail here rather than
|
||||
* passing on the old count-only rule.
|
||||
*
|
||||
* @param {object[]} records
|
||||
*/
|
||||
export function activationCheck(records) {
|
||||
const high = records.filter((r) => r.effort === 'high');
|
||||
if (high.length === 0) {
|
||||
return { ok: false, reason: 'no `effort: high` run found in stats', discovered_dimensions: null };
|
||||
}
|
||||
const last = high[high.length - 1];
|
||||
const d = Number(last.dimensions);
|
||||
const b = Number(last.dimensions_baseline);
|
||||
if (!Number.isFinite(d) || !Number.isFinite(b)) {
|
||||
return { ok: false, reason: 'latest high run lacks dimensions/dimensions_baseline', discovered_dimensions: null };
|
||||
}
|
||||
const discovered = d - b;
|
||||
const preserved = last.dimensions_baseline_preserved;
|
||||
|
||||
const base = {
|
||||
ts: last.ts ?? null,
|
||||
dimensions: d,
|
||||
dimensions_baseline: b,
|
||||
discovered_dimensions: discovered,
|
||||
dimensions_baseline_preserved: preserved ?? null,
|
||||
conv_turns: Number(last.conv_turns ?? 0),
|
||||
empty_turns: Number(last.empty_turns ?? 0),
|
||||
};
|
||||
|
||||
if (typeof preserved !== 'boolean') {
|
||||
return {
|
||||
...base,
|
||||
ok: false,
|
||||
reason:
|
||||
'latest high run does not attest `dimensions_baseline_preserved`; the SC needs a true ' +
|
||||
'superset of the interview dimensions, and a count delta cannot show membership',
|
||||
};
|
||||
}
|
||||
if (preserved === false) {
|
||||
return {
|
||||
...base,
|
||||
ok: false,
|
||||
reason:
|
||||
`latest high run discovered ${discovered} dimension(s) but did NOT preserve its interview ` +
|
||||
'baseline, so the final list is not a superset of it',
|
||||
};
|
||||
}
|
||||
if (discovered < 1) {
|
||||
return { ...base, ok: false, reason: 'latest high run discovered no dimensions beyond its baseline' };
|
||||
}
|
||||
return { ...base, ok: true };
|
||||
}
|
||||
|
||||
// ---- CLI shim ----------------------------------------------------------------
|
||||
|
||||
function pct(x) {
|
||||
return x === null ? 'n/a' : `${(x * 100).toFixed(1)}%`;
|
||||
}
|
||||
|
||||
function defaultStatsPath(env = process.env) {
|
||||
const dir = env.CLAUDE_PLUGIN_DATA;
|
||||
return dir ? join(dir, STATS_FILENAME) : null;
|
||||
}
|
||||
|
||||
function parseArgs(argv) {
|
||||
const o = { stats: null, json: false, activation: false, help: false };
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
const a = argv[i];
|
||||
if (a === '--stats') o.stats = argv[++i];
|
||||
else if (a === '--json') o.json = true;
|
||||
else if (a === '--activation-check') o.activation = true;
|
||||
else if (a === '--help' || a === '-h') o.help = true;
|
||||
else { process.stderr.write(`Unknown argument: ${a}\n`); process.exit(2); }
|
||||
}
|
||||
return o;
|
||||
}
|
||||
|
||||
function mainCli() {
|
||||
const o = parseArgs(process.argv.slice(2));
|
||||
if (o.help) {
|
||||
process.stdout.write(
|
||||
'Usage: storm-measure.mjs [--stats FILE] [--activation-check] [--json]\n' +
|
||||
' Default --stats: ${CLAUDE_PLUGIN_DATA}/' + STATS_FILENAME + '\n' +
|
||||
' Thresholds (pre-registered, docs/storm-measurement.md): ' +
|
||||
`adopt >= ${ADOPT_THRESHOLD * 100}%, decline < ${DECLINE_THRESHOLD * 100}%\n`,
|
||||
);
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
const statsPath = o.stats || defaultStatsPath();
|
||||
if (!statsPath) {
|
||||
process.stderr.write('storm-measure: CLAUDE_PLUGIN_DATA is not set and no --stats FILE was given.\n');
|
||||
process.exit(2);
|
||||
}
|
||||
if (!existsSync(statsPath)) {
|
||||
process.stderr.write(`storm-measure: stats file not found: ${statsPath}\n`);
|
||||
process.exit(2);
|
||||
}
|
||||
|
||||
let parsed;
|
||||
try {
|
||||
parsed = parseStats(readFileSync(statsPath, 'utf-8'));
|
||||
} catch (e) {
|
||||
process.stderr.write(`${e.message}\n`);
|
||||
process.exit(2);
|
||||
}
|
||||
|
||||
if (o.activation) {
|
||||
const res = activationCheck(parsed.records);
|
||||
process.stdout.write(JSON.stringify(res, null, 2) + '\n');
|
||||
process.exit(res.ok ? 0 : 1);
|
||||
}
|
||||
|
||||
const m = measure(parsed.records);
|
||||
|
||||
if (o.json) {
|
||||
process.stdout.write(JSON.stringify({ statsPath, ...m, malformed: parsed.malformed, legacy: parsed.legacy }, null, 2) + '\n');
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
const L = [];
|
||||
L.push(`STORM adoption gate — ${statsPath}`);
|
||||
L.push(` control (effort != high): n=${m.control.n} median unique_sources=${m.control.sources ?? 'n/a'}`);
|
||||
L.push(` treatment (effort = high): n=${m.treatment.n} median unique_sources=${m.treatment.sources ?? 'n/a'}`);
|
||||
L.push(` excluded (empty_turns > 0): ${m.excluded}`);
|
||||
if (parsed.legacy) L.push(` legacy rows without \`effort\`: ${parsed.legacy}`);
|
||||
if (parsed.malformed) L.push(` malformed lines: ${parsed.malformed}`);
|
||||
L.push('');
|
||||
L.push(` median gain, unique_sources: ${pct(m.sources.gain)}`);
|
||||
L.push(` median gain, dimensions over baseline: ${pct(m.dimensions.gain)} (n=${m.dimensions.n})`);
|
||||
L.push('');
|
||||
L.push(` thresholds: adopt >= ${pct(ADOPT_THRESHOLD)} on EITHER · decline < ${pct(DECLINE_THRESHOLD)} on EITHER · adopt wins ties`);
|
||||
L.push(` VERDICT: ${m.verdict}`);
|
||||
process.stdout.write(L.join('\n') + '\n');
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
if (import.meta.url === `file://${process.argv[1]}`) {
|
||||
mainCli();
|
||||
}
|
||||
|
|
@ -141,7 +141,7 @@ introduced. This section bridges sessions — it's the "baton" in a relay race.}
|
|||
- **Master plan:** `{plan file path}`
|
||||
- **Steps from plan:** {step N}–{step M}
|
||||
- **Estimated complexity:** {low | medium | high}
|
||||
- **Model recommendation:** {opus | sonnet | fable} — {rationale}
|
||||
- **Model recommendation:** {opus | sonnet} — {rationale}
|
||||
|
||||
## Recovery Metadata
|
||||
|
||||
|
|
|
|||
|
|
@ -17,9 +17,8 @@ source: {interview | manual}
|
|||
# plan polishing a wrong premise after a rejected iteration).
|
||||
framing: {preserve | refine | replace | new-direction}
|
||||
# v5.1 — per-phase effort + model signal (Phase 3.5).
|
||||
# `effort` ∈ {low, standard, high}; `model` ∈ {sonnet, opus, fable} (v5.9).
|
||||
# Omit `model:` for `standard` so composition falls through to profile
|
||||
# resolver. Force-stop alternative is the commented
|
||||
# `effort` ∈ {low, standard, high}. Omit `model:` for `standard` so composition
|
||||
# falls through to profile resolver. Force-stop alternative is the commented
|
||||
# `phase_signals_partial: true` below (mutually exclusive with `phase_signals`).
|
||||
phase_signals:
|
||||
- phase: research
|
||||
|
|
|
|||
|
|
@ -16,7 +16,6 @@ import { dirname, join } from 'node:path';
|
|||
import { fileURLToPath } from 'node:url';
|
||||
import { resolvePhaseSignal } from '../../lib/profiles/phase-signal-resolver.mjs';
|
||||
import { validateBriefContent, PHASE_SIGNAL_PHASES, EFFORT_LEVELS } from '../../lib/validators/brief-validator.mjs';
|
||||
import { BASE_ALLOWED_MODELS } from '../../lib/validators/profile-validator.mjs';
|
||||
import { parseDocument } from '../../lib/util/frontmatter.mjs';
|
||||
|
||||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
|
|
@ -85,8 +84,8 @@ test('trekbrief — SC1: each of 4 phases has both effort AND model on full-sign
|
|||
assert.ok(EFFORT_LEVELS.includes(r.effort),
|
||||
`phase=${phase}: effort "${r.effort}" not in EFFORT_LEVELS`);
|
||||
if ('model' in r) {
|
||||
assert.ok(BASE_ALLOWED_MODELS.includes(r.model),
|
||||
`phase=${phase}: model "${r.model}" not in [${BASE_ALLOWED_MODELS.join(', ')}]`);
|
||||
assert.ok(['sonnet', 'opus'].includes(r.model),
|
||||
`phase=${phase}: model "${r.model}" not in [sonnet, opus]`);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
|
@ -100,19 +99,6 @@ test('trekbrief — SC1: missing phase_signals + brief_version 2.1 triggers BRIE
|
|||
);
|
||||
});
|
||||
|
||||
// --- v5.9 — fable tier option in the Phase 3.5 loop ---
|
||||
|
||||
test('trekbrief — v5.9 Phase 3.5 canonical mapping contains the fable row and offers 4 options', () => {
|
||||
const text = read();
|
||||
const startIdx = text.indexOf('## Phase 3.5');
|
||||
assert.ok(startIdx >= 0, 'Phase 3.5 not found');
|
||||
const section = text.slice(startIdx, text.indexOf('## Phase 4', startIdx));
|
||||
assert.ok(section.includes('fable → {effort: high, model: fable}'),
|
||||
'Phase 3.5 canonical mapping must contain the fable tier row');
|
||||
assert.ok(section.includes('with 4 options'),
|
||||
'Phase 3.5 loop must offer 4 options (AskUserQuestion maxItems: 4)');
|
||||
});
|
||||
|
||||
// --- v5.5 — framing enforcement + TL;DR + memory-alignment prose-pins ---
|
||||
|
||||
test('trekbrief — v5.5 Phase 2.5 framing declaration heading present', () => {
|
||||
|
|
|
|||
|
|
@ -1,127 +0,0 @@
|
|||
// tests/commands/trekendsession.test.mjs
|
||||
// Regression tests for /trekendsession (commands/trekendsession.md).
|
||||
//
|
||||
// Bug (2026-07-03): two of the three !`...` eager-exec blocks contained
|
||||
// unresolved placeholders (<project-dir> etc.). The harness executes
|
||||
// eager-exec blocks at command LOAD time, so zsh parsed <project-dir> as
|
||||
// input redirection and the command aborted before the model saw a single
|
||||
// instruction. Eager-exec is only valid for self-contained commands.
|
||||
//
|
||||
// Pattern D (markdown structure) — assertions against command prose.
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { readFileSync, readdirSync } from 'node:fs';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
|
||||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
const ROOT = join(HERE, '..', '..');
|
||||
const COMMANDS_DIR = join(ROOT, 'commands');
|
||||
const COMMAND_FILE = join(COMMANDS_DIR, 'trekendsession.md');
|
||||
|
||||
function readCommand() {
|
||||
return readFileSync(COMMAND_FILE, 'utf8');
|
||||
}
|
||||
|
||||
function extractPhase(commandText, phaseHeader) {
|
||||
const startIdx = commandText.indexOf(phaseHeader);
|
||||
if (startIdx === -1) return '';
|
||||
const rest = commandText.slice(startIdx);
|
||||
const nextPhase = rest.search(/\n## (?:Phase |Hard )/);
|
||||
if (nextPhase === -1) return rest;
|
||||
return rest.slice(0, nextPhase);
|
||||
}
|
||||
|
||||
// Extract all eager-exec blocks (!`...`) from a command/skill file,
|
||||
// including multi-line blocks. Returns [{ content, line }].
|
||||
function extractEagerBlocks(text) {
|
||||
const blocks = [];
|
||||
const re = /!`([^`]+)`/g;
|
||||
let m;
|
||||
while ((m = re.exec(text)) !== null) {
|
||||
const line = text.slice(0, m.index).split('\n').length;
|
||||
blocks.push({ content: m[1], line });
|
||||
}
|
||||
return blocks;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------
|
||||
// Marketplace-wide regression guard: eager-exec blocks must be
|
||||
// self-contained. An unresolved placeholder (<angle> or {curly}) in an
|
||||
// eager block is executed verbatim by the shell at load time — <x> is
|
||||
// parsed as input redirection and aborts the whole command load.
|
||||
// ---------------------------------------------------------------
|
||||
|
||||
test('eager-exec guard — no !`-block in commands/ contains an unresolved placeholder', () => {
|
||||
const offenders = [];
|
||||
for (const file of readdirSync(COMMANDS_DIR).filter((f) => f.endsWith('.md'))) {
|
||||
const text = readFileSync(join(COMMANDS_DIR, file), 'utf8');
|
||||
for (const { content, line } of extractEagerBlocks(text)) {
|
||||
// Placeholder conventions: <angle-word> or {curly_word}. Curly must
|
||||
// contain a separator (- or _) so JS destructuring like {join} in a
|
||||
// legitimate self-contained script does not false-positive; angle
|
||||
// placeholders are unambiguous (shell would parse them as redirects).
|
||||
if (/<[a-z][a-z0-9_-]*>/.test(content) || /\{[a-z][a-z0-9]*([_-][a-z0-9]+)+\}/.test(content)) {
|
||||
offenders.push(`${file}:${line}`);
|
||||
}
|
||||
}
|
||||
}
|
||||
assert.deepEqual(
|
||||
offenders,
|
||||
[],
|
||||
`eager-exec !\`-blocks run at command LOAD time and must be self-contained; ` +
|
||||
`placeholder found in: ${offenders.join(', ')}`,
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------
|
||||
// trekendsession-specific: exactly one eager block (Phase 1 project
|
||||
// discovery — self-contained, legitimate); Phases 3 and 4 are runtime
|
||||
// Bash-tool commands with model-substituted values, never eager.
|
||||
// ---------------------------------------------------------------
|
||||
|
||||
test('trekendsession — exactly one eager-exec block remains (Phase 1 discovery)', () => {
|
||||
const cmd = readCommand();
|
||||
const blocks = extractEagerBlocks(cmd);
|
||||
assert.equal(
|
||||
blocks.length,
|
||||
1,
|
||||
`expected exactly 1 eager-exec block (Phase 1 discovery), got ${blocks.length} at line(s) ${blocks.map((b) => b.line).join(', ')}`,
|
||||
);
|
||||
assert.match(
|
||||
blocks[0].content,
|
||||
/readdirSync\(root\)/,
|
||||
'the surviving eager block must be the self-contained Phase 1 discovery script',
|
||||
);
|
||||
});
|
||||
|
||||
test('trekendsession Phase 3 — atomic-write block is runtime Bash (no eager prefix) with plugin-root import', () => {
|
||||
const phase3 = extractPhase(readCommand(), '## Phase 3 ');
|
||||
assert.doesNotMatch(phase3, /!`/, 'Phase 3 must not use eager-exec — values exist only at runtime');
|
||||
assert.match(
|
||||
phase3,
|
||||
/\$\{CLAUDE_PLUGIN_ROOT\}\/lib\/util\/atomic-write\.mjs/,
|
||||
'Phase 3 import must use the absolute ${CLAUDE_PLUGIN_ROOT} path — cwd is the user repo, not the plugin root',
|
||||
);
|
||||
assert.doesNotMatch(
|
||||
phase3,
|
||||
/['"]\.\/lib\/util\/atomic-write\.mjs['"]/,
|
||||
'Phase 3 must not import atomic-write.mjs via a cwd-relative path',
|
||||
);
|
||||
});
|
||||
|
||||
test('trekendsession Phase 4 — validator call is runtime Bash (no eager prefix) with plugin-root path', () => {
|
||||
const phase4 = extractPhase(readCommand(), '## Phase 4 ');
|
||||
assert.doesNotMatch(phase4, /!`/, 'Phase 4 must not use eager-exec — the state-file path exists only at runtime');
|
||||
assert.match(
|
||||
phase4,
|
||||
/\$\{CLAUDE_PLUGIN_ROOT\}\/lib\/validators\/session-state-validator\.mjs/,
|
||||
'Phase 4 validator path must use the absolute ${CLAUDE_PLUGIN_ROOT} convention',
|
||||
);
|
||||
assert.doesNotMatch(
|
||||
phase4,
|
||||
/<[a-z][a-z0-9_-]*>/,
|
||||
'Phase 4 must not use <angle> placeholders in commands — zsh parses <x> as input redirection',
|
||||
);
|
||||
});
|
||||
|
|
@ -1,35 +0,0 @@
|
|||
// tests/commands/trekresearch-engine.test.mjs
|
||||
// Step 1 (deep-research-engine): pin the contract the `--engine deep-research`
|
||||
// adapter must hit. The adapted in-context `/deep-research` report, reduced into
|
||||
// the research-brief schema, must pass research-validator under the strict
|
||||
// default; and a brief missing a required section must fail. This is the one
|
||||
// genuinely automatable slice of SC2 (schema, not provenance).
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
import { validateResearchContent } from '../../lib/validators/research-validator.mjs';
|
||||
|
||||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
const ROOT = join(HERE, '..', '..');
|
||||
const FIXTURE = join(ROOT, 'tests', 'fixtures', 'research-deep-research-adapted.md');
|
||||
|
||||
test('deep-research adapter output contract — valid brief passes, missing section fails', () => {
|
||||
const text = readFileSync(FIXTURE, 'utf-8');
|
||||
|
||||
// (a) positive: the adapter's target output passes the validator (default = strict).
|
||||
const okResult = validateResearchContent(text);
|
||||
assert.equal(okResult.valid, true, JSON.stringify(okResult.errors));
|
||||
|
||||
// (b) negative: stripping a required section makes it fail with RESEARCH_MISSING_SECTION,
|
||||
// giving the contract teeth (a fixture that always passes proves nothing).
|
||||
const mutated = text.replace('## Dimensions', '## Removed');
|
||||
const badResult = validateResearchContent(mutated);
|
||||
assert.equal(badResult.valid, false);
|
||||
assert.ok(
|
||||
badResult.errors.find(e => e.code === 'RESEARCH_MISSING_SECTION'),
|
||||
'expected RESEARCH_MISSING_SECTION; got ' + JSON.stringify(badResult.errors),
|
||||
);
|
||||
});
|
||||
|
|
@ -34,277 +34,10 @@ test('trekresearch — sequencing-gate surface mentions BRIEF_V51_MISSING_SIGNAL
|
|||
|
||||
test('trekresearch — low-effort path references --quick equivalent', () => {
|
||||
const text = read();
|
||||
// Bound the Composition rule section by the next `###` heading rather than a
|
||||
// magic 2000-character window: a fixed count silently drops the match as soon
|
||||
// as prose is inserted above it, turning a real pin into a no-op.
|
||||
const sectionOf = (doc) => {
|
||||
const compIdx = doc.indexOf('## Composition rule (v5.1)');
|
||||
assert.ok(compIdx >= 0, 'Composition rule (v5.1) section missing');
|
||||
const nextHeading = doc.indexOf('\n### ', compIdx);
|
||||
return nextHeading > compIdx ? doc.slice(compIdx, nextHeading) : doc.slice(compIdx);
|
||||
};
|
||||
|
||||
// (a) positive: the low-effort path is documented inside the bounded section.
|
||||
assert.match(sectionOf(text), /--quick/, 'Low-effort path must mention --quick equivalent');
|
||||
|
||||
// (b) negative: an actual removal must still be caught — a bound that can
|
||||
// never fail proves nothing.
|
||||
const mutated = text.replace(/--quick/g, '--removed');
|
||||
assert.doesNotMatch(sectionOf(mutated), /--quick/,
|
||||
'heading-bounded slice must still fail on a genuine removal');
|
||||
});
|
||||
|
||||
// --- Step 7: Phase 5 bounded conversation loop (heading-bounded slices) ---
|
||||
|
||||
// Same bounding discipline as the Composition-rule pin above: slice from the
|
||||
// phase heading to the NEXT phase heading, never a fixed character window.
|
||||
function phaseSlice(doc, startHeading, endHeading) {
|
||||
const start = doc.indexOf(startHeading);
|
||||
assert.ok(start >= 0, `${startHeading} missing`);
|
||||
const end = doc.indexOf(endHeading, start);
|
||||
assert.ok(end > start, `${endHeading} missing — could not bound ${startHeading}`);
|
||||
return doc.slice(start, end);
|
||||
}
|
||||
|
||||
function phase5(doc) {
|
||||
return phaseSlice(doc, '## Phase 5 —', '## Phase 6 —');
|
||||
}
|
||||
|
||||
test('trekresearch — Phase 5 loop is gated on effort == high and names both primitives', () => {
|
||||
const p5 = phase5(read());
|
||||
assert.match(p5, /effort == 'high'/, 'Phase 5 loop must be gated on effort == \'high\'');
|
||||
assert.match(p5, /research-loop-cap\.mjs/, 'Phase 5 must call the loop-cap shim per turn');
|
||||
assert.match(p5, /query-privacy-gate\.mjs/, 'Phase 5 must route outbound queries through the privacy gate');
|
||||
assert.match(p5, /\$\{CLAUDE_PLUGIN_ROOT\}/, 'shim invocations must use the ${CLAUDE_PLUGIN_ROOT} path form');
|
||||
});
|
||||
|
||||
// CLAUDE_PLUGIN_DATA and CLAUDE_PLUGIN_ROOT are substituted in this command's
|
||||
// TEXT but are EMPTY in the Bash tool's process env. Every snippet below runs
|
||||
// in that env, so each needs a resolution that does not depend on it.
|
||||
test('trekresearch — the scope-marker snippets resolve a root instead of requiring CLAUDE_PLUGIN_DATA', () => {
|
||||
const p5 = phase5(read());
|
||||
const blocks = [...p5.matchAll(/```bash\n([\s\S]*?)```/g)].map((m) => m[1]);
|
||||
const write = blocks.find((b) => b.includes('trekresearch-loop-scope') && b.includes('printf'));
|
||||
const remove = blocks.find((b) => b.includes('trekresearch-loop-scope') && b.includes('rm -f'));
|
||||
assert.ok(write, 'Phase 5 must carry the scope-marker write snippet');
|
||||
assert.ok(remove, 'Phase 5 must carry the scope-marker removal snippet');
|
||||
|
||||
for (const [name, block] of [['write', write], ['remove', remove]]) {
|
||||
assert.match(
|
||||
block,
|
||||
/\$\{CLAUDE_PLUGIN_DATA:-\$HOME\/\.claude\/voyage\}/,
|
||||
`the ${name} snippet must fall back to the same root research-loop-cap.mjs resolves`,
|
||||
);
|
||||
assert.match(
|
||||
block,
|
||||
/case .* in\s*\n?\s*\/\*\)/,
|
||||
`the ${name} snippet must guard on ONE absolute-path test — write and remove must not disagree on what counts as usable`,
|
||||
);
|
||||
}
|
||||
|
||||
// Unset, ${CLAUDE_CODE_SESSION_ID} composes a marker named `.json`, which no
|
||||
// hook lookup and no TTL sweep ever matches or cleans up.
|
||||
assert.match(
|
||||
write,
|
||||
/-n "\$\{?CLAUDE_CODE_SESSION_ID/,
|
||||
'the write snippet must require a non-empty CLAUDE_CODE_SESSION_ID before composing the marker path',
|
||||
);
|
||||
});
|
||||
|
||||
test('trekresearch — the per-turn gates separate "gate could not run" from "gate says no"', () => {
|
||||
const p5 = phase5(read());
|
||||
assert.match(
|
||||
p5,
|
||||
/VOYAGE_ROOT/,
|
||||
'the gate snippet must resolve a plugin root rather than interpolating ${CLAUDE_PLUGIN_ROOT} straight into `node`',
|
||||
);
|
||||
assert.match(
|
||||
p5,
|
||||
/exit 2|could not run/i,
|
||||
'an unresolvable gate must be distinguishable from a denial — otherwise every query reads as a privacy violation no rewrite can clear',
|
||||
);
|
||||
assert.match(
|
||||
p5,
|
||||
/plugins\/cache/,
|
||||
'the fallback must name the plugin cache location it searches',
|
||||
);
|
||||
});
|
||||
|
||||
test('trekresearch — Phase 5 declares the loop bound and all three exits', () => {
|
||||
const p5 = phase5(read());
|
||||
assert.match(p5, /### Loop bound/, 'Phase 5 must carry a `### Loop bound` sub-heading');
|
||||
assert.match(
|
||||
p5,
|
||||
/\*\*Maximum 3 turns per under-illuminated dimension\.\*\*/,
|
||||
'the bound must be stated verbatim',
|
||||
);
|
||||
// Three exits — converged / cap exhausted / operator stop.
|
||||
assert.match(p5, /converged/i, 'exit 1 (converged) must be documented');
|
||||
assert.match(p5, /exhaust/i, 'exit 2 (cap exhausted) must be documented');
|
||||
assert.match(p5, /operator stop/i, 'exit 3 (operator stop) must be documented');
|
||||
// Exhaustion must reach the operator — a silent cap is indistinguishable
|
||||
// from convergence, which is the failure this loop exists to avoid.
|
||||
assert.match(
|
||||
p5,
|
||||
/visibl|visible|print/i,
|
||||
'cap exhaustion must be written visibly to the operator',
|
||||
);
|
||||
});
|
||||
|
||||
test('trekresearch — Phase 5 marks empty turns without re-targeting the same dimension', () => {
|
||||
const p5 = phase5(read());
|
||||
assert.match(p5, /`empty`/, 'a finding-less or citation-less turn must be marked `empty`');
|
||||
assert.match(p5, /empty_turns/, 'empty turns must be counted (empty_turns)');
|
||||
assert.match(
|
||||
p5,
|
||||
/does NOT re-target|not re-target/i,
|
||||
'an empty turn must not re-target the same dimension',
|
||||
);
|
||||
});
|
||||
|
||||
test('trekresearch — Phase 5 states the no-brief default and the moot precedence matrix', () => {
|
||||
const p5 = phase5(read());
|
||||
// (a) no-brief default
|
||||
assert.match(p5, /effort = 'standard'/, 'no-brief default effort must be stated');
|
||||
assert.match(
|
||||
p5,
|
||||
/--project/,
|
||||
'the no-brief default must be anchored to the absence of --project/brief.md',
|
||||
);
|
||||
// (b) precedence matrix — each entry independently makes the loop moot,
|
||||
// mirroring the --engine moot gate in Phase 4.
|
||||
for (const token of ['--quick', '--local', 'external_research_enabled']) {
|
||||
assert.ok(p5.includes(token), `moot matrix must name ${token}`);
|
||||
}
|
||||
assert.match(p5, /moot/i, 'the matrix must use the same moot vocabulary as the engine gate');
|
||||
// (c) interaction rule — effort: high without model under a cheap profile.
|
||||
assert.match(
|
||||
p5,
|
||||
/effort: high/,
|
||||
'the interaction rule for a brief carrying effort: high without model must be stated',
|
||||
);
|
||||
});
|
||||
|
||||
test('trekresearch — Phase 5 restates the honesty rule for loop output', () => {
|
||||
const p5 = phase5(read());
|
||||
// Whitespace-tolerant: the pin is on the sentence, not on where the
|
||||
// paragraph happens to wrap.
|
||||
assert.match(
|
||||
p5,
|
||||
/more\s+turns\s+do\s+not\s+make\s+a\s+finding\s+more\s+credible/i,
|
||||
'the honesty hard rule must be restated for the loop output',
|
||||
);
|
||||
});
|
||||
|
||||
test('trekresearch — Phase 5 pins survive only while the prose does (mutation control)', () => {
|
||||
const text = read();
|
||||
const mutated = text.replace(/research-loop-cap\.mjs/g, 'removed-cap.mjs');
|
||||
assert.doesNotMatch(
|
||||
phase5(mutated),
|
||||
/research-loop-cap\.mjs/,
|
||||
'heading-bounded Phase 5 slice must still fail on a genuine removal',
|
||||
);
|
||||
});
|
||||
|
||||
test('trekresearch — High-effort behavior keeps the standard/low effort sentences verbatim', () => {
|
||||
const text = read();
|
||||
assert.ok(
|
||||
text.includes('Standard effort (or absent): use the existing conditional triggers.'),
|
||||
'the standard-effort sentence must survive the Phase 5 rewrite verbatim',
|
||||
);
|
||||
assert.ok(
|
||||
text.includes('Low effort: inline research only, no agent swarm'),
|
||||
'the low-effort sentence must survive the Phase 5 rewrite verbatim',
|
||||
);
|
||||
});
|
||||
|
||||
// --- Step 8: Phase 4.5 dimension discovery + Independence amendment ---
|
||||
|
||||
const ORCHESTRATOR_FILE = join(ROOT, 'agents', 'research-orchestrator.md');
|
||||
function readOrchestrator() { return readFileSync(ORCHESTRATOR_FILE, 'utf8'); }
|
||||
|
||||
test('trekresearch — Phase 4.5 exists between Phase 4 and Phase 5 with the effort skip-guard', () => {
|
||||
const text = read();
|
||||
const p45 = text.indexOf('## Phase 4.5 —');
|
||||
assert.ok(p45 >= 0, 'Phase 4.5 heading missing');
|
||||
const p4 = text.indexOf('## Phase 4 —');
|
||||
const p5 = text.indexOf('## Phase 5 —');
|
||||
assert.ok(p4 >= 0 && p5 > p45 && p45 > p4, 'Phase 4.5 must sit between Phase 4 and Phase 5');
|
||||
|
||||
const slice = text.slice(p45, p5);
|
||||
assert.match(
|
||||
slice,
|
||||
/\*\*Skip this phase entirely unless `phase_signal_result\.effort == 'high'`/,
|
||||
'Phase 4.5 must carry the bolded skip-guard in the Phase 3.5 form',
|
||||
);
|
||||
assert.match(slice, /query-privacy-gate\.mjs/,
|
||||
'Phase 4.5 must name the privacy gate as its compensating control');
|
||||
assert.match(slice, /maxDimensions: 8|maxDimensions` *: *8/,
|
||||
'Phase 4.5 must augment under the existing maxDimensions ceiling, not raise it');
|
||||
});
|
||||
|
||||
// Phase 4.5 never invokes research-loop-cap.mjs, so the cap's own flag check
|
||||
// does not reach it. Gating on effort alone means unsetting VOYAGE_STORM_ENABLED
|
||||
// leaves half the mechanism live and the decline branch unreachable — while
|
||||
// CLAUDE.md claims both phases go inert. The guard has to name both conditions.
|
||||
test('trekresearch — Phase 4.5 skip-guard is gated on VOYAGE_STORM_ENABLED as well as effort', () => {
|
||||
const text = read();
|
||||
const p45 = text.indexOf('## Phase 4.5 —');
|
||||
const p5 = text.indexOf('## Phase 5 —');
|
||||
const slice = text.slice(p45, p5);
|
||||
|
||||
const guard = slice.slice(0, slice.indexOf('\n\n', slice.indexOf('**Skip this phase')));
|
||||
assert.match(guard, /VOYAGE_STORM_ENABLED/,
|
||||
'the Phase 4.5 skip-guard must name VOYAGE_STORM_ENABLED, not effort alone');
|
||||
assert.match(guard, /\bAND\b|\*\*and\*\*/,
|
||||
'the guard must be a conjunction — both conditions, not either');
|
||||
assert.match(slice, /decline/i,
|
||||
'Phase 4.5 must say why the flag gates it: the decline branch has to stay reachable');
|
||||
});
|
||||
|
||||
test('trekresearch — Independence hard rule carries an explicit Phase 4.5 amendment', () => {
|
||||
const text = read();
|
||||
const rulesIdx = text.indexOf('## Hard rules');
|
||||
assert.ok(rulesIdx >= 0, 'Hard rules section missing');
|
||||
const rules = text.slice(rulesIdx);
|
||||
const indIdx = rules.indexOf('**Independence:**');
|
||||
assert.ok(indIdx >= 0, 'Independence hard rule missing');
|
||||
// Bound the rule at the next bullet so the amendment must live inside it.
|
||||
const nextBullet = rules.indexOf('\n- **', indIdx);
|
||||
const independence = nextBullet > indIdx ? rules.slice(indIdx, nextBullet) : rules.slice(indIdx);
|
||||
assert.match(independence, /Amend(ed|ment)/i,
|
||||
'Independence must be explicitly amended, not silently contradicted');
|
||||
assert.match(independence, /Phase 4\.5/, 'the amendment must name Phase 4.5 as the crossing');
|
||||
assert.match(independence, /query-privacy-gate\.mjs/,
|
||||
'the amendment must name the compensating control');
|
||||
});
|
||||
|
||||
test('trekresearch — orchestrator phase map is correct, has no Phase 9, and carries Phase 4.5', () => {
|
||||
const doc = readOrchestrator();
|
||||
const start = doc.indexOf('<!-- Phase mapping');
|
||||
assert.ok(start >= 0, 'phase mapping comment missing');
|
||||
const end = doc.indexOf('-->', start);
|
||||
assert.ok(end > start, 'phase mapping comment not terminated');
|
||||
const map = doc.slice(start, end);
|
||||
|
||||
assert.doesNotMatch(map, /Command Phase 9/,
|
||||
'the command ends at Phase 8 — a Command Phase 9 row is a fiction');
|
||||
|
||||
// Six orchestrator rows, each pointing at the phase the command actually has.
|
||||
const expected = [
|
||||
[1, '4'],
|
||||
[2, '4'],
|
||||
[3, '5'],
|
||||
[4, '6'],
|
||||
[5, '7'],
|
||||
[6, '8'],
|
||||
];
|
||||
for (const [orch, cmd] of expected) {
|
||||
const re = new RegExp(`Orchestrator Phase ${orch}\\s+= Command Phase ${cmd.replace('.', '\\.')}\\b`);
|
||||
assert.match(map, re, `map row for Orchestrator Phase ${orch} must point at Command Phase ${cmd}`);
|
||||
}
|
||||
|
||||
assert.match(map, /Command Phase 4\.5/, 'the map must carry the new Phase 4.5 row');
|
||||
const compIdx = text.indexOf('## Composition rule (v5.1)');
|
||||
assert.ok(compIdx >= 0, 'Composition rule (v5.1) section missing');
|
||||
const section = text.slice(compIdx, compIdx + 2000);
|
||||
assert.match(section, /--quick/, 'Low-effort path must mention --quick equivalent');
|
||||
});
|
||||
|
||||
// --- v5.1.1 runtime SC4 + SC7 ---
|
||||
|
|
|
|||
|
|
@ -1,17 +0,0 @@
|
|||
[
|
||||
{
|
||||
"reviewer": "brief-conformance-reviewer",
|
||||
"findings": [
|
||||
{ "severity": "BLOCKER", "rule_key": "UNIMPLEMENTED_CRITERION", "file": "lib/handlers/login.mjs", "line": 17 },
|
||||
{ "severity": "MAJOR", "rule_key": "PLAN_EXECUTE_DRIFT", "file": "lib/handlers/login.mjs", "line": 13 }
|
||||
]
|
||||
},
|
||||
{
|
||||
"reviewer": "code-correctness-reviewer",
|
||||
"findings": [
|
||||
{ "severity": "BLOCKER", "rule_key": "SECURITY_INJECTION", "file": "lib/auth/jwt.mjs", "line": 19 },
|
||||
{ "severity": "MAJOR", "rule_key": "MISSING_TEST", "file": "lib/auth/refresh.mjs", "line": 0 },
|
||||
{ "severity": "MINOR", "rule_key": "MISSING_ERROR_HANDLING", "file": "lib/auth/refresh.mjs", "line": 10 }
|
||||
]
|
||||
}
|
||||
]
|
||||
45
tests/fixtures/brief-effort-fable.md
vendored
45
tests/fixtures/brief-effort-fable.md
vendored
|
|
@ -1,45 +0,0 @@
|
|||
---
|
||||
type: trekbrief
|
||||
brief_version: "2.1"
|
||||
created: 2026-07-02
|
||||
task: "Fixture: high-effort all phases on fable (v5.9 allowlist test)"
|
||||
slug: brief-effort-fable
|
||||
project_dir: .claude/projects/2026-07-02-brief-effort-fable/
|
||||
research_topics: 0
|
||||
research_status: complete
|
||||
auto_research: false
|
||||
interview_turns: 4
|
||||
source: fixture
|
||||
phase_signals:
|
||||
- phase: research
|
||||
effort: high
|
||||
model: fable
|
||||
- phase: plan
|
||||
effort: high
|
||||
model: fable
|
||||
- phase: execute
|
||||
effort: high
|
||||
model: fable
|
||||
- phase: review
|
||||
effort: high
|
||||
model: fable
|
||||
---
|
||||
|
||||
# Task: High-effort fable fixture
|
||||
|
||||
## Intent
|
||||
|
||||
Test fixture for the v5.9 fable model tier — all 4 phases at the
|
||||
high effort tier with explicit fable model overrides. Mirrors
|
||||
brief-effort-high.md with `model: opus` replaced by `model: fable`.
|
||||
|
||||
## Goal
|
||||
|
||||
Resolver returns `{effort: 'high', model: 'fable'}` for each of the 4
|
||||
PHASE_SIGNAL_PHASES.
|
||||
|
||||
## Success Criteria
|
||||
|
||||
- Validator passes with no BRIEF_INVALID_MODEL.
|
||||
- resolvePhaseSignal(fm, phase).effort === 'high' for all 4 phases.
|
||||
- resolvePhaseSignal(fm, phase).model === 'fable' for all 4 phases.
|
||||
22
tests/fixtures/expected.prom
vendored
22
tests/fixtures/expected.prom
vendored
|
|
@ -33,31 +33,19 @@ voyage_trekplan_deep_dives{_schema_id="trekplan",slug="add-auth",mode="default",
|
|||
voyage_trekplan_research_briefs_used{_schema_id="trekplan",slug="add-auth",mode="default",profile="premium",profile_source="flag"} 3
|
||||
# HELP voyage_trekresearch_agents_external voyage stats — trekresearch_agents_external
|
||||
# TYPE voyage_trekresearch_agents_external gauge
|
||||
voyage_trekresearch_agents_external{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",effort="high",profile="premium",profile_source="default"} 3
|
||||
voyage_trekresearch_agents_external{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",profile="premium",profile_source="default"} 3
|
||||
# HELP voyage_trekresearch_agents_local voyage stats — trekresearch_agents_local
|
||||
# TYPE voyage_trekresearch_agents_local gauge
|
||||
voyage_trekresearch_agents_local{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",effort="high",profile="premium",profile_source="default"} 5
|
||||
voyage_trekresearch_agents_local{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",profile="premium",profile_source="default"} 5
|
||||
# HELP voyage_trekresearch_contradictions voyage stats — trekresearch_contradictions
|
||||
# TYPE voyage_trekresearch_contradictions gauge
|
||||
voyage_trekresearch_contradictions{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",effort="high",profile="premium",profile_source="default"} 1
|
||||
# HELP voyage_trekresearch_conv_turns voyage stats — trekresearch_conv_turns
|
||||
# TYPE voyage_trekresearch_conv_turns gauge
|
||||
voyage_trekresearch_conv_turns{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",effort="high",profile="premium",profile_source="default"} 5
|
||||
voyage_trekresearch_contradictions{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",profile="premium",profile_source="default"} 1
|
||||
# HELP voyage_trekresearch_dimensions voyage stats — trekresearch_dimensions
|
||||
# TYPE voyage_trekresearch_dimensions gauge
|
||||
voyage_trekresearch_dimensions{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",effort="high",profile="premium",profile_source="default"} 4
|
||||
# HELP voyage_trekresearch_dimensions_baseline voyage stats — trekresearch_dimensions_baseline
|
||||
# TYPE voyage_trekresearch_dimensions_baseline gauge
|
||||
voyage_trekresearch_dimensions_baseline{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",effort="high",profile="premium",profile_source="default"} 3
|
||||
# HELP voyage_trekresearch_empty_turns voyage stats — trekresearch_empty_turns
|
||||
# TYPE voyage_trekresearch_empty_turns gauge
|
||||
voyage_trekresearch_empty_turns{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",effort="high",profile="premium",profile_source="default"} 1
|
||||
voyage_trekresearch_dimensions{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",profile="premium",profile_source="default"} 4
|
||||
# HELP voyage_trekresearch_open_questions voyage stats — trekresearch_open_questions
|
||||
# TYPE voyage_trekresearch_open_questions gauge
|
||||
voyage_trekresearch_open_questions{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",effort="high",profile="premium",profile_source="default"} 2
|
||||
# HELP voyage_trekresearch_unique_sources voyage stats — trekresearch_unique_sources
|
||||
# TYPE voyage_trekresearch_unique_sources gauge
|
||||
voyage_trekresearch_unique_sources{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",effort="high",profile="premium",profile_source="default"} 17
|
||||
voyage_trekresearch_open_questions{_schema_id="trekresearch",slug="add-auth",mode="default",scope="both",profile="premium",profile_source="default"} 2
|
||||
# HELP voyage_trekreview_duration_ms voyage stats — trekreview_duration_ms
|
||||
# TYPE voyage_trekreview_duration_ms histogram
|
||||
voyage_trekreview_duration_ms{_schema_id="trekreview",slug="add-auth",verdict="ALLOW",mode="default",profile="balanced",profile_source="flag"} 4521
|
||||
|
|
|
|||
2
tests/fixtures/jsonl-schemas.md
vendored
2
tests/fixtures/jsonl-schemas.md
vendored
|
|
@ -20,7 +20,7 @@
|
|||
| schema_id | fields | writer_path | line_ref | v4.1 additive | PII |
|
||||
|-----------|--------|-------------|----------|---------------|-----|
|
||||
| trekbrief-stats | ts, task, slug, mode, interview_turns, review_iterations, brief_quality, research_topics, auto_research, auto_result, project_dir | commands/trekbrief.md (orchestrator-emit Phase 7) | trekbrief.md:657-672 | profile, phase_models, profile_source | none |
|
||||
| trekresearch-stats | ts, question, mode, scope, engine, slug, project_dir, brief_path, dimensions, dimensions_baseline, dimensions_baseline_preserved, effort, conv_turns, empty_turns, unique_sources, agents_local, agents_external, gemini_used, confidence, contradictions, open_questions | commands/trekresearch.md (orchestrator-emit Stats tracking) | trekresearch.md:634-676 | profile, phase_models, parallel_agents, external_research_enabled, profile_source | none |
|
||||
| trekresearch-stats | ts, question, mode, scope, slug, project_dir, brief_path, dimensions, agents_local, agents_external, gemini_used, confidence, contradictions, open_questions | commands/trekresearch.md (orchestrator-emit Stats tracking) | trekresearch.md:388-410 | profile, phase_models, parallel_agents, external_research_enabled, profile_source | none |
|
||||
| trekplan-stats | ts, task, mode, slug, brief_path, project_dir, codebase_size, codebase_files, agents_deployed, deep_dives, research_briefs_used, research_scout_used, critic_verdict, guardian_verdict, outcome | commands/trekplan.md (orchestrator-emit Phase 12) | trekplan.md:805-826 | profile, phase_models, parallel_agents, profile_source | none |
|
||||
| trekexecute-stats (Phase 9 record) | ts, plan, plan_type, mode, result, steps_total, steps_passed, steps_failed, steps_skipped, failed_at_step | commands/trekexecute.md (orchestrator-emit Phase 9) | trekexecute.md:1479-1494 | profile, phase_models, profile_source | none |
|
||||
| trekexecute-stats (autonomy events) | ts, event, known_event, payload | lib/stats/event-emit.mjs `emit()` | event-emit.mjs:64-86 | payload.profile, payload.phase_models, payload.profile_source | none |
|
||||
|
|
|
|||
49
tests/fixtures/research-deep-research-adapted.md
vendored
49
tests/fixtures/research-deep-research-adapted.md
vendored
|
|
@ -1,49 +0,0 @@
|
|||
---
|
||||
type: trekresearch-brief
|
||||
created: 2026-06-30
|
||||
question: "Should /trekresearch delegate its external phase to the built-in /deep-research workflow?"
|
||||
confidence: 0.8
|
||||
dimensions: 2
|
||||
mcp_servers_used: []
|
||||
local_agents_used: []
|
||||
external_agents_used:
|
||||
- deep-research
|
||||
---
|
||||
|
||||
# Deep-research engine adapter output
|
||||
|
||||
> Fixture: a `/deep-research` in-context report reduced into the research-brief
|
||||
> schema by the `--engine deep-research` adapter (Step 4). Models the target the
|
||||
> adapter must hit; not real engine output.
|
||||
|
||||
## Executive Summary
|
||||
|
||||
Delegating the external phase to the built-in `/deep-research` workflow is a
|
||||
viable opt-in engine that supplies fan-out and cited claim-verification for free.
|
||||
Confidence is medium-high on the mechanism but lower on availability, because the
|
||||
workflow exposes no positive "is-enabled" probe. The load-bearing caveat is
|
||||
provenance: structural validity does not certify that the cited URLs are real, so
|
||||
a swarm fallback plus a human URL spot-check stay mandatory.
|
||||
|
||||
## Dimensions
|
||||
|
||||
### Engine mechanism -- Confidence: high
|
||||
|
||||
**External findings:**
|
||||
- `/deep-research` is a built-in dynamic workflow reachable only by prose instruction, with no programmatic API (https://code.claude.com/docs/workflows).
|
||||
- Its report lands in-context with no on-disk artifact, so the adapter must transform what is already in the turn (https://code.claude.com/docs/commands).
|
||||
|
||||
### Fallback ergonomics -- Confidence: high
|
||||
|
||||
**External findings:**
|
||||
- There is no positive availability probe; only `disableWorkflows` / `CLAUDE_CODE_DISABLE_WORKFLOWS` off-switches and a 2.1.154 version floor are documented (https://code.claude.com/docs/skills).
|
||||
- Disabled-workflow behavior under `claude -p` is undocumented, so the post-hoc presence check must be robust to every failure manifestation (https://github.com/anthropics/claude-code/issues/52272).
|
||||
|
||||
## Sources
|
||||
|
||||
| # | Source | Type | Quality | Used in |
|
||||
|---|--------|------|---------|---------|
|
||||
| 1 | https://code.claude.com/docs/workflows | official | high | Engine mechanism |
|
||||
| 2 | https://code.claude.com/docs/commands | official | high | Engine mechanism |
|
||||
| 3 | https://code.claude.com/docs/skills | official | high | Fallback ergonomics |
|
||||
| 4 | https://github.com/anthropics/claude-code/issues/52272 | community | medium | Fallback ergonomics |
|
||||
2
tests/fixtures/stats-sample.jsonl
vendored
2
tests/fixtures/stats-sample.jsonl
vendored
|
|
@ -2,4 +2,4 @@
|
|||
{"_schema_id":"trekexecute","ts":"2026-05-09T08:30:00.000Z","plan":"trekplan-add-auth.md","plan_type":"plan","mode":"execute","result":"completed","steps_total":12,"steps_passed":12,"steps_failed":0,"steps_skipped":0,"profile":"premium","profile_source":"inheritance"}
|
||||
{"_schema_id":"trekreview","ts":"2026-05-09T09:00:00.000Z","slug":"add-auth","verdict":"ALLOW","reviewed_files_count":18,"mode":"default","duration_ms":4521,"profile":"balanced","profile_source":"flag"}
|
||||
{"_schema_id":"trekbrief","ts":"2026-05-09T07:00:00.000Z","slug":"add-auth","mode":"default","interview_turns":7,"review_iterations":2,"research_topics":3,"profile":"economy","profile_source":"env"}
|
||||
{"_schema_id":"trekresearch","ts":"2026-05-09T07:30:00.000Z","slug":"add-auth","mode":"default","scope":"both","dimensions":4,"dimensions_baseline":3,"effort":"high","conv_turns":5,"empty_turns":1,"unique_sources":17,"agents_local":5,"agents_external":3,"contradictions":1,"open_questions":2,"profile":"premium","profile_source":"default"}
|
||||
{"_schema_id":"trekresearch","ts":"2026-05-09T07:30:00.000Z","slug":"add-auth","mode":"default","scope":"both","dimensions":4,"agents_local":5,"agents_external":3,"contradictions":1,"open_questions":2,"profile":"premium","profile_source":"default"}
|
||||
|
|
|
|||
|
|
@ -1,556 +0,0 @@
|
|||
// tests/hooks/agent-cap.test.mjs
|
||||
// Step 10 — pins hooks/scripts/pre-agent-cap.mjs, the PreToolUse enforcement
|
||||
// of the /trekresearch Phase 5 loop bound.
|
||||
//
|
||||
// The spike (docs/spike-pretooluse-subagent-reach.md, RESULT: FIRES) proved a
|
||||
// plugin PreToolUse hook observes sub-agent tool calls, so the cap can be
|
||||
// enforced rather than merely documented. This file pins the two properties
|
||||
// that matter in opposite directions:
|
||||
//
|
||||
// (a) it DENIES (exit 2) once the ledger shows the budget spent, and
|
||||
// (b) it does NOT over-block — an unrelated session, unparsable stdin, a
|
||||
// stale marker, or the kill switch all exit 0.
|
||||
//
|
||||
// Pattern: tests/hooks/bash-guard.test.mjs (child process via runHook).
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, existsSync } from 'node:fs';
|
||||
import { execFileSync } from 'node:child_process';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { runHookWithEnv } from '../helpers/hook-helper.mjs';
|
||||
|
||||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
const ROOT = join(HERE, '..', '..');
|
||||
const CAP_HOOK = join(ROOT, 'hooks', 'scripts', 'pre-agent-cap.mjs');
|
||||
const HOOKS_JSON = join(ROOT, 'hooks', 'hooks.json');
|
||||
|
||||
const SESSION = 'sess-abc123';
|
||||
const RUN_ID = 'run-xyz789';
|
||||
|
||||
/**
|
||||
* Build a throwaway CLAUDE_PLUGIN_DATA dir holding a scope marker for
|
||||
* `sessionId` and `turns` ledger entries for RUN_ID.
|
||||
*/
|
||||
function fixture({ turns = 0, sessionId = SESSION, startedAt = new Date(), exhausted = false } = {}) {
|
||||
const dir = mkdtempSync(join(tmpdir(), 'voyage-cap-'));
|
||||
mkdirSync(join(dir, 'trekresearch-loop-scope'), { recursive: true });
|
||||
writeFileSync(
|
||||
join(dir, 'trekresearch-loop-scope', `${sessionId}.json`),
|
||||
JSON.stringify({ runId: RUN_ID, startedAt: startedAt.toISOString() }),
|
||||
);
|
||||
const lines = Array.from({ length: turns }, (_, i) =>
|
||||
JSON.stringify({ ts: new Date().toISOString(), runId: RUN_ID, dimension: `d${i}`, effort: 'high', slot: i + 1 }),
|
||||
);
|
||||
// The tombstone research-loop-cap.mjs appends when it denies a turn for
|
||||
// budget. Its presence is what tells this hook "the gate already said no".
|
||||
if (exhausted) lines.push(JSON.stringify({ ts: new Date().toISOString(), runId: RUN_ID, exhausted: true }));
|
||||
writeFileSync(join(dir, 'trekresearch-loop-ledger.jsonl'), lines.length ? lines.join('\n') + '\n' : '');
|
||||
return dir;
|
||||
}
|
||||
|
||||
// TREKRESEARCH_MAX_CONV_TURNS=1 => budget = 1 * MAX_TOTAL_DIMENSIONS (8).
|
||||
const CAPPED_ENV = { VOYAGE_STORM_ENABLED: '1', TREKRESEARCH_MAX_CONV_TURNS: '1' };
|
||||
const BUDGET = 8;
|
||||
|
||||
function searchInput(sessionId = SESSION) {
|
||||
return {
|
||||
session_id: sessionId,
|
||||
hook_event_name: 'PreToolUse',
|
||||
tool_name: 'WebSearch',
|
||||
tool_input: { query: 'claude code hooks reference' },
|
||||
agent_id: 'aa6d19525a4680fe0',
|
||||
agent_type: 'general-purpose',
|
||||
};
|
||||
}
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// DENY — budget spent
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-agent-cap DENIES once the budget gate has denied a turn (tombstone present)', async () => {
|
||||
const dir = fixture({ turns: BUDGET, exhausted: true });
|
||||
const { code, stderr } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 2);
|
||||
assert.match(stderr, /loop cap/i, 'stderr must name the cap it enforced');
|
||||
assert.match(stderr, new RegExp(`${BUDGET}`), 'stderr must state the budget');
|
||||
});
|
||||
|
||||
test('pre-agent-cap DENIES above the budget too (a breached ledger, whatever caused it)', async () => {
|
||||
const dir = fixture({ turns: BUDGET + 5 });
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 2);
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// ALLOW — under the cap, and ON the cap.
|
||||
//
|
||||
// allowTurn appends BEFORE the turn runs, so during the FINAL granted turn the
|
||||
// ledger already holds `budget` records. Denying at `used >= budget` therefore
|
||||
// blocked that turn's own tool calls: the primitive granted B turns, the
|
||||
// harness permitted B-1, and an exhausted run always ended through an exit-2
|
||||
// denial rather than the graceful "cap exhausted" exit the prose defines. The
|
||||
// boundary belongs one turn later, and the tombstone above — not the count —
|
||||
// is what marks a run actually finished.
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-agent-cap ALLOWS a loop turn under the budget', async () => {
|
||||
const dir = fixture({ turns: BUDGET - 1 });
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
test('pre-agent-cap ALLOWS the FINAL granted turn — its own record is already on the ledger', async () => {
|
||||
const dir = fixture({ turns: BUDGET });
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(
|
||||
code, 0,
|
||||
'turn B is granted and in flight; denying it makes the harness permit B-1 turns and forces the wrong exit',
|
||||
);
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// DOES NOT OVER-BLOCK — the property that keeps this hook safe to wire
|
||||
// globally. A broken PreToolUse hook would brick every session on the box.
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-agent-cap ALLOWS an unrelated session even when a loop is exhausted', async () => {
|
||||
const dir = fixture({ turns: BUDGET });
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput('some-other-session'), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0, 'no scope marker for this session_id => out of scope');
|
||||
});
|
||||
|
||||
test('pre-agent-cap ALLOWS when no scope marker directory exists at all', async () => {
|
||||
const dir = mkdtempSync(join(tmpdir(), 'voyage-cap-empty-'));
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
test('pre-agent-cap ALLOWS on unparsable stdin', async () => {
|
||||
const dir = fixture({ turns: BUDGET });
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, 'not json at all', {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
test('pre-agent-cap ALLOWS when the input carries no session_id', async () => {
|
||||
const dir = fixture({ turns: BUDGET });
|
||||
const input = searchInput();
|
||||
delete input.session_id;
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, input, {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// TTL / auto-reset — a marker left behind by a crashed run must not deny
|
||||
// tool calls forever.
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-agent-cap ALLOWS when the scope marker is older than the TTL', async () => {
|
||||
const dir = fixture({ turns: BUDGET, startedAt: new Date(Date.now() - 48 * 3600 * 1000) });
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
VOYAGE_CAP_SCOPE_TTL_MS: '1000',
|
||||
});
|
||||
assert.strictEqual(code, 0, 'a stale marker must auto-reset, not deny forever');
|
||||
});
|
||||
|
||||
// The TTL runs from marker.startedAt, not from last activity, and `claude
|
||||
// --resume` keeps the same session_id — so a run that died leaving its marker
|
||||
// behind hands the resumed session whatever deny window is left. Two things
|
||||
// bound that: the window is hours, not the machine's life (below), and it only
|
||||
// opens at all once the budget gate has actually denied a turn.
|
||||
test('pre-agent-cap ALLOWS a resumed session whose crashed run never exhausted its budget', async () => {
|
||||
// Marker still fresh, ledger part-spent, no tombstone: the run died mid-loop.
|
||||
const dir = fixture({ turns: 5 });
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(
|
||||
code, 0,
|
||||
'a part-spent run leaves no denial record, so resuming its session must not brick unrelated work',
|
||||
);
|
||||
});
|
||||
|
||||
test('pre-agent-cap uses a default TTL of hours, not a day — a 3h-old marker auto-resets', async () => {
|
||||
const dir = fixture({
|
||||
turns: BUDGET,
|
||||
exhausted: true,
|
||||
startedAt: new Date(Date.now() - 3 * 3600 * 1000),
|
||||
});
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
}); // no VOYAGE_CAP_SCOPE_TTL_MS — this is the built-in default
|
||||
assert.strictEqual(code, 0, 'no real research run lasts 3h; a marker that old is debris');
|
||||
});
|
||||
|
||||
test('pre-agent-cap names the marker path when it denies, so the operator has a remedy', async () => {
|
||||
const dir = fixture({ turns: BUDGET, exhausted: true });
|
||||
const { code, stderr } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 2);
|
||||
assert.ok(
|
||||
stderr.includes(join(dir, 'trekresearch-loop-scope', `${SESSION}.json`)),
|
||||
`stderr must name the marker to delete; got:\n${stderr}`,
|
||||
);
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// Fail CLOSED once in scope — the hook's own header says a budget control
|
||||
// that cannot count must not grant. The unreadable-ledger branch returned 0
|
||||
// and therefore ALLOWED, which is the opposite. A directory standing where
|
||||
// the ledger file belongs reproduces it portably (EISDIR).
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-agent-cap DENIES when the ledger cannot be read at all (fail closed)', async () => {
|
||||
const dir = fixture({ turns: 0 });
|
||||
const ledgerPath = join(dir, 'trekresearch-loop-ledger.jsonl');
|
||||
rmSync(ledgerPath, { force: true });
|
||||
mkdirSync(ledgerPath, { recursive: true });
|
||||
const { code, stderr } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 2, 'an in-scope run whose ledger cannot be counted must not be granted');
|
||||
assert.match(stderr, /could not be read|unreadable/i, 'stderr must say counting failed, not that the budget is spent');
|
||||
});
|
||||
|
||||
test('pre-agent-cap ALLOWS an unreadable ledger when the session is OUT of scope', async () => {
|
||||
const dir = fixture({ turns: 0, sessionId: 'a-different-session' });
|
||||
const ledgerPath = join(dir, 'trekresearch-loop-ledger.jsonl');
|
||||
rmSync(ledgerPath, { force: true });
|
||||
mkdirSync(ledgerPath, { recursive: true });
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0, 'fail-closed is scoped to the loop, it must not brick unrelated sessions');
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// Kill switch + default-off
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-agent-cap kill switch VOYAGE_DISABLE_CAP_HOOK=1 allows an exhausted loop', async () => {
|
||||
const dir = fixture({ turns: BUDGET });
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
VOYAGE_DISABLE_CAP_HOOK: '1',
|
||||
});
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
test('pre-agent-cap is inert when VOYAGE_STORM_ENABLED is not 1 (default-off)', async () => {
|
||||
const dir = fixture({ turns: BUDGET });
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
VOYAGE_STORM_ENABLED: '0',
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// The measured environment — CLAUDE_PLUGIN_DATA is EMPTY in the Bash tool
|
||||
// env, so that is the environment every real run happens in. The writer (the
|
||||
// Phase 5 bash snippet) and the reader (this hook) must land on the SAME
|
||||
// fallback root, or the hook allows unconditionally while claiming to enforce.
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-agent-cap enforces via the fallback root when CLAUDE_PLUGIN_DATA is absent', async () => {
|
||||
const home = mkdtempSync(join(tmpdir(), 'voyage-cap-home-'));
|
||||
const root = join(home, '.claude', 'voyage');
|
||||
mkdirSync(join(root, 'trekresearch-loop-scope'), { recursive: true });
|
||||
writeFileSync(
|
||||
join(root, 'trekresearch-loop-scope', `${SESSION}.json`),
|
||||
JSON.stringify({ runId: RUN_ID, startedAt: new Date().toISOString() }),
|
||||
);
|
||||
writeFileSync(
|
||||
join(root, 'trekresearch-loop-ledger.jsonl'),
|
||||
[
|
||||
...Array.from({ length: BUDGET }, (_, i) =>
|
||||
JSON.stringify({ ts: new Date().toISOString(), runId: RUN_ID, dimension: `d${i}`, effort: 'high', slot: i + 1 })),
|
||||
JSON.stringify({ ts: new Date().toISOString(), runId: RUN_ID, exhausted: true }),
|
||||
].join('\n') + '\n',
|
||||
);
|
||||
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: '',
|
||||
HOME: home,
|
||||
});
|
||||
assert.strictEqual(code, 2, 'the hook must find marker AND ledger under the fallback root and deny');
|
||||
});
|
||||
|
||||
test('pre-agent-cap allows under budget in the fallback root — the fallback is not a blanket deny', async () => {
|
||||
const home = mkdtempSync(join(tmpdir(), 'voyage-cap-home-'));
|
||||
const root = join(home, '.claude', 'voyage');
|
||||
mkdirSync(join(root, 'trekresearch-loop-scope'), { recursive: true });
|
||||
writeFileSync(
|
||||
join(root, 'trekresearch-loop-scope', `${SESSION}.json`),
|
||||
JSON.stringify({ runId: RUN_ID, startedAt: new Date().toISOString() }),
|
||||
);
|
||||
writeFileSync(join(root, 'trekresearch-loop-ledger.jsonl'), '');
|
||||
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: '',
|
||||
HOME: home,
|
||||
});
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// Append-only counting — the hook reads the ledger, it never writes it.
|
||||
// Writing per tool call would make the cap count its own enforcement.
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-agent-cap never writes to the ledger', async () => {
|
||||
const dir = fixture({ turns: 2 });
|
||||
const ledgerPath = join(dir, 'trekresearch-loop-ledger.jsonl');
|
||||
const before = readFileSync(ledgerPath, 'utf-8');
|
||||
await runHookWithEnv(CAP_HOOK, searchInput(), { ...CAPPED_ENV, CLAUDE_PLUGIN_DATA: dir });
|
||||
assert.strictEqual(readFileSync(ledgerPath, 'utf-8'), before);
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// Crash-time marker states. These are the states the TTL discussion in the
|
||||
// hook header anticipates, and none of them had a test: a marker written
|
||||
// half-way, and a marker whose runId never made it. Both must ALLOW — a
|
||||
// marker we cannot read cannot tell us which run we are in, and guessing
|
||||
// would deny tool calls in a session we know nothing about.
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-agent-cap ALLOWS on a partially written (corrupt) scope marker', async () => {
|
||||
const dir = fixture({ turns: BUDGET, exhausted: true });
|
||||
// Exactly what an interrupted printf leaves behind: valid prefix, no close.
|
||||
writeFileSync(join(dir, 'trekresearch-loop-scope', `${SESSION}.json`), '{"runId":"run-xyz789","star');
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0, 'an unparsable marker is not evidence of a loop turn');
|
||||
});
|
||||
|
||||
test('pre-agent-cap ALLOWS a marker that carries no runId', async () => {
|
||||
const dir = fixture({ turns: BUDGET, exhausted: true });
|
||||
writeFileSync(
|
||||
join(dir, 'trekresearch-loop-scope', `${SESSION}.json`),
|
||||
JSON.stringify({ startedAt: new Date().toISOString() }),
|
||||
);
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0, 'without a runId there are no ledger lines to count against');
|
||||
});
|
||||
|
||||
test('pre-agent-cap ALLOWS a marker whose runId is not a string', async () => {
|
||||
const dir = fixture({ turns: BUDGET, exhausted: true });
|
||||
// Truthy, so it clears the `!marker?.runId` guard and the session counts as
|
||||
// in scope — but readLedger compares runId with ===, so a number matches no
|
||||
// record and the run reads as 0 turns spent. Allow is the right answer either
|
||||
// way, which is why this stays a pin on the OUTCOME and not an argument for a
|
||||
// type guard: no writer emits a non-string runId, and the two routes are
|
||||
// indistinguishable from outside.
|
||||
writeFileSync(
|
||||
join(dir, 'trekresearch-loop-scope', `${SESSION}.json`),
|
||||
JSON.stringify({ runId: 5, startedAt: new Date().toISOString() }),
|
||||
);
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
test('pre-agent-cap ALLOWS a marker whose runId is the empty string', async () => {
|
||||
const dir = fixture({ turns: BUDGET, exhausted: true });
|
||||
writeFileSync(
|
||||
join(dir, 'trekresearch-loop-scope', `${SESSION}.json`),
|
||||
JSON.stringify({ runId: '', startedAt: new Date().toISOString() }),
|
||||
);
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// Malformed ledger lines — a truncated final write must not be counted as a
|
||||
// turn, and must not stop the well-formed lines around it from counting.
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-agent-cap does not count a malformed ledger line as a turn', async () => {
|
||||
const dir = fixture({ turns: BUDGET - 1 });
|
||||
const ledgerPath = join(dir, 'trekresearch-loop-ledger.jsonl');
|
||||
writeFileSync(ledgerPath, readFileSync(ledgerPath, 'utf-8') + '{"runId":"run-xyz789","dimen\n');
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 0, 'a half-written line is not a spent turn');
|
||||
});
|
||||
|
||||
test('pre-agent-cap still finds the tombstone with malformed lines around it', async () => {
|
||||
const dir = fixture({ turns: BUDGET, exhausted: true });
|
||||
const ledgerPath = join(dir, 'trekresearch-loop-ledger.jsonl');
|
||||
writeFileSync(ledgerPath, '{ garbage\n' + readFileSync(ledgerPath, 'utf-8') + 'also garbage\n');
|
||||
const { code } = await runHookWithEnv(CAP_HOOK, searchInput(), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: dir,
|
||||
});
|
||||
assert.strictEqual(code, 2, 'skipping bad lines must not mean skipping the run’s denial record');
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// The marker snippet is EXECUTED, not asserted about.
|
||||
//
|
||||
// Every existing pin on the marker lifecycle is a substring assertion on the
|
||||
// prose in commands/trekresearch.md. A snippet that emitted invalid JSON, or
|
||||
// wrote to a path the hook never looks at, would keep the whole suite green
|
||||
// while the hook silently allowed everything — which is the exact failure S82
|
||||
// found by hand. So these tests run the real shell blocks out of the command
|
||||
// file, with CLAUDE_PLUGIN_DATA stripped and HOME sandboxed, and then run the
|
||||
// real hook against what they produced.
|
||||
// -----------------------------------------------------------------------
|
||||
const CMD_FILE = join(ROOT, 'commands', 'trekresearch.md');
|
||||
|
||||
/** Pull the ```bash block that contains `needle` out of the command file. */
|
||||
function bashBlockContaining(needle) {
|
||||
const text = readFileSync(CMD_FILE, 'utf-8');
|
||||
const at = text.indexOf(needle);
|
||||
assert.ok(at > -1, `commands/trekresearch.md no longer contains ${JSON.stringify(needle)}`);
|
||||
const open = text.lastIndexOf('```bash', at);
|
||||
assert.ok(open > -1, `no \`\`\`bash fence opens before ${JSON.stringify(needle)}`);
|
||||
const bodyStart = text.indexOf('\n', open) + 1;
|
||||
const close = text.indexOf('```', bodyStart);
|
||||
assert.ok(close > bodyStart, 'unterminated bash fence');
|
||||
return text.slice(bodyStart, close);
|
||||
}
|
||||
|
||||
function runSnippet(snippet, env) {
|
||||
return execFileSync('bash', ['-c', snippet], {
|
||||
encoding: 'utf-8',
|
||||
env: { PATH: process.env.PATH, ...env },
|
||||
});
|
||||
}
|
||||
|
||||
test('the marker WRITE snippet lands valid JSON exactly where the hook looks for it', () => {
|
||||
const home = mkdtempSync(join(tmpdir(), 'voyage-snippet-'));
|
||||
const sessionId = 'snippet-session-1';
|
||||
const snippet = bashBlockContaining('Arms the PreToolUse cap').replace(/\{run_id\}/g, 'snippet-run-1');
|
||||
|
||||
runSnippet(snippet, { HOME: home, CLAUDE_CODE_SESSION_ID: sessionId });
|
||||
|
||||
const markerPath = join(home, '.claude', 'voyage', 'trekresearch-loop-scope', `${sessionId}.json`);
|
||||
assert.ok(existsSync(markerPath), `snippet wrote no marker at ${markerPath}`);
|
||||
const marker = JSON.parse(readFileSync(markerPath, 'utf-8')); // throws if the printf emits bad JSON
|
||||
assert.strictEqual(marker.runId, 'snippet-run-1', 'runId must be the same id passed to --run-id');
|
||||
assert.ok(Number.isFinite(Date.parse(marker.startedAt)), `startedAt must parse, got ${marker.startedAt}`);
|
||||
});
|
||||
|
||||
test('the marker snippet writes NO file when CLAUDE_CODE_SESSION_ID is empty', () => {
|
||||
const home = mkdtempSync(join(tmpdir(), 'voyage-snippet-'));
|
||||
const snippet = bashBlockContaining('Arms the PreToolUse cap').replace(/\{run_id\}/g, 'snippet-run-2');
|
||||
|
||||
const out = runSnippet(snippet, { HOME: home });
|
||||
|
||||
const scopeDir = join(home, '.claude', 'voyage', 'trekresearch-loop-scope');
|
||||
assert.ok(!existsSync(join(scopeDir, '.json')), 'an empty session id must not produce a `.json` marker');
|
||||
assert.match(out, /stays inert/i, 'the snippet must say the harness cap is inert, not fail silently');
|
||||
});
|
||||
|
||||
test('write snippet then real hook: the loop’s own writer arms the enforcement end to end', async () => {
|
||||
const home = mkdtempSync(join(tmpdir(), 'voyage-snippet-'));
|
||||
const sessionId = 'snippet-session-3';
|
||||
const runId = 'snippet-run-3';
|
||||
|
||||
runSnippet(
|
||||
bashBlockContaining('Arms the PreToolUse cap').replace(/\{run_id\}/g, runId),
|
||||
{ HOME: home, CLAUDE_CODE_SESSION_ID: sessionId },
|
||||
);
|
||||
// A spent, tombstoned ledger for that same runId, under the same resolved root.
|
||||
writeFileSync(
|
||||
join(home, '.claude', 'voyage', 'trekresearch-loop-ledger.jsonl'),
|
||||
[
|
||||
...Array.from({ length: BUDGET }, (_, i) =>
|
||||
JSON.stringify({ ts: new Date().toISOString(), runId, dimension: `d${i}`, effort: 'high', slot: i + 1 })),
|
||||
JSON.stringify({ ts: new Date().toISOString(), runId, exhausted: true }),
|
||||
].join('\n') + '\n',
|
||||
);
|
||||
|
||||
const denied = await runHookWithEnv(CAP_HOOK, searchInput(sessionId), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: '',
|
||||
HOME: home,
|
||||
});
|
||||
assert.strictEqual(denied.code, 2, 'the hook must find the snippet’s marker and enforce against it');
|
||||
|
||||
// And the removal snippet must disarm it again — same root, same guard.
|
||||
runSnippet(
|
||||
bashBlockContaining('Removal — idempotent').replace(/\{run_id\}/g, runId),
|
||||
{ HOME: home, CLAUDE_CODE_SESSION_ID: sessionId },
|
||||
);
|
||||
assert.ok(
|
||||
!existsSync(join(home, '.claude', 'voyage', 'trekresearch-loop-scope', `${sessionId}.json`)),
|
||||
'the removal snippet must delete the marker the write snippet created',
|
||||
);
|
||||
const allowed = await runHookWithEnv(CAP_HOOK, searchInput(sessionId), {
|
||||
...CAPPED_ENV,
|
||||
CLAUDE_PLUGIN_DATA: '',
|
||||
HOME: home,
|
||||
});
|
||||
assert.strictEqual(allowed.code, 0, 'a removed marker must take the session back out of scope');
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// Wiring — pattern from tests/hooks/hooks-json-stop-wired.test.mjs
|
||||
// -----------------------------------------------------------------------
|
||||
function invocationOf(h) {
|
||||
return [h.command || '', ...(h.args || [])].join(' ').trim();
|
||||
}
|
||||
|
||||
test('hooks.json wires pre-agent-cap.mjs on PreToolUse with ${CLAUDE_PLUGIN_ROOT}', () => {
|
||||
const cfg = JSON.parse(readFileSync(HOOKS_JSON, 'utf8'));
|
||||
const invocations = (cfg.hooks.PreToolUse || []).flatMap((entry) =>
|
||||
(entry.hooks || []).map(invocationOf),
|
||||
);
|
||||
const capInvocation = invocations.find((cmd) => cmd.includes('pre-agent-cap.mjs'));
|
||||
assert.ok(capInvocation, `no PreToolUse hook references pre-agent-cap.mjs. Found: ${JSON.stringify(invocations)}`);
|
||||
assert.match(capInvocation, /\$\{CLAUDE_PLUGIN_ROOT\}/, 'relative paths fail in headless sessions');
|
||||
assert.match(capInvocation, /^node\s+/);
|
||||
});
|
||||
|
||||
test('hooks.json matcher for pre-agent-cap covers the loop’s outbound surface', () => {
|
||||
const cfg = JSON.parse(readFileSync(HOOKS_JSON, 'utf8'));
|
||||
const entry = (cfg.hooks.PreToolUse || []).find((e) =>
|
||||
(e.hooks || []).some((h) => invocationOf(h).includes('pre-agent-cap.mjs')),
|
||||
);
|
||||
assert.ok(entry, 'pre-agent-cap entry missing from PreToolUse');
|
||||
for (const tool of ['WebSearch', 'WebFetch', 'Task']) {
|
||||
assert.match(entry.matcher, new RegExp(tool), `matcher must cover ${tool}`);
|
||||
}
|
||||
});
|
||||
|
|
@ -129,100 +129,6 @@ test('pre-bash-executor BLOCKS system shutdown command', async () => {
|
|||
assert.strictEqual(code, 2);
|
||||
});
|
||||
|
||||
test('pre-bash-executor BLOCKS a privileged halt at command position', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput('sudo shutdown -h now'));
|
||||
assert.strictEqual(code, 2);
|
||||
});
|
||||
|
||||
test('pre-bash-executor BLOCKS a destructive keyword after a separator', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput('echo done && poweroff'));
|
||||
assert.strictEqual(code, 2);
|
||||
});
|
||||
|
||||
// Bypasses opened when the rule was anchored to command position: whitespace
|
||||
// was collapsed BEFORE the pattern ran (killing the newline branch), `&` was
|
||||
// missing from the separator class, and a keyword handed to a shell wrapper
|
||||
// sits at command position without any separator in front of it.
|
||||
test('pre-bash-executor BLOCKS a destructive keyword after a newline separator', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput('echo done\npoweroff'));
|
||||
assert.strictEqual(code, 2);
|
||||
});
|
||||
|
||||
test('pre-bash-executor BLOCKS a destructive keyword after a background separator', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput('echo done & poweroff'));
|
||||
assert.strictEqual(code, 2);
|
||||
});
|
||||
|
||||
test('pre-bash-executor BLOCKS a destructive command wrapped in bash -c', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput('bash -c "poweroff"'));
|
||||
assert.strictEqual(code, 2);
|
||||
});
|
||||
|
||||
test('pre-bash-executor BLOCKS a destructive command wrapped in sh -c', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput("sh -c 'reboot'"));
|
||||
assert.strictEqual(code, 2);
|
||||
});
|
||||
|
||||
test('pre-bash-executor BLOCKS a destructive command handed to xargs', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput('echo x | xargs reboot'));
|
||||
assert.strictEqual(code, 2);
|
||||
});
|
||||
|
||||
test('pre-bash-executor BLOCKS a backslash-escaped destructive command', async () => {
|
||||
// `\reboot` runs reboot — the backslash suppresses alias expansion, nothing
|
||||
// else. The command-position anchor must see through it.
|
||||
const { code } = await runHook(PRE_BASH, bashInput('\\reboot'));
|
||||
assert.strictEqual(code, 2);
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// ALLOW — the same keywords as DATA, not at command position.
|
||||
// The rule matched the bare word anywhere in the string, so a quoted grep
|
||||
// pattern, ordinary prose, or a commit message that merely named the rule
|
||||
// was blocked. Anchoring to command position is what separates the two.
|
||||
// -----------------------------------------------------------------------
|
||||
test('pre-bash-executor ALLOWS the keyword inside a quoted grep pattern', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput("grep 'halt' f.mjs"));
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
// The change's own motivating case: a quoted grep alternation. Anchoring alone
|
||||
// did not reach it — the `|` inside the quotes reads as a separator unless
|
||||
// quoted spans are treated as data.
|
||||
test('pre-bash-executor ALLOWS a quoted grep alternation over the keywords', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput('grep "halt|poweroff" f.mjs'));
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
// Heredoc bodies are data too, and a newline separator is what makes them look
|
||||
// like command position. The rule's own comment names heredoc data as the
|
||||
// friction anchoring was meant to remove.
|
||||
test('pre-bash-executor ALLOWS the keyword at the start of a heredoc body line', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput('cat <<EOF\nreboot is a word here\nEOF'));
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
test('pre-bash-executor ALLOWS a commit message piped through a heredoc', async () => {
|
||||
const { code } = await runHook(
|
||||
PRE_BASH,
|
||||
bashInput("git commit -F - <<'MSG'\nhalt the loop on empty turns\nMSG"),
|
||||
);
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
test('pre-bash-executor ALLOWS the keyword inside echoed prose', async () => {
|
||||
const { code } = await runHook(PRE_BASH, bashInput('echo "we should halt here"'));
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
test('pre-bash-executor ALLOWS a commit message that names the rule', async () => {
|
||||
const { code } = await runHook(
|
||||
PRE_BASH,
|
||||
bashInput('git commit -m "fix(hooks): anchor shutdown rule to command position"'),
|
||||
);
|
||||
assert.strictEqual(code, 0);
|
||||
});
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// BLOCK — cron persistence
|
||||
// -----------------------------------------------------------------------
|
||||
|
|
|
|||
|
|
@ -31,25 +31,6 @@ test('SC #12: stats-sample.jsonl → expected.prom snapshot byte-for-byte match'
|
|||
` node scripts/gen-expected-prom.mjs > tests/fixtures/expected.prom`);
|
||||
});
|
||||
|
||||
test('Step 9: the four numeric STORM fields are metric families and effort is a label', () => {
|
||||
const expected = readFileSync(join(FIXTURES, 'expected.prom'), 'utf-8');
|
||||
for (const field of ['unique_sources', 'dimensions_baseline', 'conv_turns', 'empty_turns']) {
|
||||
assert.match(
|
||||
expected,
|
||||
new RegExp(`^# TYPE voyage_trekresearch_${field} `, 'm'),
|
||||
`${field} must appear as its own metric family — a numeric that never becomes a metric cannot be measured`,
|
||||
);
|
||||
}
|
||||
// effort is a low-cardinality string: it must ride along as a LABEL, never
|
||||
// as a metric family (a label is what makes high-vs-standard groupable).
|
||||
assert.match(expected, /effort="[a-z]+"/, 'effort must be emitted as a label');
|
||||
assert.doesNotMatch(
|
||||
expected,
|
||||
/^# TYPE voyage_trekresearch_effort /m,
|
||||
'effort must not become a metric family',
|
||||
);
|
||||
});
|
||||
|
||||
test('empty-input handling: [] returns empty string (no headers)', () => {
|
||||
assert.equal(transformToPrometheus([]), '');
|
||||
assert.equal(transformToPrometheus(null), '');
|
||||
|
|
|
|||
|
|
@ -14,7 +14,6 @@ import {
|
|||
POST_BASH_STATS_ALLOWED,
|
||||
EVENT_EMIT_PAYLOAD_ALLOWED,
|
||||
TOKEN_USAGE_ALLOWED,
|
||||
TREKRESEARCH_ALLOWED,
|
||||
} from '../../lib/exporters/field-allowlist.mjs';
|
||||
|
||||
// ---- path-validator: CWE-22 mitigation -------------------------------------
|
||||
|
|
@ -279,43 +278,6 @@ test('field-allowlist: token-usage INCLUDES numeric/label fields, EXCLUDES sessi
|
|||
assert.equal('cwd' in out, false, 'cwd MUST be stripped (CWE-212)');
|
||||
});
|
||||
|
||||
// ---- trekresearch allowlist: the `engine` field ----------------------------
|
||||
|
||||
test('field-allowlist: trekresearch INCLUDES engine, EXCLUDES question/project_dir/brief_path (two-sided)', () => {
|
||||
const record = {
|
||||
ts: '2026-08-09T12:00:00.000Z',
|
||||
question: 'which retrieval strategy survives contradiction?',
|
||||
mode: 'default',
|
||||
scope: 'both',
|
||||
engine: 'deep-research',
|
||||
slug: 'storm-upgrade',
|
||||
project_dir: '/Users/ktg/secret/project',
|
||||
brief_path: '/Users/ktg/secret/project/brief.md',
|
||||
dimensions: 4,
|
||||
agents_local: 7,
|
||||
agents_external: 4,
|
||||
gemini_used: false,
|
||||
confidence: 0.82,
|
||||
contradictions: 1,
|
||||
open_questions: 3,
|
||||
};
|
||||
const out = applyFieldAllowlist(record, 'trekresearch');
|
||||
// INCLUDED — low-cardinality label, emitted (trekresearch.md:533) and
|
||||
// promised in prose (:570-572); it was silently dropped before this pin.
|
||||
assert.equal('engine' in out, true, 'engine MUST be allowlisted — it is emitted and documented');
|
||||
assert.equal(out.engine, 'deep-research');
|
||||
assert.equal(out._schema_id, 'trekresearch');
|
||||
// EXCLUDED (CWE-212 boundary)
|
||||
assert.equal('question' in out, false, 'question MUST be stripped (prose, CWE-212)');
|
||||
assert.equal('project_dir' in out, false, 'project_dir MUST be stripped (path, CWE-212)');
|
||||
assert.equal('brief_path' in out, false, 'brief_path MUST be stripped (path, CWE-212)');
|
||||
});
|
||||
|
||||
test('field-allowlist: TREKRESEARCH_ALLOWED is frozen (drift-pin)', () => {
|
||||
assert.equal(Object.isFrozen(TREKRESEARCH_ALLOWED), true,
|
||||
'TREKRESEARCH_ALLOWED must be frozen — runtime mutation prevention');
|
||||
});
|
||||
|
||||
test('field-allowlist: null/undefined record handled safely', () => {
|
||||
assert.deepEqual(applyFieldAllowlist(null, 'trekplan'), {});
|
||||
assert.deepEqual(applyFieldAllowlist(undefined, 'trekplan'), {});
|
||||
|
|
|
|||
|
|
@ -127,39 +127,3 @@ test('non-orchestrator agents do NOT include the Agent tool (no recursive swarmi
|
|||
);
|
||||
}
|
||||
});
|
||||
|
||||
// M4 (v5.7.1): examples-relocation invariant. The 34 <example> blocks belong in
|
||||
// agent BODIES, not in the always-loaded `description:` frontmatter (voyage
|
||||
// launches its agents by name, so the auto-selection examples are cost without
|
||||
// function there). This pins the migration: examples cannot regress into
|
||||
// frontmatter, and cannot be silently lost during the move.
|
||||
function bodyOf(text) {
|
||||
// Slice the file content AFTER the closing frontmatter `---`.
|
||||
// Do NOT split on /^---$/m — a horizontal rule in the body would mis-split.
|
||||
const m = text.match(/^---\r?\n[\s\S]*?\r?\n---\r?\n?/);
|
||||
return m ? text.slice(m[0].length) : text;
|
||||
}
|
||||
|
||||
test('no agents/*.md frontmatter contains an <example> block (M4: examples live in the body)', () => {
|
||||
for (const f of agentFiles) {
|
||||
const fm = extractFrontmatter(read(`agents/${f}`));
|
||||
assert.ok(fm !== null, `agents/${f}: missing frontmatter`);
|
||||
assert.equal(
|
||||
(fm.match(/<example>/g) || []).length,
|
||||
0,
|
||||
`agents/${f}: <example> blocks must live in the body, not the always-loaded description: frontmatter (M4)`,
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('agent bodies retain at least 34 <example> blocks (M4: relocation moves, never deletes)', () => {
|
||||
let total = 0;
|
||||
for (const f of agentFiles) {
|
||||
total += (bodyOf(read(`agents/${f}`)).match(/<example>/g) || []).length;
|
||||
}
|
||||
assert.ok(
|
||||
total >= 34,
|
||||
`expected >= 34 <example> blocks across agent bodies (17 agents x 2), got ${total} ` +
|
||||
`— examples may have been deleted instead of relocated (M4)`,
|
||||
);
|
||||
});
|
||||
|
|
|
|||
|
|
@ -26,7 +26,6 @@ import { fileURLToPath } from 'node:url';
|
|||
import { parseDocument } from '../../lib/util/frontmatter.mjs';
|
||||
import { resolveProfile, loadProfile } from '../../lib/profiles/resolver.mjs';
|
||||
import { STATES } from '../../lib/util/autonomy-gate.mjs';
|
||||
import { TREKRESEARCH_ALLOWED } from '../../lib/exporters/field-allowlist.mjs';
|
||||
|
||||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
const ROOT = join(HERE, '..', '..');
|
||||
|
|
@ -919,7 +918,7 @@ test('S15: default-profile name is consistent across resolver + all profile docs
|
|||
assert.equal(profile_source, 'default', 'resolveProfile({}, {}) must report source=default');
|
||||
assert.equal(def, 'premium', 'resolver hardcoded default is premium (operator decision 2026-05-13, commit 40d8742)');
|
||||
|
||||
const OTHERS = ['economy', 'balanced', 'premium', 'fable'].filter((p) => p !== def);
|
||||
const OTHERS = ['economy', 'balanced', 'premium'].filter((p) => p !== def);
|
||||
for (const doc of PROFILE_DOCS) {
|
||||
const body = read(doc);
|
||||
assert.ok(
|
||||
|
|
@ -942,7 +941,7 @@ test('S15: default-profile name is consistent across resolver + all profile docs
|
|||
test('S15: profile tables encode each built-in yaml phase_models exactly', () => {
|
||||
// Column order in every profile table: Profile | Brief | Research | Plan | Execute | Review | Continue | Use case
|
||||
const PHASES = ['brief', 'research', 'plan', 'execute', 'review', 'continue'];
|
||||
for (const name of ['economy', 'balanced', 'premium', 'fable']) {
|
||||
for (const name of ['economy', 'balanced', 'premium']) {
|
||||
const pm = loadProfile(name).phase_models; // {brief:'opus', ...}
|
||||
const expected = PHASES.map((ph) => pm[ph]);
|
||||
for (const doc of PROFILE_DOCS) {
|
||||
|
|
@ -962,31 +961,6 @@ test('S15: profile tables encode each built-in yaml phase_models exactly', () =>
|
|||
}
|
||||
});
|
||||
|
||||
// STRUCTURAL pin: the exporter allowlist and the authoring fixture must agree.
|
||||
// `engine` was emitted, documented in prose, and still dropped at the export
|
||||
// boundary because nothing tied the two together. Derives one side from the
|
||||
// frozen Set, so it survives rewording of the fixture row.
|
||||
test('S74: every TREKRESEARCH_ALLOWED name is declared in the jsonl-schemas fixture row', () => {
|
||||
const row = read('tests/fixtures/jsonl-schemas.md')
|
||||
.split('\n')
|
||||
.find((l) => l.startsWith('| trekresearch-stats '));
|
||||
assert.ok(row, 'jsonl-schemas.md is missing the `trekresearch-stats` row');
|
||||
// Columns: '' | schema_id | fields | writer_path | line_ref | v4.1 additive | PII | ''
|
||||
const cells = row.split('|').map((c) => c.trim());
|
||||
// Both columns are required: profile/profile_source/parallel_agents live in
|
||||
// the `v4.1 additive` column only, so checking `fields` alone fails at once.
|
||||
const declared = new Set(
|
||||
[cells[2], cells[5]].flatMap((c) => c.split(',').map((f) => f.trim())).filter(Boolean),
|
||||
);
|
||||
for (const name of TREKRESEARCH_ALLOWED) {
|
||||
assert.ok(
|
||||
declared.has(name),
|
||||
`\`${name}\` is allowlisted in field-allowlist.mjs but absent from the fixture row's `
|
||||
+ '`fields` + `v4.1 additive` columns — fix the SOURCE, not this pin',
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
// --- S34 (V30) — economy is self-declared experimental until the cross-tier
|
||||
// Jaccard floor (0.55) is empirically calibrated (Step-17 calibration deferred
|
||||
// to v4.2). The status must be visible in BOTH the profile data
|
||||
|
|
@ -1091,264 +1065,6 @@ test('S18: --min-brief-version is documented at the trekplan + trekresearch boun
|
|||
}
|
||||
});
|
||||
|
||||
test('deep-research-engine: --engine is documented + consistent across surfaces', () => {
|
||||
// Cross-doc pin mirroring S18 (:1057). The opt-in external-research engine flag
|
||||
// must be discoverable wherever /trekresearch flags live: the command itself
|
||||
// plus the three reference surfaces.
|
||||
for (const f of ['commands/trekresearch.md', 'docs/command-modes.md', 'CLAUDE.md', 'README.md']) {
|
||||
assert.ok(
|
||||
read(f).includes('--engine'),
|
||||
`${f} must document the --engine flag (deep-research-engine)`,
|
||||
);
|
||||
}
|
||||
// README documents it specifically as an **Engine** mode-table row.
|
||||
assert.ok(
|
||||
/\*\*Engine\*\*/.test(read('README.md')),
|
||||
'README.md must document --engine as an **Engine** mode row',
|
||||
);
|
||||
// The command prose must name both engine values and the swarm fallback, so the
|
||||
// opt-in + graceful-degradation contract is pinned — not merely the flag string.
|
||||
const research = read('commands/trekresearch.md');
|
||||
assert.ok(/\bswarm\b/.test(research), 'trekresearch.md must name the swarm engine value');
|
||||
assert.ok(/\bdeep-research\b/.test(research), 'trekresearch.md must name the deep-research engine value');
|
||||
assert.ok(
|
||||
/fall back|falls back/.test(research),
|
||||
'trekresearch.md must document the swarm fallback (graceful degradation)',
|
||||
);
|
||||
});
|
||||
|
||||
// ── STORM bounded loop — env-vars documented across the four surfaces ──────
|
||||
// Same cross-doc shape as the --engine pin above. An operator-facing switch
|
||||
// documented on one surface is a switch most operators never find; and the
|
||||
// three below decide cost, so they are the ones worth pinning.
|
||||
|
||||
const STORM_SURFACES = ['docs/command-modes.md', 'CLAUDE.md', 'README.md', 'docs/architecture.md'];
|
||||
const STORM_ENV_VARS = ['VOYAGE_STORM_ENABLED', 'TREKRESEARCH_MAX_CONV_TURNS', 'VOYAGE_DISABLE_CAP_HOOK'];
|
||||
|
||||
for (const envVar of STORM_ENV_VARS) {
|
||||
test(`STORM: ${envVar} is documented on all four reference surfaces`, () => {
|
||||
for (const f of STORM_SURFACES) {
|
||||
assert.ok(read(f).includes(envVar), `${f} must document ${envVar} (STORM bounded-loop env-vars)`);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
test('STORM: VOYAGE_STORM_ENABLED is documented WITH its default-off contract', () => {
|
||||
// Naming a switch without its default is how a default-on mechanism ships by
|
||||
// accident. Default-off is also the decline branch: declining costs nothing.
|
||||
for (const f of STORM_SURFACES) {
|
||||
const t = read(f);
|
||||
const i = t.indexOf('VOYAGE_STORM_ENABLED');
|
||||
const window = t.slice(Math.max(0, i - 400), i + 400);
|
||||
assert.match(
|
||||
window,
|
||||
/default-off|default off|opt-in/i,
|
||||
`${f}: VOYAGE_STORM_ENABLED must be documented together with its default-off contract`,
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
// The flag gates BOTH STORM phases, not just the Phase 5 loop. A surface that
|
||||
// scopes it to "the loop" tells an operator that unsetting it still leaves
|
||||
// Phase 4.5 discovery running — which was true until the guard was fixed, and
|
||||
// is the half-off state the decline branch cannot survive.
|
||||
test('STORM: VOYAGE_STORM_ENABLED is documented as gating BOTH phases, not the loop alone', () => {
|
||||
for (const f of STORM_SURFACES) {
|
||||
const t = read(f);
|
||||
const i = t.indexOf('VOYAGE_STORM_ENABLED');
|
||||
const window = t.slice(Math.max(0, i - 400), i + 400);
|
||||
assert.match(
|
||||
window,
|
||||
/4\.5|discovery|dimension discovery|both/i,
|
||||
`${f}: VOYAGE_STORM_ENABLED must be documented as gating Phase 4.5 too, not only the Phase 5 loop`,
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
// The variable is empty in the Bash tool env, so "missing CLAUDE_PLUGIN_DATA
|
||||
// denies" described a loop that could never spend turn 1. The root is resolved
|
||||
// in code now; a doc that still promises the deny describes a mechanism the
|
||||
// code does not have.
|
||||
test('STORM: no surface claims a missing CLAUDE_PLUGIN_DATA denies — the root falls back', () => {
|
||||
for (const f of ['docs/architecture.md', 'CLAUDE.md', 'README.md', 'docs/command-modes.md']) {
|
||||
const t = read(f);
|
||||
assert.doesNotMatch(
|
||||
t,
|
||||
/(missing|no|absent) `?CLAUDE_PLUGIN_DATA`?[^.\n]*(denies|fails closed)/i,
|
||||
`${f}: CLAUDE_PLUGIN_DATA absence no longer denies — it resolves to ~/.claude/voyage`,
|
||||
);
|
||||
}
|
||||
assert.match(
|
||||
read('docs/architecture.md'),
|
||||
/~\/\.claude\/voyage/,
|
||||
'docs/architecture.md must name the fallback data root the cap and the hook share',
|
||||
);
|
||||
});
|
||||
|
||||
test('STORM: README research-dimension prose stays at the existing 3–8 ceiling', () => {
|
||||
// Phase 4.5 discovers dimensions UNDER settings.json:16's maxDimensions: 8.
|
||||
// Rewriting this prose upward would raise a ceiling the brief asked us to hold.
|
||||
assert.ok(
|
||||
read('README.md').includes('3–8 research dimensions'),
|
||||
'README.md must keep the "3–8 research dimensions" prose — augmentation happens under the existing cap, it does not raise it',
|
||||
);
|
||||
});
|
||||
|
||||
// The Phase 4.5 amendment to the Independence hard rule crosses that rule
|
||||
// deliberately: discovered dimensions are mined from Phase-4 output, which holds
|
||||
// local-agent findings, so a discovered dimension can steer an external query.
|
||||
// The crossing is defensible — bounded, disclosed, and resolving a tension the
|
||||
// brief created itself. What was not defensible was naming query-privacy-gate.mjs
|
||||
// as its compensating control: that gate inspects outbound query CONTENT for
|
||||
// paths, repo identifiers and secret-shaped strings. It compensates the EGRESS
|
||||
// risk the crossing creates. It cannot stop a local finding from shaping an
|
||||
// external agent's question, so the bias risk was left with no named control
|
||||
// while the text read as though it had one.
|
||||
test('STORM: the Independence amendment does not name the privacy gate as the BIAS control', () => {
|
||||
const t = read('commands/trekresearch.md');
|
||||
const amendment = t.slice(t.indexOf('**Independence:**'), t.indexOf('**Graceful degradation:**'));
|
||||
assert.ok(amendment.length > 100, 'the Independence hard rule and its amendment must still be present');
|
||||
assert.ok(
|
||||
!/compensating\s+control\s+is\s+`query-privacy-gate/.test(amendment),
|
||||
'query-privacy-gate.mjs compensates egress, not bias — naming it as THE compensating control for the ' +
|
||||
'Independence crossing claims a control the gate cannot provide',
|
||||
);
|
||||
// The bias risk must carry a control that actually bears on bias.
|
||||
assert.match(
|
||||
amendment,
|
||||
/contrarian-researcher/,
|
||||
'the amendment must name the control that does bear on bias — contrarian-researcher runs unconditionally ' +
|
||||
'at effort: high, the only effort at which the crossing happens',
|
||||
);
|
||||
assert.match(
|
||||
amendment,
|
||||
/egress/i,
|
||||
'the privacy gate should still be named, as the control for the egress risk the same crossing creates',
|
||||
);
|
||||
});
|
||||
|
||||
test('STORM: no banned Sonnet-swarm phrase introduced on any STORM surface', () => {
|
||||
const BANNED = [
|
||||
'Sonnet exploration',
|
||||
'Sonnet runs the exploration',
|
||||
'front-loads cheap Sonnet',
|
||||
'exploration agents stay on Sonnet',
|
||||
];
|
||||
for (const f of STORM_SURFACES) {
|
||||
const t = read(f);
|
||||
for (const phrase of BANNED) {
|
||||
assert.ok(!t.includes(phrase), `${f} must not claim "${phrase}" — sub-agents are opus-pinned`);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
// ── STORM bounded loop — Phase 5 scope-marker wiring (S79) ────────────────
|
||||
// hooks/scripts/pre-agent-cap.mjs enforces the loop bound ONLY while a scope
|
||||
// marker exists for the calling session. Nothing wrote that marker, so the
|
||||
// hook shipped correct-but-latent. These pins bind the two ends of one
|
||||
// contract: the reader (the hook) and the writer (Phase 5 prose). Both sides
|
||||
// are derived from the hook source where possible, so drift in EITHER
|
||||
// direction fails — renaming the directory in the hook breaks the prose pin
|
||||
// just as rewriting the prose does.
|
||||
|
||||
const CAP_HOOK_SRC = read('hooks/scripts/pre-agent-cap.mjs');
|
||||
const RESEARCH_CMD = read('commands/trekresearch.md');
|
||||
|
||||
function phase5Section(text) {
|
||||
const start = text.indexOf('## Phase 5');
|
||||
const end = text.indexOf('## Phase 6', start);
|
||||
assert.ok(start > 0 && end > start, 'commands/trekresearch.md must keep ## Phase 5 … ## Phase 6');
|
||||
return text.slice(start, end);
|
||||
}
|
||||
|
||||
test('STORM marker: the directory the hook reads is the directory Phase 5 writes', () => {
|
||||
const m = CAP_HOOK_SRC.match(/const SCOPE_DIRNAME = '([^']+)'/);
|
||||
assert.ok(m, 'pre-agent-cap.mjs must keep SCOPE_DIRNAME as a single-quoted literal');
|
||||
const scopeDir = m[1];
|
||||
assert.ok(
|
||||
phase5Section(RESEARCH_CMD).includes(scopeDir),
|
||||
`commands/trekresearch.md Phase 5 must write the scope marker under ${scopeDir}/ — a hook keyed on a directory nobody writes is latent, not enforcing`,
|
||||
);
|
||||
assert.ok(
|
||||
phase5Section(RESEARCH_CMD).includes('CLAUDE_PLUGIN_DATA'),
|
||||
'Phase 5 must root the marker at CLAUDE_PLUGIN_DATA — the same root the hook resolves',
|
||||
);
|
||||
});
|
||||
|
||||
test('STORM marker: both payload fields the hook reads are named in Phase 5', () => {
|
||||
// The hook rejects a marker without runId, and treats an unparsable
|
||||
// startedAt as stale. A writer that emits neither name produces a marker
|
||||
// that is silently ignored.
|
||||
const phase5 = phase5Section(RESEARCH_CMD);
|
||||
for (const field of ['runId', 'startedAt']) {
|
||||
assert.ok(
|
||||
CAP_HOOK_SRC.includes(`marker.${field}`) || CAP_HOOK_SRC.includes(`marker?.${field}`),
|
||||
`pre-agent-cap.mjs must still read marker.${field}`,
|
||||
);
|
||||
assert.ok(
|
||||
phase5.includes(field),
|
||||
`commands/trekresearch.md Phase 5 must write the ${field} field the hook reads`,
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('STORM marker: Phase 5 keys the marker filename by the harness session id', () => {
|
||||
// Verified 2026-08-12: the session_id on a PreToolUse payload equals
|
||||
// $CLAUDE_CODE_SESSION_ID for the same session. The marker filename is the
|
||||
// scope key — key it by anything else and the hook never matches.
|
||||
assert.ok(
|
||||
phase5Section(RESEARCH_CMD).includes('CLAUDE_CODE_SESSION_ID'),
|
||||
'Phase 5 must name CLAUDE_CODE_SESSION_ID as the marker filename key (== the hook payload session_id)',
|
||||
);
|
||||
});
|
||||
|
||||
test('STORM marker: only Phase 5 writes it — Phase 4.5 does not', () => {
|
||||
// The hook contract says "a marker file only the Phase 5 loop writes".
|
||||
// Phase 4.5 mines already-retrieved Phase-4 results and spends no loop
|
||||
// turns; scoping enforcement there widens the window for nothing.
|
||||
const m = CAP_HOOK_SRC.match(/const SCOPE_DIRNAME = '([^']+)'/);
|
||||
const scopeDir = m[1];
|
||||
const start = RESEARCH_CMD.indexOf('## Phase 4.5');
|
||||
const end = RESEARCH_CMD.indexOf('## Phase 5', start);
|
||||
assert.ok(start > 0 && end > start, 'commands/trekresearch.md must keep ## Phase 4.5 … ## Phase 5');
|
||||
assert.ok(
|
||||
!RESEARCH_CMD.slice(start, end).includes(scopeDir),
|
||||
'Phase 4.5 must not write the scope marker — only the Phase 5 loop does',
|
||||
);
|
||||
});
|
||||
|
||||
test('STORM marker: every one of the three loop exits removes the marker', () => {
|
||||
// A marker left behind after the cap is spent denies WebSearch/WebFetch/Task
|
||||
// for the REST of the session — Phase 6 synthesis spawns agents. Cleanup on
|
||||
// the exhausted exit is what keeps enforcement from becoming a session brick.
|
||||
const phase5 = phase5Section(RESEARCH_CMD);
|
||||
const exitsStart = phase5.indexOf('### Exits');
|
||||
assert.ok(exitsStart > 0, 'Phase 5 must keep the ### Exits section');
|
||||
const exits = phase5.slice(exitsStart, phase5.indexOf('###', exitsStart + 5));
|
||||
for (const [n, label] of [['1.', 'converged'], ['2.', 'cap exhausted'], ['3.', 'operator stop']]) {
|
||||
const from = exits.indexOf(`\n${n}`);
|
||||
assert.ok(from > 0, `Exits must keep numbered entry ${n} (${label})`);
|
||||
const nextMarker = exits.indexOf(`\n${Number(n[0]) + 1}.`, from);
|
||||
const entry = exits.slice(from, nextMarker > from ? nextMarker : undefined);
|
||||
assert.match(
|
||||
entry,
|
||||
/remove the (scope )?marker|rm -f/i,
|
||||
`Exit ${n} (${label}) must remove the scope marker — a marker outliving the loop blocks the rest of the session`,
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
test('STORM marker: the crash path is documented as TTL auto-reset, not cleanup', () => {
|
||||
// A crashed session runs no cleanup at all. The honest statement is that the
|
||||
// hook's TTL covers it; claiming cleanup covers crashes would be false.
|
||||
const phase5 = phase5Section(RESEARCH_CMD);
|
||||
assert.match(
|
||||
phase5,
|
||||
/TTL/,
|
||||
'Phase 5 must state that a crashed run is covered by the hook TTL, not by exit cleanup',
|
||||
);
|
||||
});
|
||||
|
||||
test('S18: HANDOVER-CONTRACTS documents the pre-2.2 zero-framing-enforcement hole', () => {
|
||||
// The framing defense is producer-elective: a brief declaring ≤ 2.1 sidesteps
|
||||
// it entirely. Handover 1 (PUBLIC CONTRACT) must disclose this and name the remedy.
|
||||
|
|
@ -1550,45 +1266,3 @@ test('S38: forward-guard — no product-facing doc asserts main-context relief u
|
|||
offenders.join('\n'),
|
||||
);
|
||||
});
|
||||
|
||||
// ── v5.9 (fable-tier step 6) — composed-resolver wiring pin ────────────────
|
||||
// The four pipeline commands must resolve {effort, model} via the composed
|
||||
// CLI (`resolver.mjs --resolve-phase-model`, brief > profile > default). A
|
||||
// direct `phase-signal-resolver.mjs --brief` invocation in a command Bash
|
||||
// block is the brief-only CLI: it silently drops the profile layer (the
|
||||
// AP2-1 footgun — `--profile <x>` would never reach sub-agent spawns again).
|
||||
// If this pin fails, re-wire the command to the composed CLI — do not relax
|
||||
// the pin.
|
||||
|
||||
for (const cmd of ['trekresearch', 'trekplan', 'trekreview', 'trekexecute']) {
|
||||
test(`v5.9: commands/${cmd}.md invokes the composed resolver, not the brief-only CLI`, () => {
|
||||
const text = read(`commands/${cmd}.md`);
|
||||
assert.ok(
|
||||
text.includes('--resolve-phase-model'),
|
||||
`commands/${cmd}.md must invoke resolver.mjs --resolve-phase-model (composed brief > profile > default)`,
|
||||
);
|
||||
assert.ok(
|
||||
!/phase-signal-resolver\.mjs --brief/.test(text),
|
||||
`commands/${cmd}.md must not invoke the brief-only phase-signal-resolver CLI — the profile layer would be silently dropped`,
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
// ── v5.9 (fable-tier step 7) — orchestrator model-pin ABSENCE pin ───────────
|
||||
// Command frontmatter deliberately omits `model:` — omission is the only
|
||||
// session-inheritance spelling documented by BOTH official surfaces (AP2-2:
|
||||
// the `inherit` literal is disputed between skills.md and the plugin-dev
|
||||
// command-frontmatter reference). Re-adding a pin locks the orchestrator to a
|
||||
// fixed model in every session; if deterministic pinning is ever wanted again,
|
||||
// do it consciously and update this pin's rationale.
|
||||
|
||||
test('v5.9: no commands/*.md frontmatter carries a model: key (session inheritance by omission)', () => {
|
||||
const offenders = [];
|
||||
for (const f of listMd('commands')) {
|
||||
const doc = parseDocument(read(`commands/${f}`));
|
||||
const fm = doc.parsed && doc.parsed.frontmatter;
|
||||
if (fm && 'model' in fm) offenders.push(f);
|
||||
}
|
||||
assert.deepEqual(offenders, [],
|
||||
`command frontmatter must omit model: (orchestrator follows the session model); offenders: ${offenders.join(', ')}`);
|
||||
});
|
||||
|
|
|
|||
|
|
@ -1,49 +0,0 @@
|
|||
// tests/lib/gold-eval.test.mjs
|
||||
// SKAL-1·4b — the offline gold-scored output eval (the scoring RUN).
|
||||
//
|
||||
// This is the third test-census category (see lib/util/test-census.mjs):
|
||||
// neither a behavior unit test nor a doc-consistency pin, but a SCORING RUN —
|
||||
// it feeds a committed agent-run fixture through the deterministic coordinator
|
||||
// contract (4a) and scores the result against the golden corpus at
|
||||
// (file, rule_key) granularity. Offline: committed reviewer payloads, no live
|
||||
// agent spawn, no LLM, no network. The all-agree foundation; the
|
||||
// LLM-in-the-loop eval is the separate 4c tier.
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { join, dirname } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
import { runContract } from '../../lib/review/coordinator-contract.mjs';
|
||||
import { scoreFindings, scoreVerdict } from '../../lib/review/gold-scorer.mjs';
|
||||
|
||||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
const ROOT = join(HERE, '..', '..');
|
||||
const gold = JSON.parse(readFileSync(join(ROOT, 'tests/fixtures/bakeoff-rich/gold.json'), 'utf-8'));
|
||||
const runPerfect = JSON.parse(
|
||||
readFileSync(join(ROOT, 'tests/fixtures/bakeoff-rich/runs/run-perfect.json'), 'utf-8'),
|
||||
);
|
||||
|
||||
test('gold-eval — run-perfect reproduces gold at (file,rule_key): precision/recall/f1 = 1.0', () => {
|
||||
const result = runContract(runPerfect);
|
||||
const score = scoreFindings(result.findings, gold.findings);
|
||||
assert.equal(score.precision, 1, `precision ${score.precision}; spurious=${JSON.stringify(score.spurious)}`);
|
||||
assert.equal(score.recall, 1, `recall ${score.recall}; missed=${JSON.stringify(score.missed)}`);
|
||||
assert.equal(score.f1, 1);
|
||||
assert.equal(score.tp, gold.findings.length);
|
||||
});
|
||||
|
||||
test('gold-eval — run-perfect coordinator verdict matches gold expected_verdict (BLOCK)', () => {
|
||||
const result = runContract(runPerfect);
|
||||
assert.equal(result.verdict, gold.expected_verdict);
|
||||
assert.ok(scoreVerdict(result.verdict, gold.expected_verdict));
|
||||
});
|
||||
|
||||
test('gold-eval — run-perfect drops nothing (no suppressed, no schema-skipped payloads)', () => {
|
||||
// A clean run is the regression guard: any future contract change that
|
||||
// silently suppresses or skips one of the 5 seeded findings breaks here.
|
||||
const result = runContract(runPerfect);
|
||||
assert.equal(result.skipped.length, 0, `skipped payloads: ${JSON.stringify(result.skipped)}`);
|
||||
assert.equal(result.suppressed.length, 0, `suppressed findings: ${result.suppressed.length}`);
|
||||
assert.equal(result.findings.length, gold.findings.length);
|
||||
});
|
||||
|
|
@ -1,126 +0,0 @@
|
|||
// tests/lib/gold-scorer.test.mjs
|
||||
// SKAL-1·4b — behavior test for the gold scorer.
|
||||
//
|
||||
// scoreFindings compares a recorded agent-run's findings against a golden
|
||||
// corpus record at (file, rule_key) granularity (line + severity deliberately
|
||||
// ignored). This pins the precision/recall/f1 math AND the discriminating
|
||||
// paths (false negatives + false positives), not merely the all-match case —
|
||||
// a scorer that always returned 1.0 would pass an all-match-only test.
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { scoreFindings, scoreVerdict } from '../../lib/review/gold-scorer.mjs';
|
||||
|
||||
// A minimal gold set: two distinct (file, rule_key) pairs.
|
||||
const GOLD = [
|
||||
{ file: 'a.mjs', line: 10, rule_key: 'SECURITY_INJECTION', severity: 'BLOCKER' },
|
||||
{ file: 'b.mjs', line: 0, rule_key: 'MISSING_TEST', severity: 'MAJOR' },
|
||||
];
|
||||
|
||||
test('scoreFindings — perfect match scores precision/recall/f1 = 1', () => {
|
||||
// Same pairs, different line/severity (must be ignored at (file,rule_key) granularity).
|
||||
const run = [
|
||||
{ file: 'a.mjs', line: 99, rule_key: 'SECURITY_INJECTION', severity: 'MINOR' },
|
||||
{ file: 'b.mjs', line: 7, rule_key: 'MISSING_TEST', severity: 'BLOCKER' },
|
||||
];
|
||||
const s = scoreFindings(run, GOLD);
|
||||
assert.equal(s.tp, 2);
|
||||
assert.equal(s.fp, 0);
|
||||
assert.equal(s.fn, 0);
|
||||
assert.equal(s.precision, 1);
|
||||
assert.equal(s.recall, 1);
|
||||
assert.equal(s.f1, 1);
|
||||
});
|
||||
|
||||
test('scoreFindings — a missed gold pair is a false negative (recall < 1)', () => {
|
||||
const run = [{ file: 'a.mjs', line: 10, rule_key: 'SECURITY_INJECTION', severity: 'BLOCKER' }];
|
||||
const s = scoreFindings(run, GOLD);
|
||||
assert.equal(s.tp, 1);
|
||||
assert.equal(s.fp, 0);
|
||||
assert.equal(s.fn, 1);
|
||||
assert.equal(s.precision, 1);
|
||||
assert.equal(s.recall, 0.5);
|
||||
assert.deepEqual(s.missed, ['b.mjs MISSING_TEST']);
|
||||
});
|
||||
|
||||
test('scoreFindings — a spurious pair is a false positive (precision < 1)', () => {
|
||||
const run = [
|
||||
{ file: 'a.mjs', line: 10, rule_key: 'SECURITY_INJECTION', severity: 'BLOCKER' },
|
||||
{ file: 'b.mjs', line: 0, rule_key: 'MISSING_TEST', severity: 'MAJOR' },
|
||||
{ file: 'c.mjs', line: 3, rule_key: 'PLACEHOLDER_IN_CODE', severity: 'MAJOR' },
|
||||
];
|
||||
const s = scoreFindings(run, GOLD);
|
||||
assert.equal(s.tp, 2);
|
||||
assert.equal(s.fp, 1);
|
||||
assert.equal(s.fn, 0);
|
||||
assert.equal(s.recall, 1);
|
||||
assert.equal(s.precision, 2 / 3);
|
||||
assert.deepEqual(s.spurious, ['c.mjs PLACEHOLDER_IN_CODE']);
|
||||
});
|
||||
|
||||
test('scoreFindings — mixed FN + FP', () => {
|
||||
const run = [
|
||||
{ file: 'a.mjs', line: 10, rule_key: 'SECURITY_INJECTION', severity: 'BLOCKER' }, // match
|
||||
{ file: 'c.mjs', line: 3, rule_key: 'PLACEHOLDER_IN_CODE', severity: 'MAJOR' }, // spurious
|
||||
];
|
||||
const s = scoreFindings(run, GOLD);
|
||||
assert.equal(s.tp, 1);
|
||||
assert.equal(s.fp, 1);
|
||||
assert.equal(s.fn, 1);
|
||||
assert.equal(s.precision, 0.5);
|
||||
assert.equal(s.recall, 0.5);
|
||||
assert.equal(s.f1, 0.5);
|
||||
});
|
||||
|
||||
test('scoreFindings — duplicate (file,rule_key) pairs collapse (set semantics)', () => {
|
||||
const run = [
|
||||
{ file: 'a.mjs', line: 10, rule_key: 'SECURITY_INJECTION', severity: 'BLOCKER' },
|
||||
{ file: 'a.mjs', line: 22, rule_key: 'SECURITY_INJECTION', severity: 'BLOCKER' }, // same pair
|
||||
{ file: 'b.mjs', line: 0, rule_key: 'MISSING_TEST', severity: 'MAJOR' },
|
||||
];
|
||||
const s = scoreFindings(run, GOLD);
|
||||
assert.equal(s.tp, 2);
|
||||
assert.equal(s.fp, 0);
|
||||
assert.equal(s.fn, 0);
|
||||
});
|
||||
|
||||
test('scoreFindings — empty run: recall 0 (nothing found), precision vacuously 1, f1 0', () => {
|
||||
const s = scoreFindings([], GOLD);
|
||||
assert.equal(s.tp, 0);
|
||||
assert.equal(s.fp, 0);
|
||||
assert.equal(s.fn, 2);
|
||||
assert.equal(s.recall, 0);
|
||||
assert.equal(s.precision, 1);
|
||||
assert.equal(s.f1, 0);
|
||||
});
|
||||
|
||||
test('scoreFindings — empty gold: precision 0 (all spurious), recall vacuously 1, f1 0', () => {
|
||||
const run = [{ file: 'a.mjs', line: 10, rule_key: 'SECURITY_INJECTION', severity: 'BLOCKER' }];
|
||||
const s = scoreFindings(run, []);
|
||||
assert.equal(s.tp, 0);
|
||||
assert.equal(s.fp, 1);
|
||||
assert.equal(s.fn, 0);
|
||||
assert.equal(s.precision, 0);
|
||||
assert.equal(s.recall, 1);
|
||||
assert.equal(s.f1, 0);
|
||||
});
|
||||
|
||||
test('scoreFindings — both empty: precision/recall/f1 = 1 (matched nothing perfectly)', () => {
|
||||
const s = scoreFindings([], []);
|
||||
assert.equal(s.precision, 1);
|
||||
assert.equal(s.recall, 1);
|
||||
assert.equal(s.f1, 1);
|
||||
});
|
||||
|
||||
test('scoreFindings — tolerates null/undefined finding arrays', () => {
|
||||
const s = scoreFindings(null, undefined);
|
||||
assert.equal(s.tp, 0);
|
||||
assert.equal(s.fp, 0);
|
||||
assert.equal(s.fn, 0);
|
||||
});
|
||||
|
||||
test('scoreVerdict — exact verdict match is true, mismatch is false', () => {
|
||||
assert.equal(scoreVerdict('BLOCK', 'BLOCK'), true);
|
||||
assert.equal(scoreVerdict('WARN', 'BLOCK'), false);
|
||||
assert.equal(scoreVerdict('ALLOW', 'ALLOW'), true);
|
||||
});
|
||||
|
|
@ -50,7 +50,7 @@ test('resolvePhaseSignal — defensive: null/non-object input returns null', ()
|
|||
|
||||
test('resolvePhaseSignal — drops model not in BASE_ALLOWED_MODELS (defense-in-depth gate)', () => {
|
||||
// MAJOR fix (S4): line that copies `model` must gate against the same
|
||||
// allowlist brief-validator uses (BASE_ALLOWED_MODELS = ['sonnet','opus','fable']),
|
||||
// allowlist brief-validator uses (BASE_ALLOWED_MODELS = ['sonnet','opus']),
|
||||
// mirroring how effort is gated against EFFORT_LEVELS. A brief that slipped
|
||||
// validation (hand-edited, validation skipped) must not hand a junk model
|
||||
// string to a command that then spawns an agent with `model: <junk>`.
|
||||
|
|
@ -70,17 +70,15 @@ test('resolvePhaseSignal — drops model not in BASE_ALLOWED_MODELS (defense-in-
|
|||
assert.ok(!('model' in review), 'model key absent for haiku');
|
||||
});
|
||||
|
||||
test('resolvePhaseSignal — keeps valid models (sonnet, opus, fable) after gating', () => {
|
||||
test('resolvePhaseSignal — keeps valid models (sonnet, opus) after gating', () => {
|
||||
const fm = {
|
||||
phase_signals: [
|
||||
{ phase: 'research', effort: 'low', model: 'sonnet' },
|
||||
{ phase: 'execute', effort: 'high', model: 'opus' },
|
||||
{ phase: 'review', effort: 'high', model: 'fable' },
|
||||
],
|
||||
};
|
||||
assert.equal(resolvePhaseSignal(fm, 'research').model, 'sonnet');
|
||||
assert.equal(resolvePhaseSignal(fm, 'execute').model, 'opus');
|
||||
assert.equal(resolvePhaseSignal(fm, 'review').model, 'fable');
|
||||
});
|
||||
|
||||
test('resolvePhaseSignalFromFile + CLI shim — writes JSON to stdout, exit 0', () => {
|
||||
|
|
|
|||
|
|
@ -50,20 +50,6 @@ test('SC #5: loadProfile("premium") returns all-opus', () => {
|
|||
}
|
||||
});
|
||||
|
||||
test('SC #5: loadProfile("fable") returns all-fable (BUILTIN_NAMES canary)', () => {
|
||||
// Throws PROFILE_NOT_FOUND if 'fable' regresses out of BUILTIN_NAMES —
|
||||
// the silent-fallback-to-premium failure mode this pin exists to catch.
|
||||
const p = loadProfile('fable');
|
||||
assert.equal(p.name, 'fable');
|
||||
for (const phase of ['brief', 'research', 'plan', 'execute', 'review', 'continue']) {
|
||||
assert.equal(p.phase_models[phase], 'fable', `fable ${phase} should be fable`);
|
||||
}
|
||||
assert.equal(p.parallel_agents_min, 6);
|
||||
assert.equal(p.parallel_agents_max, 8);
|
||||
assert.equal(p.external_research_enabled, true);
|
||||
assert.equal(p.brief_reviewer_iter_cap, 3);
|
||||
});
|
||||
|
||||
test('SC #5: loadProfile throws PROFILE_NOT_FOUND for unknown profile', () => {
|
||||
try {
|
||||
loadProfile('does-not-exist-xyz');
|
||||
|
|
|
|||
|
|
@ -60,30 +60,3 @@ test('resolvePhaseModel — Case 6 (defensive): null briefPath falls through to
|
|||
assert.equal(r.model, 'opus', 'premium.plan default = opus');
|
||||
assert.equal(r.source, 'default');
|
||||
});
|
||||
|
||||
// v5.9 — fable profile composition (brief SC 4) + effort passthrough coherence
|
||||
|
||||
test('resolvePhaseModel — --profile fable: all six phases resolve model fable (no brief signal)', () => {
|
||||
for (const phase of ['brief', 'research', 'plan', 'execute', 'review', 'continue']) {
|
||||
const r = resolvePhaseModel(phase, null, ['--profile', 'fable'], {});
|
||||
assert.equal(r.model, 'fable', `fable.${phase} should be fable; got ${JSON.stringify(r)}`);
|
||||
assert.equal(r.source, 'flag');
|
||||
}
|
||||
});
|
||||
|
||||
test('resolvePhaseModel — brief signal (opus) beats --profile fable (composition precedence pin)', () => {
|
||||
// brief-effort-high.md pins execute to model: opus. Brief must beat the fable profile.
|
||||
const r = resolvePhaseModel('execute', FIXTURE('brief-effort-high.md'), ['--profile', 'fable'], {});
|
||||
assert.equal(r.model, 'opus', `brief signal should beat fable profile; got ${JSON.stringify(r)}`);
|
||||
assert.equal(r.source, 'brief-signal');
|
||||
});
|
||||
|
||||
test('resolvePhaseModel — composed output carries {effort, model} atomically (v5.9 passthrough)', () => {
|
||||
// brief-effort-high.md: execute → {effort: high, model: opus}. The composed
|
||||
// resolver must return BOTH fields from one call — the split-CLI design this
|
||||
// passthrough replaced could drift effort and model apart.
|
||||
const r = resolvePhaseModel('execute', FIXTURE('brief-effort-high.md'), [], {});
|
||||
assert.equal(r.effort, 'high', `effort must pass through; got ${JSON.stringify(r)}`);
|
||||
assert.equal(r.model, 'opus');
|
||||
assert.equal(r.source, 'brief-signal');
|
||||
});
|
||||
|
|
|
|||
|
|
@ -94,81 +94,6 @@ test('SC #11(b): commands/trekplan.md prose mentions phase_models + parallel_age
|
|||
'trekplan.md prose must mention parallel_agents (additive stats field)');
|
||||
});
|
||||
|
||||
// --- Step 9: the five STORM measurement fields -----------------------------
|
||||
// Four numerics + one low-cardinality label. Without `effort` there is no axis
|
||||
// to group high-vs-standard runs on, and the measurement gate cannot be
|
||||
// computed at all.
|
||||
|
||||
const STORM_FIELDS = [
|
||||
'effort',
|
||||
'unique_sources',
|
||||
'dimensions_baseline',
|
||||
'conv_turns',
|
||||
'empty_turns',
|
||||
];
|
||||
|
||||
test('Step 9: a standard-run trekresearch record parses and survives applyFieldAllowlist', async () => {
|
||||
const { applyFieldAllowlist } = await import('../../lib/exporters/field-allowlist.mjs');
|
||||
// A standard run: loop never armed, so no discovered dimensions and no turns.
|
||||
const raw = JSON.parse(JSON.stringify({
|
||||
_schema_id: 'trekresearch',
|
||||
ts: '2026-08-09T12:00:00.000Z',
|
||||
slug: 'add-auth',
|
||||
mode: 'default',
|
||||
scope: 'both',
|
||||
engine: 'swarm',
|
||||
question: 'free prose that must never reach the exporter',
|
||||
project_dir: '/Users/somebody/repos/x',
|
||||
brief_path: '/Users/somebody/repos/x/brief.md',
|
||||
dimensions: 4,
|
||||
dimensions_baseline: 4,
|
||||
conv_turns: 0,
|
||||
empty_turns: 0,
|
||||
unique_sources: 11,
|
||||
effort: 'standard',
|
||||
agents_local: 5,
|
||||
agents_external: 4,
|
||||
gemini_used: false,
|
||||
confidence: 0.8,
|
||||
contradictions: 1,
|
||||
open_questions: 2,
|
||||
profile: 'premium',
|
||||
profile_source: 'default',
|
||||
}));
|
||||
|
||||
assert.equal(raw.dimensions_baseline, raw.dimensions,
|
||||
'a standard run discovers no dimensions — baseline must equal the final count');
|
||||
|
||||
const out = applyFieldAllowlist(raw, 'trekresearch');
|
||||
for (const field of STORM_FIELDS) {
|
||||
assert.ok(field in out, `${field} must survive the trekresearch allowlist`);
|
||||
}
|
||||
assert.equal(out.conv_turns, 0);
|
||||
assert.equal(out.empty_turns, 0);
|
||||
assert.equal(out.effort, 'standard');
|
||||
// Deny-by-omission must still hold for the PII-ish fields.
|
||||
for (const denied of ['question', 'project_dir', 'brief_path']) {
|
||||
assert.equal(denied in out, false, `${denied} must NOT reach the exporter`);
|
||||
}
|
||||
});
|
||||
|
||||
test('Step 9: commands/trekresearch.md prose names all five measurement fields', () => {
|
||||
const content = readFileSync(join(REPO_ROOT, 'commands', 'trekresearch.md'), 'utf-8');
|
||||
for (const field of STORM_FIELDS) {
|
||||
assert.ok(content.includes(field),
|
||||
`trekresearch.md prose must name ${field} — an emitted-but-undocumented field is unauditable`);
|
||||
}
|
||||
});
|
||||
|
||||
test('Step 9: tests/fixtures/jsonl-schemas.md trekresearch row lists the five fields', () => {
|
||||
const doc = readFileSync(join(REPO_ROOT, 'tests', 'fixtures', 'jsonl-schemas.md'), 'utf-8');
|
||||
const row = doc.split('\n').find(l => l.startsWith('| trekresearch-stats '));
|
||||
assert.ok(row, 'jsonl-schemas.md is missing the trekresearch-stats row');
|
||||
for (const field of STORM_FIELDS) {
|
||||
assert.ok(row.includes(field), `authoring reference must list ${field}`);
|
||||
}
|
||||
});
|
||||
|
||||
test('SC #11(b): commands/trekresearch.md prose mentions external_research_enabled', () => {
|
||||
const content = readFileSync(join(REPO_ROOT, 'commands', 'trekresearch.md'), 'utf-8');
|
||||
assert.match(content, /external_research_enabled/,
|
||||
|
|
|
|||
|
|
@ -1,597 +0,0 @@
|
|||
// tests/lib/research-loop-cap.test.mjs
|
||||
// Cover lib/util/research-loop-cap.mjs: default-off, worst-case arithmetic,
|
||||
// anti-dead-data (different caps → different denial points), statefulness
|
||||
// (identical args → different answers once the budget is hit), env
|
||||
// coercion, fail-closed on missing CLAUDE_PLUGIN_DATA, and the CLI shim.
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { execFileSync, execFile } from 'node:child_process';
|
||||
import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, existsSync } from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
import {
|
||||
allowTurn,
|
||||
isStormEnabled,
|
||||
resolveMaxConvTurns,
|
||||
resolveLedgerPath,
|
||||
resolveDataRoot,
|
||||
readLedger,
|
||||
checkDimensionCeiling,
|
||||
MAX_CONV_TURNS,
|
||||
MAX_TOTAL_DIMENSIONS,
|
||||
} from '../../lib/util/research-loop-cap.mjs';
|
||||
|
||||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
const SHIM = join(HERE, '..', '..', 'lib', 'util', 'research-loop-cap.mjs');
|
||||
|
||||
function withTmpDataDir(fn) {
|
||||
const dir = mkdtempSync(join(tmpdir(), 'research-loop-cap-'));
|
||||
try {
|
||||
return fn(dir);
|
||||
} finally {
|
||||
rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
}
|
||||
|
||||
function runShim(args, env) {
|
||||
try {
|
||||
const out = execFileSync(process.execPath, [SHIM, ...args], {
|
||||
encoding: 'utf-8',
|
||||
stdio: ['ignore', 'pipe', 'pipe'],
|
||||
env: { ...process.env, ...env },
|
||||
});
|
||||
return { code: 0, out };
|
||||
} catch (e) {
|
||||
return { code: e.status ?? 1, out: e.stdout?.toString() ?? '' };
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Run the shim in a BUILT env rather than an inherited one. Spreading
|
||||
* process.env means no test can express "CLAUDE_PLUGIN_DATA is absent" — the
|
||||
* exact condition that holds in every real run — so the shim's behaviour there
|
||||
* went uncovered while the module was denying turn 1.
|
||||
*/
|
||||
function runShimStripped(args, env = {}) {
|
||||
try {
|
||||
const out = execFileSync(process.execPath, [SHIM, ...args], {
|
||||
encoding: 'utf-8',
|
||||
stdio: ['ignore', 'pipe', 'pipe'],
|
||||
env: { PATH: process.env.PATH, ...env },
|
||||
});
|
||||
return { code: 0, out };
|
||||
} catch (e) {
|
||||
return { code: e.status ?? 1, out: e.stdout?.toString() ?? '' };
|
||||
}
|
||||
}
|
||||
|
||||
// ---- (a) default-off --------------------------------------------------------
|
||||
|
||||
test('allowTurn — VOYAGE_STORM_ENABLED unset denies with budget 0, regardless of effort', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir };
|
||||
const r = allowTurn({ runId: 'r1', dimension: 'd1', effort: 'high' }, { env });
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.reason, 'storm_disabled');
|
||||
assert.equal(r.budget, 0);
|
||||
});
|
||||
});
|
||||
|
||||
test('allowTurn — VOYAGE_STORM_ENABLED=0 denies same as unset', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '0' };
|
||||
const r = allowTurn({ runId: 'r1', dimension: 'd1', effort: 'high' }, { env });
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.reason, 'storm_disabled');
|
||||
});
|
||||
});
|
||||
|
||||
test('allowTurn — enabled but effort !== high denies with budget 0', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1' };
|
||||
const r = allowTurn({ runId: 'r1', dimension: 'd1', effort: 'standard' }, { env });
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.reason, 'effort_not_high');
|
||||
assert.equal(r.budget, 0);
|
||||
});
|
||||
});
|
||||
|
||||
// ---- (f) CLAUDE_PLUGIN_DATA unset => documented fallback root ----------------
|
||||
//
|
||||
// CLAUDE_PLUGIN_DATA is EMPTY in the Bash tool's process env (measured in a
|
||||
// live plugin-enabled session), and the Bash snippet in commands/trekresearch.md
|
||||
// is the module's only caller. Denying on its absence therefore denied turn 1
|
||||
// of every real run: the loop could never spend a turn, and the pre-registered
|
||||
// measurement could not be run at all. The root is resolved in code, not
|
||||
// demanded of the environment.
|
||||
|
||||
test('allowTurn — CLAUDE_PLUGIN_DATA unset falls back to the documented root and grants', () => {
|
||||
withTmpDataDir((home) => {
|
||||
const env = { VOYAGE_STORM_ENABLED: '1', HOME: home }; // no CLAUDE_PLUGIN_DATA
|
||||
const r = allowTurn({ runId: 'r1', dimension: 'd1', effort: 'high' }, { env });
|
||||
assert.equal(r.ok, true, 'the loop must be able to spend turn 1 without CLAUDE_PLUGIN_DATA');
|
||||
assert.equal(r.used, 1);
|
||||
assert.equal(r.budget, MAX_CONV_TURNS * MAX_TOTAL_DIMENSIONS);
|
||||
assert.ok(
|
||||
existsSync(join(home, '.claude', 'voyage', 'trekresearch-loop-ledger.jsonl')),
|
||||
'the ledger must be written under the fallback root',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
test('resolveDataRoot — CLAUDE_PLUGIN_DATA wins; empty or unset falls back to ~/.claude/voyage', () => {
|
||||
assert.equal(resolveDataRoot({ CLAUDE_PLUGIN_DATA: '/tmp/plugin-data' }), '/tmp/plugin-data');
|
||||
assert.equal(resolveDataRoot({ CLAUDE_PLUGIN_DATA: '', HOME: '/home/x' }), join('/home/x', '.claude', 'voyage'));
|
||||
assert.equal(resolveDataRoot({ HOME: '/home/x' }), join('/home/x', '.claude', 'voyage'));
|
||||
});
|
||||
|
||||
// ---- (c) worst-case arithmetic ----------------------------------------------
|
||||
|
||||
test('allowTurn — budget is max_conv_turns × max_total_dimensions (default 3×8=24)', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1' };
|
||||
const r = allowTurn({ runId: 'r1', dimension: 'd1', effort: 'high' }, { env });
|
||||
assert.equal(r.ok, true);
|
||||
assert.equal(r.budget, 24);
|
||||
assert.equal(r.used, 1);
|
||||
});
|
||||
});
|
||||
|
||||
test('allowTurn — grants exactly `budget` turns then denies the next one (default 24)', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1' };
|
||||
let last;
|
||||
for (let i = 0; i < 24; i++) {
|
||||
last = allowTurn({ runId: 'r-exhaust', dimension: `d${i % 8}`, effort: 'high' }, { env });
|
||||
assert.equal(last.ok, true, `turn ${i + 1} should be granted`);
|
||||
}
|
||||
const denied = allowTurn({ runId: 'r-exhaust', dimension: 'd0', effort: 'high' }, { env });
|
||||
assert.equal(denied.ok, false);
|
||||
assert.equal(denied.reason, 'budget_exhausted');
|
||||
assert.equal(denied.used, 24);
|
||||
assert.equal(denied.budget, 24);
|
||||
});
|
||||
});
|
||||
|
||||
// ---- (c)/(anti-dead-data) — different caps → observably different denial points
|
||||
|
||||
test('allowTurn — TREKRESEARCH_MAX_CONV_TURNS=1 denies after 8 turns (1×8), not 24', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1', TREKRESEARCH_MAX_CONV_TURNS: '1' };
|
||||
let last;
|
||||
for (let i = 0; i < 8; i++) {
|
||||
last = allowTurn({ runId: 'r-narrow', dimension: `d${i}`, effort: 'high' }, { env });
|
||||
assert.equal(last.ok, true, `turn ${i + 1} should be granted`);
|
||||
}
|
||||
const denied = allowTurn({ runId: 'r-narrow', dimension: 'd8', effort: 'high' }, { env });
|
||||
assert.equal(denied.ok, false);
|
||||
assert.equal(denied.budget, 8);
|
||||
assert.notEqual(denied.budget, 24, 'a narrower cap must produce a different denial point than the default');
|
||||
});
|
||||
});
|
||||
|
||||
// ---- (d) stateful — identical args give different answers once exhausted ---
|
||||
|
||||
test('allowTurn — identical {runId, dimension, effort} args diverge once the budget is hit (proves statefulness)', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1', TREKRESEARCH_MAX_CONV_TURNS: '1' };
|
||||
const args = { runId: 'r-identical', dimension: 'same-dim', effort: 'high' };
|
||||
const results = [];
|
||||
for (let i = 0; i < 9; i++) results.push(allowTurn(args, { env }));
|
||||
// First 8 (budget = 1*8) granted, 9th denied — same exact input object each time.
|
||||
assert.deepEqual(results.slice(0, 8).map(r => r.ok), Array(8).fill(true));
|
||||
assert.equal(results[8].ok, false);
|
||||
assert.equal(results[8].reason, 'budget_exhausted');
|
||||
});
|
||||
});
|
||||
|
||||
// ---- (e) env coercion --------------------------------------------------------
|
||||
|
||||
test('resolveMaxConvTurns — NaN string falls back to default', () => {
|
||||
assert.equal(resolveMaxConvTurns({ TREKRESEARCH_MAX_CONV_TURNS: 'abc' }), MAX_CONV_TURNS);
|
||||
});
|
||||
|
||||
test('resolveMaxConvTurns — empty string falls back to default', () => {
|
||||
assert.equal(resolveMaxConvTurns({ TREKRESEARCH_MAX_CONV_TURNS: '' }), MAX_CONV_TURNS);
|
||||
});
|
||||
|
||||
test('resolveMaxConvTurns — negative value falls back to default', () => {
|
||||
assert.equal(resolveMaxConvTurns({ TREKRESEARCH_MAX_CONV_TURNS: '-5' }), MAX_CONV_TURNS);
|
||||
});
|
||||
|
||||
test('resolveMaxConvTurns — zero falls back to default (never unbounded)', () => {
|
||||
assert.equal(resolveMaxConvTurns({ TREKRESEARCH_MAX_CONV_TURNS: '0' }), MAX_CONV_TURNS);
|
||||
});
|
||||
|
||||
test('resolveMaxConvTurns — unset falls back to default', () => {
|
||||
assert.equal(resolveMaxConvTurns({}), MAX_CONV_TURNS);
|
||||
});
|
||||
|
||||
test('resolveMaxConvTurns — valid positive integer string is honored', () => {
|
||||
assert.equal(resolveMaxConvTurns({ TREKRESEARCH_MAX_CONV_TURNS: '2' }), 2);
|
||||
});
|
||||
|
||||
// A fractional value passed the `n <= 0` guard and only THEN floored, so 0.5 and
|
||||
// 0.9 became 0 and the budget became 0 × 8 = 0 — every turn denied, the loop
|
||||
// silently dead, while README.md and docs/architecture.md both promise a
|
||||
// fallback of 3. The guard has to see the floored value, not the raw one.
|
||||
test('resolveMaxConvTurns — a fractional value below 1 falls back to the default, never 0', () => {
|
||||
assert.equal(resolveMaxConvTurns({ TREKRESEARCH_MAX_CONV_TURNS: '0.5' }), MAX_CONV_TURNS);
|
||||
assert.equal(resolveMaxConvTurns({ TREKRESEARCH_MAX_CONV_TURNS: '0.9' }), MAX_CONV_TURNS);
|
||||
});
|
||||
|
||||
test('resolveMaxConvTurns — a fractional value above 1 still floors (2.7 → 2)', () => {
|
||||
assert.equal(resolveMaxConvTurns({ TREKRESEARCH_MAX_CONV_TURNS: '2.7' }), 2);
|
||||
});
|
||||
|
||||
test('resolveMaxConvTurns — Infinity is not a cap and falls back to the default', () => {
|
||||
assert.equal(resolveMaxConvTurns({ TREKRESEARCH_MAX_CONV_TURNS: 'Infinity' }), MAX_CONV_TURNS);
|
||||
});
|
||||
|
||||
test('allowTurn — a fractional cap below 1 cannot produce a budget of 0', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1', TREKRESEARCH_MAX_CONV_TURNS: '0.5' };
|
||||
const r = allowTurn({ runId: 'r-frac', dimension: 'd1', effort: 'high' }, { env });
|
||||
assert.equal(r.ok, true, 'a budget of 0 would make the loop silently dead, not bounded');
|
||||
assert.equal(r.budget, MAX_CONV_TURNS * MAX_TOTAL_DIMENSIONS);
|
||||
});
|
||||
});
|
||||
|
||||
// ---- pure-core unit coverage --------------------------------------------------
|
||||
|
||||
test('isStormEnabled — only the literal string "1" enables', () => {
|
||||
assert.equal(isStormEnabled({ VOYAGE_STORM_ENABLED: '1' }), true);
|
||||
assert.equal(isStormEnabled({ VOYAGE_STORM_ENABLED: 'true' }), false);
|
||||
assert.equal(isStormEnabled({}), false);
|
||||
});
|
||||
|
||||
test('resolveLedgerPath — falls back under ~/.claude/voyage when CLAUDE_PLUGIN_DATA is unset or empty', () => {
|
||||
const expected = join('/home/x', '.claude', 'voyage', 'trekresearch-loop-ledger.jsonl');
|
||||
assert.equal(resolveLedgerPath({ HOME: '/home/x' }), expected);
|
||||
assert.equal(resolveLedgerPath({ CLAUDE_PLUGIN_DATA: '', HOME: '/home/x' }), expected);
|
||||
});
|
||||
|
||||
test('resolveLedgerPath — joins CLAUDE_PLUGIN_DATA with the ledger filename', () => {
|
||||
const p = resolveLedgerPath({ CLAUDE_PLUGIN_DATA: '/tmp/plugin-data' });
|
||||
assert.equal(p, join('/tmp/plugin-data', 'trekresearch-loop-ledger.jsonl'));
|
||||
});
|
||||
|
||||
test('allowTurn — missing runId or dimension denies with missing_args', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1' };
|
||||
const r1 = allowTurn({ dimension: 'd1', effort: 'high' }, { env });
|
||||
assert.equal(r1.ok, false);
|
||||
assert.equal(r1.reason, 'missing_args');
|
||||
const r2 = allowTurn({ runId: 'r1', effort: 'high' }, { env });
|
||||
assert.equal(r2.ok, false);
|
||||
assert.equal(r2.reason, 'missing_args');
|
||||
});
|
||||
});
|
||||
|
||||
// ---- fail-closed on an unreadable ledger ------------------------------------
|
||||
//
|
||||
// The module's own header states that a budget control "must never silently
|
||||
// grant unlimited turns just because the data dir is missing". The
|
||||
// missing-DIRECTORY case already failed closed; the unreadable-FILE case
|
||||
// returned 0 from countTurns and therefore re-granted the full budget on every
|
||||
// call, unbounded — a fail-open in the same module that argues against one.
|
||||
// A directory standing where the ledger file belongs reproduces it portably
|
||||
// (EISDIR), with no chmod that a root test runner would ignore.
|
||||
|
||||
function withUnreadableLedger(fn) {
|
||||
return withTmpDataDir((dir) => {
|
||||
mkdirSync(join(dir, 'trekresearch-loop-ledger.jsonl'), { recursive: true });
|
||||
return fn(dir);
|
||||
});
|
||||
}
|
||||
|
||||
test('readLedger — a missing ledger is 0 turns, not an error (turn 1 must be grantable)', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
assert.equal(readLedger(join(dir, 'nope.jsonl'), 'r1').granted, 0);
|
||||
});
|
||||
});
|
||||
|
||||
test('readLedger — an unreadable ledger throws rather than reporting 0 turns spent', () => {
|
||||
withUnreadableLedger((dir) => {
|
||||
assert.throws(
|
||||
() => readLedger(join(dir, 'trekresearch-loop-ledger.jsonl'), 'r1'),
|
||||
/unreadable/i,
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
test('readLedger — malformed lines are skipped, well-formed ones for the run still count', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const p = join(dir, 'trekresearch-loop-ledger.jsonl');
|
||||
writeFileSync(p, [
|
||||
JSON.stringify({ runId: 'r1', dimension: 'd1' }),
|
||||
'{ not json',
|
||||
'',
|
||||
JSON.stringify({ runId: 'other', dimension: 'd1' }),
|
||||
JSON.stringify({ runId: 'r1', dimension: 'd2' }),
|
||||
].join('\n') + '\n');
|
||||
assert.equal(readLedger(p, 'r1').granted, 2);
|
||||
});
|
||||
});
|
||||
|
||||
test('allowTurn — an unreadable ledger DENIES the turn instead of granting a fresh budget', () => {
|
||||
withUnreadableLedger((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1' };
|
||||
const r = allowTurn({ runId: 'r-unreadable', dimension: 'd1', effort: 'high' }, { env });
|
||||
assert.equal(r.ok, false, 'a budget control that cannot count must not grant');
|
||||
assert.match(r.reason, /ledger-read-failed/);
|
||||
});
|
||||
});
|
||||
|
||||
// ---- the bound holds under concurrency --------------------------------------
|
||||
//
|
||||
// countTurns-then-appendFileSync had no atomic claim, while the comment above
|
||||
// allowTurn asserted "Append-only: never read-modify-write" and named the
|
||||
// concurrent case (Phase 4.5/5 may spawn several agents in one message) as the
|
||||
// reason. The decision path WAS read-then-write: N callers that all observe
|
||||
// used == budget-1 all grant, and the bound is exceeded by N-1.
|
||||
//
|
||||
// Two of these tests are deterministic. They do not race anything — they assert
|
||||
// the invariant the claim introduces: a slot that is already claimed is spent,
|
||||
// even when the ledger has not caught up yet, which is exactly the state a
|
||||
// mid-append competitor leaves behind. The third runs real processes.
|
||||
|
||||
function seedLedger(dir, runId, n) {
|
||||
writeFileSync(
|
||||
join(dir, 'trekresearch-loop-ledger.jsonl'),
|
||||
Array.from({ length: n }, (_, i) =>
|
||||
JSON.stringify({ ts: new Date().toISOString(), runId, dimension: `d${i}`, effort: 'high' }),
|
||||
).join('\n') + (n ? '\n' : ''),
|
||||
);
|
||||
}
|
||||
|
||||
function seedClaims(dir, runId, slots) {
|
||||
const claimDir = join(dir, 'trekresearch-loop-claims');
|
||||
mkdirSync(claimDir, { recursive: true });
|
||||
for (const s of slots) writeFileSync(join(claimDir, `${runId}-${s}.claim`), '');
|
||||
}
|
||||
|
||||
function runShimAsync(args, env) {
|
||||
return new Promise((resolve) => {
|
||||
execFile(
|
||||
process.execPath,
|
||||
[SHIM, ...args],
|
||||
{ encoding: 'utf-8', env: { PATH: process.env.PATH, ...env } },
|
||||
(err, stdout) => resolve({ code: err ? (err.code ?? 1) : 0, out: stdout ?? '' }),
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
test('allowTurn — a slot already claimed is spent even when the ledger has not caught up', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1', TREKRESEARCH_MAX_CONV_TURNS: '1' };
|
||||
// budget = 1 × 8 = 8. Ledger shows 7 turns; a competitor already claimed
|
||||
// slot 8 and has not appended yet. Counting the ledger alone says "one slot
|
||||
// free" and grants a 9th turn overall — the breach this closes.
|
||||
seedLedger(dir, 'r-race', 7);
|
||||
seedClaims(dir, 'r-race', [1, 2, 3, 4, 5, 6, 7, 8]);
|
||||
const r = allowTurn({ runId: 'r-race', dimension: 'd0', effort: 'high' }, { env });
|
||||
assert.equal(r.ok, false, 'every slot up to the budget is claimed, so there is nothing to grant');
|
||||
assert.equal(r.reason, 'budget_exhausted');
|
||||
assert.equal(r.budget, 8);
|
||||
});
|
||||
});
|
||||
|
||||
test('allowTurn — it takes the first FREE slot and claims it, so a repeat call cannot retake it', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1', TREKRESEARCH_MAX_CONV_TURNS: '1' };
|
||||
seedLedger(dir, 'r-slot', 7);
|
||||
seedClaims(dir, 'r-slot', [1, 2, 3, 4, 5, 6, 7]);
|
||||
const first = allowTurn({ runId: 'r-slot', dimension: 'd0', effort: 'high' }, { env });
|
||||
assert.equal(first.ok, true, 'slot 8 is free and must be grantable');
|
||||
assert.equal(first.used, 8, 'used is the slot number, so it never double-counts a claimed slot');
|
||||
assert.ok(
|
||||
existsSync(join(dir, 'trekresearch-loop-claims', 'r-slot-8.claim')),
|
||||
'the grant must leave the claim behind as the atomic record of the slot',
|
||||
);
|
||||
const second = allowTurn({ runId: 'r-slot', dimension: 'd0', effort: 'high' }, { env });
|
||||
assert.equal(second.ok, false);
|
||||
assert.equal(second.reason, 'budget_exhausted');
|
||||
});
|
||||
});
|
||||
|
||||
test('allowTurn — a different runId is unaffected by another run’s claims', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1', TREKRESEARCH_MAX_CONV_TURNS: '1' };
|
||||
seedClaims(dir, 'r-other', [1, 2, 3, 4, 5, 6, 7, 8]);
|
||||
const r = allowTurn({ runId: 'r-mine', dimension: 'd0', effort: 'high' }, { env });
|
||||
assert.equal(r.ok, true, 'claims are per-run; one run must not exhaust another');
|
||||
});
|
||||
});
|
||||
|
||||
test('allowTurn — parallel processes at the boundary cannot exceed the budget', async () => {
|
||||
const dir = mkdtempSync(join(tmpdir(), 'research-loop-cap-par-'));
|
||||
try {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1', TREKRESEARCH_MAX_CONV_TURNS: '1' };
|
||||
seedLedger(dir, 'r-par', 7); // budget 8 → exactly one turn left
|
||||
const results = await Promise.all(
|
||||
Array.from({ length: 6 }, (_, i) =>
|
||||
runShimAsync(['--run-id', 'r-par', '--dimension', `p${i}`, '--effort', 'high'], env),
|
||||
),
|
||||
);
|
||||
const winners = results.filter((r) => r.code === 0).length;
|
||||
assert.equal(winners, 1, `exactly one of six concurrent callers may take the last slot, got ${winners}`);
|
||||
|
||||
// Granted turns, not raw lines: the five denied callers also record
|
||||
// exhaustion tombstones, and a tombstone is not a turn.
|
||||
const { granted } = readLedger(join(dir, 'trekresearch-loop-ledger.jsonl'), 'r-par');
|
||||
assert.equal(granted, 8, `granted turns must never exceed the budget, got ${granted}`);
|
||||
} finally {
|
||||
rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
// ---- the exhaustion tombstone -----------------------------------------------
|
||||
//
|
||||
// allowTurn appends BEFORE the turn runs, so during granted turn N the ledger
|
||||
// holds N records. The hook denied at `used >= budget`, which blocked every
|
||||
// tool call of the FINAL granted turn: the primitive granted B turns and the
|
||||
// harness permitted B-1. Worse, an exhausted run then always terminated through
|
||||
// an exit-2 tool denial instead of the graceful "cap exhausted" exit at
|
||||
// commands/trekresearch.md, which is the only exit the prose teaches.
|
||||
//
|
||||
// Letting the hook allow at `used == budget` fixes the count but would leave it
|
||||
// unable to catch the one case it exists for — the loop consults the gate, is
|
||||
// denied, and issues the tool call anyway. So the denial itself becomes a
|
||||
// record: a tombstone the hook can see.
|
||||
|
||||
test('allowTurn — denying for budget writes an exhaustion tombstone', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1', TREKRESEARCH_MAX_CONV_TURNS: '1' };
|
||||
for (let i = 0; i < 8; i++) {
|
||||
assert.equal(allowTurn({ runId: 'r-tomb', dimension: `d${i}`, effort: 'high' }, { env }).ok, true);
|
||||
}
|
||||
const ledgerPath = join(dir, 'trekresearch-loop-ledger.jsonl');
|
||||
assert.equal(readLedger(ledgerPath, 'r-tomb').exhausted, 0, 'no tombstone before the gate has denied anything');
|
||||
|
||||
const denied = allowTurn({ runId: 'r-tomb', dimension: 'd0', effort: 'high' }, { env });
|
||||
assert.equal(denied.ok, false);
|
||||
assert.equal(denied.reason, 'budget_exhausted');
|
||||
|
||||
const after = readLedger(ledgerPath, 'r-tomb');
|
||||
assert.equal(after.exhausted, 1, 'the denial must leave a record the harness gate can read');
|
||||
assert.equal(after.granted, 8, 'a tombstone is not a granted turn and must not count as one');
|
||||
});
|
||||
});
|
||||
|
||||
test('allowTurn — the tombstone is written once, not once per repeated denial', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const env = { CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1', TREKRESEARCH_MAX_CONV_TURNS: '1' };
|
||||
for (let i = 0; i < 8; i++) allowTurn({ runId: 'r-once', dimension: `d${i}`, effort: 'high' }, { env });
|
||||
for (let i = 0; i < 5; i++) allowTurn({ runId: 'r-once', dimension: 'd0', effort: 'high' }, { env });
|
||||
const after = readLedger(join(dir, 'trekresearch-loop-ledger.jsonl'), 'r-once');
|
||||
assert.equal(after.exhausted, 1, 'a hammered gate must not grow the ledger without bound');
|
||||
assert.equal(after.granted, 8);
|
||||
});
|
||||
});
|
||||
|
||||
test('readLedger — a tombstone is reported separately and never as a granted turn', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const p = join(dir, 'trekresearch-loop-ledger.jsonl');
|
||||
writeFileSync(p, [
|
||||
JSON.stringify({ runId: 'r1', dimension: 'd1', slot: 1 }),
|
||||
JSON.stringify({ runId: 'r1', exhausted: true }),
|
||||
JSON.stringify({ runId: 'other', exhausted: true }),
|
||||
].join('\n') + '\n');
|
||||
const l = readLedger(p, 'r1');
|
||||
assert.equal(l.granted, 1);
|
||||
assert.equal(l.exhausted, 1);
|
||||
assert.equal(readLedger(p, 'other').granted, 0, 'another run’s tombstone is not a granted turn either');
|
||||
});
|
||||
});
|
||||
|
||||
// ---- the discovery ceiling has a reader, not just a sentence ----------------
|
||||
//
|
||||
// The bounded-cost NFR asks for explicit ceilings on BOTH axes: max conversation
|
||||
// turns and max discovered dimensions. The turn axis got MAX_CONV_TURNS, a
|
||||
// ledger-backed reader and a PreToolUse enforcer. The discovery axis got a
|
||||
// sentence in Phase 4.5 — "append candidates only while the whole list stays at
|
||||
// or below maxDimensions: 8" — with no constant of its own, no reader, and no
|
||||
// test that a run exceeding it is caught. That is the same shape as the
|
||||
// brief_reviewer_iter_cap failure the operator decision warned about: a cap
|
||||
// nothing reads.
|
||||
//
|
||||
// The ceiling is deliberately the SAME constant that sizes the turn budget. Two
|
||||
// constants for one settings.json:16 value is how the two drift apart.
|
||||
|
||||
test('checkDimensionCeiling — a list at the ceiling is accepted', () => {
|
||||
const r = checkDimensionCeiling(Array.from({ length: MAX_TOTAL_DIMENSIONS }, (_, i) => `d${i}`));
|
||||
assert.equal(r.ok, true);
|
||||
assert.equal(r.count, MAX_TOTAL_DIMENSIONS);
|
||||
assert.equal(r.ceiling, MAX_TOTAL_DIMENSIONS);
|
||||
});
|
||||
|
||||
test('checkDimensionCeiling — one dimension over the ceiling is REJECTED', () => {
|
||||
const r = checkDimensionCeiling(Array.from({ length: MAX_TOTAL_DIMENSIONS + 1 }, (_, i) => `d${i}`));
|
||||
assert.equal(r.ok, false, 'a ceiling that accepts ceiling+1 is not a ceiling');
|
||||
assert.equal(r.reason, 'ceiling_exceeded');
|
||||
assert.equal(r.count, MAX_TOTAL_DIMENSIONS + 1);
|
||||
});
|
||||
|
||||
test('checkDimensionCeiling — a plain count works as well as a list', () => {
|
||||
assert.equal(checkDimensionCeiling(8).ok, true);
|
||||
assert.equal(checkDimensionCeiling(9).ok, false);
|
||||
assert.equal(checkDimensionCeiling('8').ok, true);
|
||||
});
|
||||
|
||||
test('checkDimensionCeiling — an unreadable count is rejected, never waved through', () => {
|
||||
for (const bad of ['abc', null, undefined, {}, -1, NaN]) {
|
||||
const r = checkDimensionCeiling(bad);
|
||||
assert.equal(r.ok, false, `${JSON.stringify(bad)} must not pass a cost ceiling`);
|
||||
assert.equal(r.reason, 'unreadable_dimension_count');
|
||||
}
|
||||
});
|
||||
|
||||
test('checkDimensionCeiling — the ceiling is the same constant that sizes the turn budget', () => {
|
||||
// Phase 4.5 and the Phase 5 budget must not be able to disagree about 8.
|
||||
assert.equal(checkDimensionCeiling(0).ceiling, MAX_TOTAL_DIMENSIONS);
|
||||
});
|
||||
|
||||
test('CLI shim — --check-dimensions exits 0 at the ceiling and 1 above it', () => {
|
||||
const at = runShim(['--check-dimensions', String(MAX_TOTAL_DIMENSIONS)], {});
|
||||
assert.equal(at.code, 0, `at the ceiling must exit 0; got ${at.out}`);
|
||||
assert.equal(JSON.parse(at.out.trim()).ok, true);
|
||||
|
||||
const over = runShim(['--check-dimensions', String(MAX_TOTAL_DIMENSIONS + 1)], {});
|
||||
assert.equal(over.code, 1, 'a run over the ceiling must be rejected by exit code, not by prose');
|
||||
const parsed = JSON.parse(over.out.trim());
|
||||
assert.equal(parsed.ok, false);
|
||||
assert.equal(parsed.reason, 'ceiling_exceeded');
|
||||
});
|
||||
|
||||
test('CLI shim — --check-dimensions needs no runId, effort or STORM flag', () => {
|
||||
// It is a cost ceiling on Phase 4.5, which never calls the budget gate, so it
|
||||
// must not inherit the budget gate's preconditions.
|
||||
const r = runShimStripped(['--check-dimensions', '3'], {});
|
||||
assert.equal(r.code, 0, `should not require --run-id/--effort; got ${r.out}`);
|
||||
});
|
||||
|
||||
// ---- (g) shim contract --------------------------------------------------------
|
||||
|
||||
test('CLI shim — grants and exits 0 when enabled + high effort + budget available', () => {
|
||||
withTmpDataDir((dir) => {
|
||||
const r = runShim(
|
||||
['--run-id', 'shim-1', '--dimension', 'd1', '--effort', 'high'],
|
||||
{ CLAUDE_PLUGIN_DATA: dir, VOYAGE_STORM_ENABLED: '1' },
|
||||
);
|
||||
assert.equal(r.code, 0);
|
||||
const parsed = JSON.parse(r.out.trim());
|
||||
assert.equal(parsed.ok, true);
|
||||
});
|
||||
});
|
||||
|
||||
test('CLI shim — denies and exits 1 when disabled', () => {
|
||||
const r = runShim(['--run-id', 'shim-2', '--dimension', 'd1', '--effort', 'high'], { VOYAGE_STORM_ENABLED: '0' });
|
||||
assert.equal(r.code, 1);
|
||||
const parsed = JSON.parse(r.out.trim());
|
||||
assert.equal(parsed.ok, false);
|
||||
assert.equal(parsed.reason, 'storm_disabled');
|
||||
});
|
||||
|
||||
test('CLI shim — grants with CLAUDE_PLUGIN_DATA STRIPPED from the environment', () => {
|
||||
withTmpDataDir((home) => {
|
||||
const r = runShimStripped(
|
||||
['--run-id', 'shim-stripped', '--dimension', 'd1', '--effort', 'high'],
|
||||
{ VOYAGE_STORM_ENABLED: '1', HOME: home },
|
||||
);
|
||||
assert.equal(r.code, 0, `shim must grant without CLAUDE_PLUGIN_DATA; got: ${r.out}`);
|
||||
const parsed = JSON.parse(r.out.trim());
|
||||
assert.equal(parsed.ok, true);
|
||||
assert.ok(existsSync(join(home, '.claude', 'voyage', 'trekresearch-loop-ledger.jsonl')));
|
||||
});
|
||||
});
|
||||
|
||||
test('CLI shim — missing required args exits 1 with usage reason', () => {
|
||||
const r = runShim(['--run-id', 'shim-3']);
|
||||
assert.equal(r.code, 1);
|
||||
const parsed = JSON.parse(r.out.trim());
|
||||
assert.equal(parsed.ok, false);
|
||||
assert.match(parsed.reason, /usage:/);
|
||||
});
|
||||
|
|
@ -15,21 +15,20 @@ import { censusTests } from '../../lib/util/test-census.mjs';
|
|||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
const TESTS_ROOT = join(HERE, '..');
|
||||
|
||||
test('suite census splits behavior / doc-pins / gold-eval (S19 + SKAL-1·4b)', (t) => {
|
||||
test('suite census splits behavior tests from doc-consistency pins (S19)', (t) => {
|
||||
const c = censusTests(TESTS_ROOT);
|
||||
// Honest-count invariant: the three buckets must account for every top-level
|
||||
// test() declaration — no silent drift between behavior, pin, and eval counts.
|
||||
assert.equal(c.behavior + c.docPins + c.goldEval, c.total,
|
||||
// Honest-count invariant: the two buckets must account for every top-level
|
||||
// test() declaration — no silent drift between behavior and pin counts.
|
||||
assert.equal(c.behavior + c.docPins, c.total,
|
||||
'census buckets must sum to the total declaration count');
|
||||
assert.ok(c.docPins > 0, 'doc-consistency pin bucket must be non-empty (regex/glob sanity)');
|
||||
assert.ok(c.behavior > 0, 'behavior bucket must be non-empty (regex/glob sanity)');
|
||||
assert.ok(c.goldEval > 0, 'gold-eval scoring-run bucket must be non-empty (SKAL-1·4b present)');
|
||||
// Report the split so the cited count is honest (audit §Top changes #8).
|
||||
// Metric = top-level test() declarations; node:test's runtime total counts
|
||||
// subtests too and is therefore ≥ this number.
|
||||
t.diagnostic(
|
||||
`behavior=${c.behavior} doc-consistency-pins=${c.docPins} ` +
|
||||
`gold-eval=${c.goldEval} total=${c.total} (top-level test() declarations)`,
|
||||
`total=${c.total} (top-level test() declarations)`,
|
||||
);
|
||||
});
|
||||
|
||||
|
|
|
|||
|
|
@ -75,15 +75,6 @@ test('deriveCost — hand-computed value for a known model (cache-aware)', () =>
|
|||
assert.equal(is_estimate, false);
|
||||
});
|
||||
|
||||
test('deriveCost — hand-computed value for claude-fable-5 (v5.9 fable tier)', () => {
|
||||
const totals = { tokens_input: 3000, tokens_output: 500, tokens_cache_creation: 1000, tokens_cache_read: 6000 };
|
||||
const { cost_usd, is_estimate } = deriveCost(totals, 'claude-fable-5');
|
||||
// (3000×10 + 500×50 + 1000×12.5 + 6000×1) / 1e6
|
||||
// = (30000 + 25000 + 12500 + 6000) / 1e6 = 73500 / 1e6 = 0.0735
|
||||
assert.equal(cost_usd, 0.0735);
|
||||
assert.equal(is_estimate, false);
|
||||
});
|
||||
|
||||
test('deriveCost — refuse-to-estimate for an unknown model', () => {
|
||||
const totals = { tokens_input: 3000, tokens_output: 500, tokens_cache_creation: 1000, tokens_cache_read: 6000 };
|
||||
const out = deriveCost(totals, 'claude-unknown-9');
|
||||
|
|
|
|||
|
|
@ -1,324 +0,0 @@
|
|||
// tests/scripts/storm-measure.test.mjs
|
||||
// Step 11 — the STORM adoption gate's deterministic accounting core.
|
||||
//
|
||||
// The gate decides ONE thing: does the bounded Phase 5 loop buy enough extra
|
||||
// source/coverage breadth to be worth flipping VOYAGE_STORM_ENABLED on by
|
||||
// default. Thresholds are pre-registered in docs/storm-measurement.md BEFORE
|
||||
// any measurement run, so this file pins the arithmetic that turns a
|
||||
// trekresearch-stats.jsonl into a verdict — not the verdict itself.
|
||||
//
|
||||
// Two properties carry the gate's honesty:
|
||||
// - runs with empty_turns > 0 are EXCLUDED from the gain and COUNTED, so
|
||||
// adoption is never decided on a broken denominator, and
|
||||
// - a stats file with no `effort` field is a loud error, never a silently
|
||||
// empty group that reads as "no gain".
|
||||
//
|
||||
// Pattern: tests/scripts/synthesis-measure.test.mjs (pure core, no fixtures on disk).
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { readFileSync } from 'node:fs';
|
||||
import {
|
||||
median,
|
||||
parseStats,
|
||||
partitionEligible,
|
||||
measure,
|
||||
decideVerdict,
|
||||
activationCheck,
|
||||
ADOPT_THRESHOLD,
|
||||
DECLINE_THRESHOLD,
|
||||
} from '../../scripts/storm-measure.mjs';
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// helpers — synthetic JSONL, one object per line, exactly as the orchestrator emits
|
||||
// ---------------------------------------------------------------------------
|
||||
function run({ effort, unique_sources, dimensions, dimensions_baseline, empty_turns = 0, conv_turns = 0 }) {
|
||||
return JSON.stringify({
|
||||
ts: '2026-08-12T00:00:00.000Z',
|
||||
question: 'q',
|
||||
mode: 'full',
|
||||
scope: 'both',
|
||||
engine: 'swarm',
|
||||
effort,
|
||||
unique_sources,
|
||||
dimensions,
|
||||
dimensions_baseline,
|
||||
conv_turns,
|
||||
empty_turns,
|
||||
});
|
||||
}
|
||||
|
||||
function jsonl(...lines) {
|
||||
return lines.join('\n') + '\n';
|
||||
}
|
||||
|
||||
// A control arm at 10 sources / 5 dimensions, and a treatment arm at 13
|
||||
// sources / 8 dimensions: +30.0% sources, +60.0% dimensions.
|
||||
const STANDARD = [
|
||||
run({ effort: 'standard', unique_sources: 10, dimensions: 5, dimensions_baseline: 5 }),
|
||||
run({ effort: 'standard', unique_sources: 10, dimensions: 5, dimensions_baseline: 5 }),
|
||||
run({ effort: 'standard', unique_sources: 10, dimensions: 5, dimensions_baseline: 5 }),
|
||||
];
|
||||
const HIGH = [
|
||||
run({ effort: 'high', unique_sources: 13, dimensions: 8, dimensions_baseline: 5, conv_turns: 3 }),
|
||||
run({ effort: 'high', unique_sources: 13, dimensions: 8, dimensions_baseline: 5, conv_turns: 3 }),
|
||||
run({ effort: 'high', unique_sources: 13, dimensions: 8, dimensions_baseline: 5, conv_turns: 3 }),
|
||||
];
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// median
|
||||
// ---------------------------------------------------------------------------
|
||||
test('median: odd, even, single', () => {
|
||||
assert.equal(median([3, 1, 2]), 2);
|
||||
assert.equal(median([1, 2, 3, 4]), 2.5);
|
||||
assert.equal(median([7]), 7);
|
||||
});
|
||||
|
||||
test('median: empty list is null, never 0 — 0 would read as a real measurement', () => {
|
||||
assert.equal(median([]), null);
|
||||
});
|
||||
|
||||
test('median does not mutate its input', () => {
|
||||
const xs = [3, 1, 2];
|
||||
median(xs);
|
||||
assert.deepEqual(xs, [3, 1, 2]);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// parseStats — the loud-error requirement
|
||||
// ---------------------------------------------------------------------------
|
||||
test('parseStats: a file with no effort field throws, it does not yield empty groups', () => {
|
||||
const noEffort = jsonl(
|
||||
JSON.stringify({ ts: 'x', unique_sources: 10, dimensions: 5, dimensions_baseline: 5 }),
|
||||
JSON.stringify({ ts: 'y', unique_sources: 11, dimensions: 5, dimensions_baseline: 5 }),
|
||||
);
|
||||
assert.throws(() => parseStats(noEffort), /effort/i);
|
||||
});
|
||||
|
||||
test('parseStats: skips blank and malformed lines but keeps the good ones', () => {
|
||||
const text = jsonl(STANDARD[0], '', 'not json', HIGH[0]);
|
||||
const { records, malformed } = parseStats(text);
|
||||
assert.equal(records.length, 2);
|
||||
assert.equal(malformed, 1);
|
||||
});
|
||||
|
||||
test('parseStats: an empty file throws rather than reporting a zero-gain verdict', () => {
|
||||
assert.throws(() => parseStats(''), /no records/i);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// exclusion of broken runs
|
||||
// ---------------------------------------------------------------------------
|
||||
test('partitionEligible: runs with empty_turns > 0 are excluded and counted', () => {
|
||||
const { records } = parseStats(jsonl(
|
||||
...HIGH,
|
||||
run({ effort: 'high', unique_sources: 99, dimensions: 8, dimensions_baseline: 5, empty_turns: 2 }),
|
||||
));
|
||||
const { eligible, excluded } = partitionEligible(records);
|
||||
assert.equal(eligible.length, 3);
|
||||
assert.equal(excluded, 1);
|
||||
});
|
||||
|
||||
test('partitionEligible: empty_turns === 0 is eligible; a missing field counts as 0', () => {
|
||||
const { records } = parseStats(jsonl(
|
||||
STANDARD[0],
|
||||
JSON.stringify({ ts: 'z', effort: 'standard', unique_sources: 10, dimensions: 5, dimensions_baseline: 5 }),
|
||||
));
|
||||
const { eligible, excluded } = partitionEligible(records);
|
||||
assert.equal(eligible.length, 2);
|
||||
assert.equal(excluded, 0);
|
||||
});
|
||||
|
||||
// A malformed empty_turns used to land in the ELIGIBLE arm: Number('many') is
|
||||
// NaN, NaN fails the isFinite test, and the `else` branch pushed it in. The
|
||||
// exclusion is one of the two properties carrying this gate's honesty, so a
|
||||
// garbage value silently re-entering the denominator defeats it — and it does so
|
||||
// in the direction that flatters adoption, since the run that broke is the run
|
||||
// whose numbers are least trustworthy.
|
||||
test('partitionEligible: a non-numeric empty_turns is EXCLUDED, never silently eligible', () => {
|
||||
const { records } = parseStats(jsonl(
|
||||
STANDARD[0],
|
||||
JSON.stringify({
|
||||
ts: 'm', effort: 'high', unique_sources: 999,
|
||||
dimensions: 8, dimensions_baseline: 5, empty_turns: 'many',
|
||||
}),
|
||||
));
|
||||
const { eligible, excluded } = partitionEligible(records);
|
||||
assert.equal(excluded, 1, 'a value that cannot be read as a turn count is not evidence of zero empty turns');
|
||||
assert.equal(eligible.length, 1);
|
||||
});
|
||||
|
||||
test('partitionEligible: an unparsable empty_turns cannot move the median either', () => {
|
||||
const m = measure(parseStats(jsonl(
|
||||
...STANDARD,
|
||||
...HIGH,
|
||||
JSON.stringify({
|
||||
ts: 'm', effort: 'high', unique_sources: 900,
|
||||
dimensions: 8, dimensions_baseline: 5, empty_turns: {},
|
||||
}),
|
||||
)).records);
|
||||
assert.equal(m.excluded, 1);
|
||||
assert.equal(m.sources.treatment, 13, 'the 900-source malformed run must not reach the median');
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// measure — the known-answer test
|
||||
// ---------------------------------------------------------------------------
|
||||
test('measure: median gains on synthetic runs give the known answer', () => {
|
||||
const m = measure(parseStats(jsonl(...STANDARD, ...HIGH)).records);
|
||||
assert.equal(m.control.n, 3);
|
||||
assert.equal(m.treatment.n, 3);
|
||||
assert.equal(m.sources.control, 10);
|
||||
assert.equal(m.sources.treatment, 13);
|
||||
assert.ok(Math.abs(m.sources.gain - 0.30) < 1e-9, `sources gain ${m.sources.gain}`);
|
||||
// (8 - 5) / 5 = 0.60 within each treatment run.
|
||||
assert.ok(Math.abs(m.dimensions.gain - 0.60) < 1e-9, `dimensions gain ${m.dimensions.gain}`);
|
||||
});
|
||||
|
||||
test('measure: an excluded run cannot move the median', () => {
|
||||
const withBroken = jsonl(
|
||||
...STANDARD,
|
||||
...HIGH,
|
||||
run({ effort: 'high', unique_sources: 900, dimensions: 8, dimensions_baseline: 5, empty_turns: 1 }),
|
||||
);
|
||||
const m = measure(parseStats(withBroken).records);
|
||||
assert.equal(m.excluded, 1);
|
||||
assert.equal(m.sources.treatment, 13, 'the 900-source broken run must not reach the median');
|
||||
assert.ok(Math.abs(m.sources.gain - 0.30) < 1e-9);
|
||||
});
|
||||
|
||||
test('measure: reports null gain (not 0) when an arm has no eligible runs', () => {
|
||||
const m = measure(parseStats(jsonl(...HIGH)).records);
|
||||
assert.equal(m.control.n, 0);
|
||||
assert.equal(m.sources.gain, null);
|
||||
assert.equal(m.verdict, 'insufficient-data');
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// verdict mapping — both sides of both thresholds
|
||||
// ---------------------------------------------------------------------------
|
||||
test('decideVerdict: at and above the adopt threshold', () => {
|
||||
assert.equal(decideVerdict(ADOPT_THRESHOLD, ADOPT_THRESHOLD), 'adopt');
|
||||
assert.equal(decideVerdict(0.55, 0.44), 'adopt');
|
||||
});
|
||||
|
||||
test('decideVerdict: both metrics between the bars is inconclusive, not adopt', () => {
|
||||
assert.equal(decideVerdict(ADOPT_THRESHOLD - 0.0001, 0.2), 'inconclusive');
|
||||
});
|
||||
|
||||
test('decideVerdict: below the decline threshold on both metrics declines', () => {
|
||||
assert.equal(decideVerdict(0.14, 0.05), 'decline');
|
||||
assert.equal(decideVerdict(DECLINE_THRESHOLD - 0.0001, 0), 'decline');
|
||||
});
|
||||
|
||||
test('decideVerdict: at the decline threshold is inconclusive, not decline', () => {
|
||||
assert.equal(decideVerdict(DECLINE_THRESHOLD, DECLINE_THRESHOLD), 'inconclusive');
|
||||
});
|
||||
|
||||
// The brief pre-registers "median forbedring >= 30 % på (a) eller (b) → adopt.
|
||||
// < 15 % → decline." — OR on both sides, with adopt evaluated first.
|
||||
test('decideVerdict: adopt needs EITHER metric — one strong metric carries a weak one', () => {
|
||||
assert.equal(decideVerdict(0.90, 0.10), 'adopt');
|
||||
assert.equal(decideVerdict(0.10, 0.90), 'adopt');
|
||||
});
|
||||
|
||||
test('decideVerdict: either metric below the decline bar declines', () => {
|
||||
assert.equal(decideVerdict(0.02, 0.20), 'decline');
|
||||
assert.equal(decideVerdict(0.20, 0.02), 'decline');
|
||||
});
|
||||
|
||||
test('decideVerdict: adopt outranks decline when one metric clears and the other is under the decline bar', () => {
|
||||
assert.equal(decideVerdict(0.90, 0.10), 'adopt');
|
||||
});
|
||||
|
||||
test('decideVerdict: a null gain is insufficient data, never a decline', () => {
|
||||
assert.equal(decideVerdict(null, 0.4), 'insufficient-data');
|
||||
assert.equal(decideVerdict(0.4, null), 'insufficient-data');
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// activationCheck — BOTH halves of the SC, not just the count
|
||||
//
|
||||
// The SC requires that an `effort: high` run discovers at least one dimension
|
||||
// AND that the dimension list in the output brief is a TRUE SUPERSET of the
|
||||
// interview-derived ones. activationCheck only computed dimensions -
|
||||
// dimensions_baseline >= 1, so a run that replaced two interview dimensions with
|
||||
// three discovered ones passed the check while violating the SC's second half.
|
||||
// Supersetness was asserted only by Phase 4.5's prose contract that discovery
|
||||
// appends; nothing read it.
|
||||
//
|
||||
// The stats record carries counts, not names — names are free prose and
|
||||
// field-allowlist.mjs denies prose by omission — so the run attests membership
|
||||
// with a low-cardinality boolean instead, and the gate refuses to call
|
||||
// activation OK without it.
|
||||
// ---------------------------------------------------------------------------
|
||||
function highRun(over = {}) {
|
||||
return JSON.stringify({
|
||||
ts: '2026-08-12T00:00:00.000Z',
|
||||
effort: 'high',
|
||||
unique_sources: 13,
|
||||
dimensions: 8,
|
||||
dimensions_baseline: 5,
|
||||
conv_turns: 3,
|
||||
empty_turns: 0,
|
||||
dimensions_baseline_preserved: true,
|
||||
...over,
|
||||
});
|
||||
}
|
||||
|
||||
test('activationCheck: discovery plus a preserved baseline is activation', () => {
|
||||
const r = activationCheck(parseStats(jsonl(highRun())).records);
|
||||
assert.equal(r.ok, true);
|
||||
assert.equal(r.discovered_dimensions, 3);
|
||||
assert.equal(r.dimensions_baseline_preserved, true);
|
||||
});
|
||||
|
||||
test('activationCheck: a REPLACED baseline is not activation, however many were discovered', () => {
|
||||
const r = activationCheck(parseStats(jsonl(
|
||||
highRun({ dimensions: 8, dimensions_baseline: 5, dimensions_baseline_preserved: false }),
|
||||
)).records);
|
||||
assert.equal(r.ok, false, 'a count delta of +3 says nothing about which dimensions survived');
|
||||
assert.match(r.reason, /superset|baseline/i);
|
||||
});
|
||||
|
||||
test('activationCheck: a run that does not attest baseline membership cannot pass', () => {
|
||||
const rec = JSON.parse(highRun());
|
||||
delete rec.dimensions_baseline_preserved;
|
||||
const r = activationCheck(parseStats(jsonl(JSON.stringify(rec))).records);
|
||||
assert.equal(r.ok, false, 'an absent attestation is not an attestation');
|
||||
assert.match(r.reason, /dimensions_baseline_preserved/);
|
||||
});
|
||||
|
||||
test('activationCheck: no discovery is still not activation even with the baseline preserved', () => {
|
||||
const r = activationCheck(parseStats(jsonl(
|
||||
highRun({ dimensions: 5, dimensions_baseline: 5 }),
|
||||
)).records);
|
||||
assert.equal(r.ok, false);
|
||||
assert.equal(r.discovered_dimensions, 0);
|
||||
});
|
||||
|
||||
test('activationCheck: no effort:high run at all is reported as such', () => {
|
||||
const r = activationCheck(parseStats(jsonl(...STANDARD)).records);
|
||||
assert.equal(r.ok, false);
|
||||
assert.match(r.reason, /high/);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// pre-registration — the thresholds are the doc's, and the doc is committed first
|
||||
// ---------------------------------------------------------------------------
|
||||
test('thresholds are the pre-registered 30% / 15%', () => {
|
||||
assert.equal(ADOPT_THRESHOLD, 0.30);
|
||||
assert.equal(DECLINE_THRESHOLD, 0.15);
|
||||
});
|
||||
|
||||
// The human-readable summary is the only form of the rule most readers will
|
||||
// ever see. It said "on BOTH" on both sides while decideVerdict evaluated OR —
|
||||
// so the report described a stricter gate than the one that produced the verdict
|
||||
// printed one line below it.
|
||||
test('the printed threshold line states the OR rule that decideVerdict actually applies', () => {
|
||||
const src = readFileSync(new URL('../../scripts/storm-measure.mjs', import.meta.url), 'utf-8');
|
||||
const line = src.split('\n').find((l) => l.includes('thresholds: adopt'));
|
||||
assert.ok(line, 'the summary must still print its threshold rule');
|
||||
assert.doesNotMatch(line, /on BOTH/, 'the rule is OR on both sides — printing BOTH misstates the gate');
|
||||
assert.match(line, /EITHER/);
|
||||
});
|
||||
|
|
@ -1,9 +1,5 @@
|
|||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { execFileSync } from 'node:child_process';
|
||||
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import { validateBriefContent } from '../../lib/validators/brief-validator.mjs';
|
||||
|
||||
const GOOD_BRIEF = `---
|
||||
|
|
@ -255,25 +251,6 @@ test('validateBrief — v5.1.1: UNQUOTED brief_version 2.1 WITH phase_signals is
|
|||
assert.ok(!r.errors.find(e => e.code === 'BRIEF_V51_MISSING_SIGNALS'));
|
||||
});
|
||||
|
||||
// --- v5.9 — fable model tier (BASE_ALLOWED_MODELS widened to three values) ---
|
||||
|
||||
test('validateBrief — v5.9: fable phase_signals fixture accepted (no BRIEF_INVALID_MODEL)', () => {
|
||||
const t = readFileSync(new URL('../fixtures/brief-effort-fable.md', import.meta.url), 'utf-8');
|
||||
const r = validateBriefContent(t, { strict: true });
|
||||
assert.equal(r.valid, true, JSON.stringify(r.errors));
|
||||
assert.ok(!r.errors.find(e => e.code === 'BRIEF_INVALID_MODEL'));
|
||||
});
|
||||
|
||||
test('validateBrief — v5.9: unknown model gpt5 in phase_signals rejected with BRIEF_INVALID_MODEL', () => {
|
||||
const t = GOOD_BRIEF
|
||||
.replace('brief_version: "2.0"', 'brief_version: "2.1"')
|
||||
.replace('source: interview\n', `source: interview\n${SIGNALS_BLOCK.replace('model: opus', 'model: gpt5')}`);
|
||||
const r = validateBriefContent(t, { strict: true });
|
||||
assert.equal(r.valid, false);
|
||||
assert.ok(r.errors.find(e => e.code === 'BRIEF_INVALID_MODEL'),
|
||||
`expected BRIEF_INVALID_MODEL for gpt5, got: ${JSON.stringify(r.errors)}`);
|
||||
});
|
||||
|
||||
// --- v5.5 — framing enforcement + obligatory TL;DR (gated at brief_version ≥ 2.2) ---
|
||||
// Operator decision (S6, option A1): framing + TL;DR are hard BLOCKERs for briefs
|
||||
// declaring brief_version ≥ 2.2; existing 2.0/2.1 briefs stay valid (forward-compat,
|
||||
|
|
@ -416,38 +393,3 @@ test('validateBrief — S18 min-version: trekreview brief is exempt (no framing
|
|||
const r = validateBriefContent(REVIEW_AS_BRIEF, { minBriefVersion: '2.2' });
|
||||
assert.ok(!r.warnings.find(w => w.code === 'BRIEF_VERSION_BELOW_MINIMUM'));
|
||||
});
|
||||
|
||||
// S56 — CLI arg-parsing regression. The no-flag invocation `brief-validator.mjs <brief.md>`
|
||||
// used to bail to Usage/exit 2 because the --min-version skip index was 0 when the flag
|
||||
// was absent, dropping the file positional (which sits at argv index 0).
|
||||
test('CLI — no-flag invocation reaches validation, does not bail to Usage (S56 regression)', () => {
|
||||
const dir = mkdtempSync(join(tmpdir(), 'brief-cli-'));
|
||||
try {
|
||||
const file = join(dir, 'brief.md');
|
||||
writeFileSync(file, GOOD_BRIEF);
|
||||
// execFileSync throws on non-zero exit; pre-fix this bailed to Usage (exit 2).
|
||||
const out = execFileSync(process.execPath, [
|
||||
'lib/validators/brief-validator.mjs',
|
||||
file,
|
||||
], { encoding: 'utf-8' });
|
||||
assert.match(out, /PASS/);
|
||||
} finally {
|
||||
rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
test('CLI — --min-version still locates the file positional after the value token (S56)', () => {
|
||||
const dir = mkdtempSync(join(tmpdir(), 'brief-cli-'));
|
||||
try {
|
||||
const file = join(dir, 'brief.md');
|
||||
writeFileSync(file, GOOD_BRIEF);
|
||||
const out = execFileSync(process.execPath, [
|
||||
'lib/validators/brief-validator.mjs',
|
||||
'--min-version', '2.0',
|
||||
file,
|
||||
], { encoding: 'utf-8' });
|
||||
assert.match(out, /PASS/);
|
||||
} finally {
|
||||
rmSync(dir, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
|
|
|||
|
|
@ -1,5 +1,5 @@
|
|||
// tests/validators/profile-validator.test.mjs
|
||||
// SC #1, #2, #3: profile-validator validates lib/profiles/{economy,balanced,premium,fable}.yaml
|
||||
// SC #1, #2, #3: profile-validator validates lib/profiles/{economy,balanced,premium}.yaml
|
||||
// (innebygde profiler) plus rejects invalid models and invalid enum types.
|
||||
|
||||
import { test } from 'node:test';
|
||||
|
|
@ -12,15 +12,14 @@ import {
|
|||
validateProfileContent,
|
||||
PROFILE_REQUIRED_FIELDS,
|
||||
PROFILE_REQUIRED_PHASES,
|
||||
BASE_ALLOWED_MODELS,
|
||||
} from '../../lib/validators/profile-validator.mjs';
|
||||
|
||||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||
const REPO_ROOT = join(__dirname, '..', '..');
|
||||
|
||||
// SC #1: alle 4 innebygde profiler grønne
|
||||
// SC #1: alle 3 innebygde profiler grønne
|
||||
|
||||
for (const profileName of ['economy', 'balanced', 'premium', 'fable']) {
|
||||
for (const profileName of ['economy', 'balanced', 'premium']) {
|
||||
test(`SC #1: lib/profiles/${profileName}.yaml validates clean`, () => {
|
||||
const r = validateProfile(join(REPO_ROOT, 'lib', 'profiles', `${profileName}.yaml`));
|
||||
assert.equal(r.valid, true,
|
||||
|
|
@ -94,77 +93,6 @@ brief_reviewer_iter_cap: 1
|
|||
`expected valid with VOYAGE_ALLOW_HAIKU=1, got: ${JSON.stringify(allowed.errors)}`);
|
||||
});
|
||||
|
||||
// Fable tier (v5.9): fable accepted under DEFAULT env (no opt-in flag),
|
||||
// unknown models still rejected — the allowlist gate must demonstrably fire.
|
||||
|
||||
test('fable accepted in phase_models under default env (no env flag)', () => {
|
||||
const fableProfile = `---
|
||||
profile_version: "1.0"
|
||||
name: fable-inline
|
||||
phase_models:
|
||||
- phase: brief
|
||||
model: fable
|
||||
- phase: research
|
||||
model: fable
|
||||
- phase: plan
|
||||
model: fable
|
||||
- phase: execute
|
||||
model: fable
|
||||
- phase: review
|
||||
model: fable
|
||||
- phase: continue
|
||||
model: fable
|
||||
parallel_agents_min: 6
|
||||
parallel_agents_max: 8
|
||||
external_research_enabled: true
|
||||
brief_reviewer_iter_cap: 3
|
||||
---
|
||||
`;
|
||||
const r = validateProfileContent(fableProfile, { env: { /* default: no flags */ } });
|
||||
assert.equal(r.valid, true,
|
||||
`expected fable accepted under default env, got: ${JSON.stringify(r.errors)}`);
|
||||
assert.equal(r.errors.length, 0);
|
||||
});
|
||||
|
||||
test('unknown model gpt5 rejected with PROFILE_INVALID_MODEL under default env', () => {
|
||||
const gpt5Profile = `---
|
||||
profile_version: "1.0"
|
||||
name: gpt5-inline
|
||||
phase_models:
|
||||
- phase: brief
|
||||
model: gpt5
|
||||
- phase: research
|
||||
model: sonnet
|
||||
- phase: plan
|
||||
model: opus
|
||||
- phase: execute
|
||||
model: sonnet
|
||||
- phase: review
|
||||
model: opus
|
||||
- phase: continue
|
||||
model: sonnet
|
||||
parallel_agents_min: 2
|
||||
parallel_agents_max: 4
|
||||
external_research_enabled: false
|
||||
brief_reviewer_iter_cap: 1
|
||||
---
|
||||
`;
|
||||
const r = validateProfileContent(gpt5Profile, { env: { /* default: no flags */ } });
|
||||
assert.equal(r.valid, false);
|
||||
const found = r.errors.find(e => e.code === 'PROFILE_INVALID_MODEL' && /gpt5/.test(e.message));
|
||||
assert.ok(found, `expected PROFILE_INVALID_MODEL for gpt5, got: ${JSON.stringify(r.errors)}`);
|
||||
});
|
||||
|
||||
// BASE_ALLOWED_MODELS allowlist drift-pin (mirrors the PROFILE_REQUIRED_FIELDS pin)
|
||||
|
||||
test('BASE_ALLOWED_MODELS export drift-pin', () => {
|
||||
assert.deepEqual(
|
||||
[...BASE_ALLOWED_MODELS],
|
||||
['sonnet', 'opus', 'fable'],
|
||||
'BASE_ALLOWED_MODELS contract drift — pin contract',
|
||||
);
|
||||
});
|
||||
|
||||
// Required fields presence
|
||||
|
||||
test('PROFILE_MISSING_FIELD when name absent', () => {
|
||||
|
|
|
|||
|
|
@ -1,174 +0,0 @@
|
|||
// tests/validators/query-privacy-gate.test.mjs
|
||||
// Cover lib/validators/query-privacy-gate.mjs: two-sided code table
|
||||
// (absolute path / repo-internal identifier / secret-shaped token), a
|
||||
// benign query passing untouched, the opt-in env var reaching only the
|
||||
// warn tier (never the hard-block tier), strict/soft severity, and the
|
||||
// CLI shim.
|
||||
//
|
||||
// Secret-shaped fixtures are built via string concatenation/repeat, never
|
||||
// as literal tokens — the repo's own secrets pre-edit hook (correctly)
|
||||
// treats a literal AKIA/sk-/ghp_ string as a real credential.
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { execFileSync } from 'node:child_process';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
import {
|
||||
validateOutboundQuery,
|
||||
ABSOLUTE_PATH_PATTERNS,
|
||||
REPO_IDENTIFIER_PATTERNS,
|
||||
SECRET_SHAPED_PATTERNS,
|
||||
} from '../../lib/validators/query-privacy-gate.mjs';
|
||||
|
||||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
const SHIM = join(HERE, '..', '..', 'lib', 'validators', 'query-privacy-gate.mjs');
|
||||
|
||||
const FAKE_OPENAI_KEY = 'sk-' + 'a'.repeat(24);
|
||||
const FAKE_AWS_KEY = 'AKIA' + 'Q'.repeat(16);
|
||||
const FAKE_GITHUB_PAT = 'ghp_' + 'b'.repeat(36);
|
||||
|
||||
// Real-world formats whose token body contains hyphens/underscores. A run of
|
||||
// plain alphanumerics is broken by those separators, so a naive
|
||||
// `[A-Za-z0-9]{20,}` run-length pattern lets them through — the gap this file
|
||||
// pins. Anthropic Console keys are `sk-ant-api03-` + ~95 base64url chars;
|
||||
// GitHub fine-grained PATs are `github_pat_<22>_<59>`; `gho_` is the OAuth
|
||||
// sibling of the classic `ghp_` token.
|
||||
const FAKE_ANTHROPIC_KEY = 'sk-' + 'ant-' + 'api03-' + 'A1b2_-x9'.repeat(12);
|
||||
const FAKE_GITHUB_FINE_GRAINED = 'github' + '_pat_' + 'A'.repeat(22) + '_' + 'c'.repeat(59);
|
||||
const FAKE_GITHUB_OAUTH = 'gho' + '_' + 'd'.repeat(36);
|
||||
|
||||
function runShim(args) {
|
||||
try {
|
||||
const out = execFileSync(process.execPath, [SHIM, ...args], {
|
||||
encoding: 'utf-8',
|
||||
stdio: ['ignore', 'pipe', 'pipe'],
|
||||
});
|
||||
return { code: 0, out };
|
||||
} catch (e) {
|
||||
return { code: e.status ?? 1, out: e.stdout?.toString() ?? '' };
|
||||
}
|
||||
}
|
||||
|
||||
// ---- two-sided code table ----------------------------------------------------
|
||||
|
||||
const TABLE = [
|
||||
{ label: 'absolute path (/Users/...)', text: 'find every caller of foo in /Users/ktg/repos/voyage/lib/util/foo.mjs', code: 'PRIVACY_ABSOLUTE_PATH' },
|
||||
{ label: 'absolute path (/home/...)', text: 'trace /home/alice/projects/app/src/index.js for imports', code: 'PRIVACY_ABSOLUTE_PATH' },
|
||||
{ label: 'repo-internal identifier (forgejo host)', text: 'what changed recently on git.fromaitochitta.com/open/voyage', code: 'PRIVACY_REPO_IDENTIFIER' },
|
||||
{ label: 'repo-internal identifier (repo name)', text: 'search issues for ktg-plugin-marketplace regressions', code: 'PRIVACY_REPO_IDENTIFIER' },
|
||||
{ label: 'secret-shaped (OpenAI/Anthropic-style key)', text: `auth failing with key ${FAKE_OPENAI_KEY}`, code: 'PRIVACY_SECRET_SHAPED' },
|
||||
{ label: 'secret-shaped (AWS access key)', text: `rotate ${FAKE_AWS_KEY} now`, code: 'PRIVACY_SECRET_SHAPED' },
|
||||
{ label: 'secret-shaped (GitHub PAT)', text: `token leaked: ${FAKE_GITHUB_PAT}`, code: 'PRIVACY_SECRET_SHAPED' },
|
||||
{ label: 'secret-shaped (Anthropic Console key)', text: `why does ${FAKE_ANTHROPIC_KEY} 401`, code: 'PRIVACY_SECRET_SHAPED' },
|
||||
{ label: 'secret-shaped (GitHub fine-grained PAT)', text: `pushed with ${FAKE_GITHUB_FINE_GRAINED}`, code: 'PRIVACY_SECRET_SHAPED' },
|
||||
{ label: 'secret-shaped (GitHub OAuth token)', text: `oauth flow returned ${FAKE_GITHUB_OAUTH}`, code: 'PRIVACY_SECRET_SHAPED' },
|
||||
];
|
||||
|
||||
for (const { label, text, code } of TABLE) {
|
||||
test(`validateOutboundQuery — ${label} → ${code} (strict, error)`, () => {
|
||||
const r = validateOutboundQuery(text, { strict: true, env: {} });
|
||||
assert.equal(r.valid, false);
|
||||
assert.ok(r.errors.find(e => e.code === code), JSON.stringify(r.errors));
|
||||
});
|
||||
}
|
||||
|
||||
test('validateOutboundQuery — benign generic query passes untouched', () => {
|
||||
const r = validateOutboundQuery('What are the tradeoffs between optimistic and pessimistic locking?', { env: {} });
|
||||
assert.equal(r.valid, true);
|
||||
assert.deepEqual(r.errors, []);
|
||||
assert.deepEqual(r.warnings, []);
|
||||
});
|
||||
|
||||
// ---- strict vs soft (warn tier only) -----------------------------------------
|
||||
|
||||
test('validateOutboundQuery — soft mode downgrades warn-tier findings to warnings, stays valid', () => {
|
||||
const r = validateOutboundQuery('inspect /Users/ktg/repos/voyage', { strict: false, env: {} });
|
||||
assert.equal(r.valid, true);
|
||||
assert.equal(r.errors.length, 0);
|
||||
assert.ok(r.warnings.find(w => w.code === 'PRIVACY_ABSOLUTE_PATH'));
|
||||
});
|
||||
|
||||
test('validateOutboundQuery — soft mode does NOT downgrade the hard-block tier', () => {
|
||||
const r = validateOutboundQuery(`leaked ${FAKE_OPENAI_KEY}`, { strict: false, env: {} });
|
||||
assert.equal(r.valid, false);
|
||||
assert.ok(r.errors.find(e => e.code === 'PRIVACY_SECRET_SHAPED'));
|
||||
});
|
||||
|
||||
// ---- opt-in env var reaches only the warn tier -------------------------------
|
||||
|
||||
test('validateOutboundQuery — VOYAGE_QUERY_PRIVACY_ALLOW=1 bypasses the warn tier entirely', () => {
|
||||
const r = validateOutboundQuery('inspect /Users/ktg/repos/voyage', { env: { VOYAGE_QUERY_PRIVACY_ALLOW: '1' } });
|
||||
assert.equal(r.valid, true);
|
||||
assert.equal(r.errors.length, 0);
|
||||
assert.equal(r.warnings.length, 0);
|
||||
});
|
||||
|
||||
test('validateOutboundQuery — VOYAGE_QUERY_PRIVACY_ALLOW=1 does NOT open the hard-block tier', () => {
|
||||
const r = validateOutboundQuery(`leaked ${FAKE_OPENAI_KEY}`, { env: { VOYAGE_QUERY_PRIVACY_ALLOW: '1' } });
|
||||
assert.equal(r.valid, false);
|
||||
assert.ok(r.errors.find(e => e.code === 'PRIVACY_SECRET_SHAPED'), 'opt-in must never unlock the hard-block tier');
|
||||
});
|
||||
|
||||
test('validateOutboundQuery — VOYAGE_QUERY_PRIVACY_ALLOW=1 combined with a secret still denies', () => {
|
||||
const r = validateOutboundQuery(`/Users/ktg/x leaked ${FAKE_OPENAI_KEY}`, { env: { VOYAGE_QUERY_PRIVACY_ALLOW: '1' } });
|
||||
assert.equal(r.valid, false);
|
||||
assert.equal(r.errors.length, 1);
|
||||
assert.equal(r.errors[0].code, 'PRIVACY_SECRET_SHAPED');
|
||||
});
|
||||
|
||||
// The hard-block tier is the one thing no operator flag unlocks, so a format
|
||||
// it misses is a secret leaving the machine with no second gate behind it.
|
||||
// Pin every real-world format against BOTH escape hatches at once.
|
||||
for (const [label, token] of [
|
||||
['Anthropic Console key', FAKE_ANTHROPIC_KEY],
|
||||
['GitHub fine-grained PAT', FAKE_GITHUB_FINE_GRAINED],
|
||||
['GitHub OAuth token', FAKE_GITHUB_OAUTH],
|
||||
]) {
|
||||
test(`validateOutboundQuery — ${label} stays blocked under --soft and the opt-in`, () => {
|
||||
const r = validateOutboundQuery(`leaked ${token}`, {
|
||||
strict: false,
|
||||
env: { VOYAGE_QUERY_PRIVACY_ALLOW: '1' },
|
||||
});
|
||||
assert.equal(r.valid, false, 'hard-block tier must never be overridable');
|
||||
assert.ok(r.errors.find(e => e.code === 'PRIVACY_SECRET_SHAPED'));
|
||||
});
|
||||
}
|
||||
|
||||
// ---- empty input --------------------------------------------------------------
|
||||
|
||||
test('validateOutboundQuery — empty string is invalid', () => {
|
||||
const r = validateOutboundQuery('', { env: {} });
|
||||
assert.equal(r.valid, false);
|
||||
assert.ok(r.errors.find(e => e.code === 'PRIVACY_EMPTY_QUERY'));
|
||||
});
|
||||
|
||||
// ---- pattern set is frozen ----------------------------------------------------
|
||||
|
||||
test('pattern sets are Object.frozen', () => {
|
||||
assert.equal(Object.isFrozen(ABSOLUTE_PATH_PATTERNS), true);
|
||||
assert.equal(Object.isFrozen(REPO_IDENTIFIER_PATTERNS), true);
|
||||
assert.equal(Object.isFrozen(SECRET_SHAPED_PATTERNS), true);
|
||||
});
|
||||
|
||||
// ---- CLI shim -----------------------------------------------------------------
|
||||
|
||||
test('CLI shim — benign query exits 0 with valid:true', () => {
|
||||
const r = runShim(['harmless generic question about caching strategies']);
|
||||
assert.equal(r.code, 0);
|
||||
const parsed = JSON.parse(r.out.trim());
|
||||
assert.equal(parsed.valid, true);
|
||||
});
|
||||
|
||||
test('CLI shim — secret-shaped query exits 1 even with --soft', () => {
|
||||
const r = runShim(['--soft', `leaked ${FAKE_OPENAI_KEY}`]);
|
||||
assert.equal(r.code, 1);
|
||||
const parsed = JSON.parse(r.out.trim());
|
||||
assert.equal(parsed.valid, false);
|
||||
assert.ok(parsed.errors.find(e => e.code === 'PRIVACY_SECRET_SHAPED'));
|
||||
});
|
||||
|
||||
test('CLI shim — missing query argument exits 2 (usage error)', () => {
|
||||
const r = runShim([]);
|
||||
assert.equal(r.code, 2);
|
||||
});
|
||||
|
|
@ -26,9 +26,6 @@ fail() { echo "[FAIL] SC$1 — $2"; FAIL=$((FAIL + 1)); exit 1; }
|
|||
|
||||
# Tracked-file exclusions (paths preserved verbatim from old name).
|
||||
# - CHANGELOG/TRADEMARKS/MIGRATION legitimately reference the old name.
|
||||
# - cc-upgrade-2.1.181-decision-matrix.md cites CC features: `ultracode`
|
||||
# (a real CC keyword, 2.1.160) and the `ultra-cc-architect` plugin name —
|
||||
# both legitimate, not rebrand leftovers.
|
||||
# - architecture-discovery.mjs + project-discovery.mjs are Q8 exceptions
|
||||
# pointing at the upstream architect producer slot.
|
||||
# - verify.sh self-references the forbidden patterns to detect them.
|
||||
|
|
@ -36,7 +33,6 @@ fail() { echo "[FAIL] SC$1 — $2"; FAIL=$((FAIL + 1)); exit 1; }
|
|||
exclude_path() {
|
||||
case "$1" in
|
||||
*CHANGELOG.md|*TRADEMARKS.md|*MIGRATION.md) return 0 ;;
|
||||
*cc-upgrade-2.1.181-decision-matrix.md) return 0 ;;
|
||||
*lib/validators/architecture-discovery.mjs) return 0 ;;
|
||||
*lib/parsers/project-discovery.mjs) return 0 ;;
|
||||
*verify.sh) return 0 ;;
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue