Claude Code 2.1.233 stopped offering the task/todo tools on Opus 4.8, Sonnet 5, Fable 5 and newer unless CLAUDE_CODE_ENABLE_TODO_TOOLS=1 is set. /trekplan still listed TaskCreate and TaskUpdate in allowed-tools. So did the two orchestrator reference docs, which mirror the command but are never spawned. No prose in any command calls either tool. Runtime effect, measured BEFORE removal on CC 2.1.274, claude-opus-5, with CLAUDE_CODE_ENABLE_TODO_TOOLS unset: - Probe: a throwaway project command whose frontmatter is `allowed-tools: Read, TaskCreate, TaskUpdate`, run headless (`claude -p --output-format stream-json --verbose`). - The session had 151 tools. The only Task* names among them were Task, TaskOutput and TaskStop; neither TaskCreate nor TaskUpdate was there. - The command loaded (listed in slash_commands) and completed with subtype=success, is_error=false and 0 tool calls. The model reported the tool as not available. - `claude plugin validate` does not flag the unknown names either: only the one accepted warning appears, before and after. Verdict: harmless dead weight, no error anywhere, so this is a correctness-of-docs fix rather than a runtime fix. End-state gate D-05: open -> closed (defects 5 -> 4 of 7). Suite 1059 (1057/0/2), run on a clean export of the index. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
11 KiB
| name | description | model | color | tools | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| review-orchestrator | Reference document, not a spawnable capability — documents the /trekreview workflow that runs inline in main context (full rationale, CC-2.1.172 history, and phase map in body). | opus | red |
|
This document is the canonical workflow description for the trekreview
pipeline as of v3.2.0. The /trekreview command reads it as
reference and executes the phases below inline in the main command
context. It was moved out of background sub-agent mode because, before
Claude Code 2.1.172, that mode would silently lose the Agent tool and degrade
the reviewer swarm to single-context reasoning. As of CC 2.1.172 sub-agents
can spawn sub-agents (up to 5 levels deep), so this constraint no longer
holds; a delegated-orchestration redesign is under evaluation (see
docs/cc-upgrade-2.1.181-decision-matrix.md, W1/CC-26). Until then,
orchestration stays inline.
The role of the "orchestrator" now belongs to the command markdown itself: the main Opus session launches reviewer agents via the Agent tool, runs the coordinator, validates the output, and writes review.md to disk.
Design principle: independent, then bounded
Each reviewer runs independently — no cross-feeding of findings between brief-conformance-reviewer and code-correctness-reviewer. The coordinator then applies BOUNDED operations only: deduplication, severity ranking, reasonableness filter. Synthesis-level inference across files is explicitly forbidden in v1.0 (Judge Agent pattern).
Input
You will receive a prompt containing:
- Project dir — path to the trekplan project folder (the brief and
optional
progress.jsonlive here; the review will be written to{project_dir}/review.md). - Brief path —
{project_dir}/brief.md. Read it; the brief is the contract that bounds review scope. - Mode —
default,quick,validate, ordry-run.default— run the full pipeline.quick— skip the coordinator's reasonableness filter; use single reviewer (code-correctness only) for faster turnaround.validate— schema-only check on existing review.md, no LLM calls.dry-run— print the discovered scope and triage map; skip writes.
- Since-ref (optional) — explicit
--since <ref>override for the SHA range. Validated viagit rev-parse --verify <ref>. - Plugin root — for template access.
Read the brief file first. It is the contract. Parse its frontmatter and every section (Intent, Goal, Non-Goals, Constraints, Success Criteria, Open Questions, Prior Attempts).
Your workflow
Execute these phases in order. Do not skip phases.
Phase 1 — Parse mode and validate input
Run the arg-parser via Bash:
node ${CLAUDE_PLUGIN_ROOT}/lib/parsers/arg-parser.mjs --command trekreview "$@"
Pull the structured flags from its JSON output. Reject unknown flags. If
--project is missing and a brief argument was not supplied directly,
print usage and stop.
Phase 2 — Validate brief
Run the brief validator in soft mode (the brief was produced earlier in the pipeline — we accept partial grades, we just want a parseable contract):
node ${CLAUDE_PLUGIN_ROOT}/lib/validators/brief-validator.mjs --soft --json {brief_path}
If valid: false with REQUIRED-field errors: stop, ask the user to
re-run /trekbrief first.
Phase 3 — Discover scope SHA range
Determine the range of commits this review covers.
- Preferred path — read
{project_dir}/progress.jsonif it exists. Extractsession_start_sha. This is the "before" SHA. - Fallback — if no
progress.json, use the brief's mtime to find the most recent commit AT OR BEFORE the brief was written. Emit a clear warning in the review's Executive Summary noting the fallback. - Override —
--since <ref>overrides the discovered "before" SHA. Validate the ref withgit rev-parse --verify <ref>. Reject if invalid. - The "after" SHA is
git rev-parse HEAD.
Compute the diff:
git diff --name-only {before_sha}..{after_sha}
Add working-tree changes (uncommitted) with the [uncommitted] annotation
the brief contract specifies. The Coverage table marks them explicitly.
Phase 4 — Triage gate (path-pattern classifier)
The triage gate is deterministic — no LLM judgment. It runs a hardcoded path-pattern classifier over the file list from Phase 3 and produces a treatment map:
| Treatment | When |
|---|---|
skip |
Matches *.lock, *.svg, dist/**, build/**, node_modules/**, generated-file marker present in first 3 lines |
deep-review |
Matches auth/**, crypto/**, **/security/**, hooks/** |
summary-only |
Default treatment for everything else |
Hard refuse-with-suggestion gates (use AskUserQuestion):
-
100 files in the diff
-
100,000 tokens of estimated diff content (
git diffoutput size / 4)
If gated, suggest narrowing the scope with --since <closer-ref> or
splitting the review across multiple commits.
Record the treatment for every file. Files marked skip MUST appear in
the Coverage section of review.md — never silently drop them. A silent
drop is a COVERAGE_SILENT_SKIP finding emitted by the coordinator.
Phase 5 — Launch parallel reviewers
Launch two reviewer agents in parallel using the Agent tool — one message, multiple tool calls.
Reviewers run independently. Do NOT pre-feed findings between them. The coordinator handles cross-cutting decisions later.
| Agent | Purpose |
|---|---|
brief-conformance-reviewer |
Trace each Success Criterion + Non-Goal to delivered code. Flag UNIMPLEMENTED_CRITERION, NON_GOAL_VIOLATED, BROKEN_SUCCESS_CRITERION, MISSING_BRIEF_REF, SCOPE_CREEP_BUILT, PLAN_EXECUTE_DRIFT. |
code-correctness-reviewer |
7-dimension code review. Flag MISSING_ERROR_HANDLING, PLAN_EXECUTE_DRIFT, MISSING_TEST, PLACEHOLDER_IN_CODE, SECURITY_INJECTION, UNDECLARED_DEPENDENCY. |
Each reviewer receives:
- Diff context — the unified diff from Phase 3 (truncated per file
for files marked
summary-only). - Triage map — full file list with treatments. Reviewers must respect
skipdecisions — if they want to flag a skipped file they emit a COVERAGE_SILENT_SKIP finding instead. - Brief path — for re-reading; do not inline the full brief into the prompt to keep token budgets honest.
In quick mode, launch only code-correctness-reviewer. Skip the
brief-conformance pass; the coverage matrix will still appear in
review.md but it is structural, not behavioral.
Phase 6 — Coordinator dedup + verdict
Launch review-coordinator with the merged findings array from Phase 5.
The coordinator runs a 4-pass process:
- Dedup by
(file, line, rule_key)triplet — keep highest severity. - HubSpot Judge filters — drop findings failing Succinctness, Accuracy, or Actionability.
- Cloudflare reasonableness — drop speculative findings without a
file:linecitation; drop findings whoserule_keyis not inRULE_CATALOGUE. - Compute verdict —
BLOCKifBLOCKER ≥ 1,WARNifMAJOR ≥ 1, elseALLOW.
The coordinator's output is the full review.md content — frontmatter + body sections + trailing JSON block — ready to write.
In quick mode, skip pass 3 (reasonableness filter). Passes 1, 2, 4
still run.
Phase 7 — Write review.md
Use the destination from Phase 1:
- With
--project: write to{project_dir}/review.md.
Create parent directories if needed. The frontmatter findings: field
must use block-style YAML (one ID per line with - prefix). The
parser at lib/util/frontmatter.mjs does not accept flow-style arrays.
The trailing JSON block in the body must be a valid json fenced code
block, last fenced block in the file, parseable by JSON.parse().
Phase 8 — Validate + stats
Run the review validator in strict mode:
node ${CLAUDE_PLUGIN_ROOT}/lib/validators/review-validator.mjs --json {project_dir}/review.md
If validation fails, repair the file (most failures are fixable in place — missing required frontmatter field, missing body section, malformed finding-ID). Do NOT proceed if any REVIEW_REQUIRED_FRONTMATTER field is missing.
Append a stats line to ${CLAUDE_PLUGIN_DATA}/trekreview-stats.jsonl:
{"ts":"...","slug":"...","verdict":"BLOCK|WARN|ALLOW","counts":{"BLOCKER":N,"MAJOR":N,"MINOR":N,"SUGGESTION":N},"reviewed_files_count":N,"mode":"default|quick|validate|dry-run","duration_ms":N}
Hard rules
- Currently runs inline, not in background. This orchestrator file is
reference, not a runnable sub-agent. Original reason: before Claude Code
2.1.172, background mode silently lost the Agent tool and the reviewer
swarm collapsed to single-context reasoning. As of CC 2.1.172 sub-agents
can spawn sub-agents (up to 5 levels deep), so that hard block is gone — a
delegated redesign is under evaluation (see
docs/cc-upgrade-2.1.181-decision-matrix.md, W1/CC-26). Until it lands, always run review agents from the main /trekreview command context. - Reviewers run independently. No cross-feeding of findings. The coordinator is the only place where reviewer outputs are combined.
- Coordinator scope is bounded. Dedup, severity ranking, reasonableness filter only. No cross-file inference. No synthesis-level hallucination. Synthesis is a v1.1 candidate — for v1.0 it is forbidden.
- Brief is the contract. Every finding must have a
brief_reftracing back to a brief section (SC, Non-Goal, Constraint, NFR). Findings withoutbrief_refare MISSING_BRIEF_REF (MAJOR). - No silent drops. Every file in the discovered diff must appear in
the Coverage section, even if its treatment is
skip. Hidden truncation is COVERAGE_SILENT_SKIP (MAJOR). - Cost: Sub-agents use their pinned
model:frontmatter (currentlyopusfor all). Aphase_signals[<phase>].modelbrief signal or the active--profile(e.g.economy) overrides the model per-phase. The orchestrator (the /trekreview command itself) runs on Opus. - Privacy: Never log secrets, tokens, or credentials. Findings citing
files with secret-like content must redact the secret in the
detail. - Honesty: If the diff is trivially small or all-skip, say so. Do not pad findings to make the review look thorough.
- Block-style YAML for findings list. The frontmatter parser does not
support flow-style arrays.
findings: [a, b]is broken; use:findings: - <id1> - <id2>