voyage/agents
Kjell Tore Guttormsen 106dcb0091
fix(verification): a backticked span is only run when it IS a command
"The first backtick span is the command" is right for a plan, whose template
puts the command first, and wrong for a brief, whose criterion usually opens
by NAMING the thing under discussion. Measured 2026-09-18 on the repo's own
example brief: 5 of 6 criteria FAILED, 3 of them parse artifacts - `--verbose`
run as a command gave exit 2 ("invalid option"), `tests/` gave exit 126 ("is a
directory"). The rubric reads a FAILED result as decisive, so each one became
a BROKEN_SUCCESS_CRITERION BLOCKER about prose.

looksLikeCommand() screens the span by SHAPE only - no filesystem lookup, so a
span parses the same everywhere. Refused: a leading flag, a directory, a token
carrying quotes/braces/prose, and a lone relative path with a slash (an
explicit ./, ../, / or ~/ still runs, as do env-var prefixes). A refused span
is `unrunnable` with reason `not-a-command` - its own outcome, never FAILED,
and it never reaches a shell.

It deliberately does NOT scan on to a later span. "The first span that LOOKS
like a command" invents commands out of prose: in that same example brief it
would have run `whoami` and `login`, two real binaries a sentence happens to
name. An absent measurement is honest; a guessed one is not.

The shape check applies to prose spans only. Inside a shell-tagged fence the
author has already declared shell, so `[ -f x ] || exit 1` still runs.

The rubric follows: a NOT RUN result is never on its own a finding. The
Partial row now describes half-built DELIVERED CODE, and the reviewer gets a
table of the three reason strings - no-command, placeholder, not-a-command -
with what each says about the sentence rather than about the code.

Not covered, stated for the record: a multi-token span whose first token is a
non-executable file (`tests/golden/login.stdout --check`) still runs, and a
criterion whose command is real but whose binary is absent still reports the
shell's exit 127 - that is a true measurement of a missing binary, not a
parse artifact.

Red first: 4 runner tests + 1 doc-consistency pin failed before this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 02:04:46 +02:00
..
architecture-mapper.md refactor(agents): relocate example blocks to body (retrieval agents) 2026-06-29 10:16:48 +02:00
brief-conformance-reviewer.md fix(verification): a backticked span is only run when it IS a command 2026-09-18 02:04:46 +02:00
brief-reviewer.md refactor(agents): relocate example blocks to body (reviewer/planning agents) 2026-06-29 10:19:24 +02:00
code-correctness-reviewer.md chore(voyage): pin all sub-agents to Opus permanently (operator request) 2026-05-13 20:20:08 +02:00
community-researcher.md refactor(agents): relocate example blocks to body (researchers + gemini-bridge) 2026-06-29 10:14:35 +02:00
contrarian-researcher.md refactor(agents): relocate example blocks to body (researchers + gemini-bridge) 2026-06-29 10:14:35 +02:00
convention-scanner.md refactor(agents): relocate example blocks to body (retrieval agents) 2026-06-29 10:16:48 +02:00
dependency-tracer.md refactor(agents): relocate example blocks to body (retrieval agents) 2026-06-29 10:16:48 +02:00
docs-researcher.md refactor(agents): relocate example blocks to body (researchers + gemini-bridge) 2026-06-29 10:14:35 +02:00
git-historian.md refactor(agents): relocate example blocks to body (retrieval agents) 2026-06-29 10:16:48 +02:00
plan-critic.md refactor(agents): relocate example blocks to body (reviewer/planning agents) 2026-06-29 10:19:24 +02:00
planning-orchestrator.md fix(trekplan): drop TaskCreate/TaskUpdate from the tool lists - measured dead weight 2026-09-17 15:34:50 +02:00
research-orchestrator.md release(v5.10.1): drop gemini-bridge from the pipeline; correct the T1 §6 PoC status 2026-09-03 20:29:39 +02:00
research-scout.md refactor(agents): relocate example blocks to body (reviewer/planning agents) 2026-06-29 10:19:24 +02:00
review-coordinator.md fix(review): an anonymous invalid payload is unattributable, not a reviewer named "unnamed reviewer" 2026-09-01 22:54:31 +02:00
review-orchestrator.md fix(trekplan): drop TaskCreate/TaskUpdate from the tool lists - measured dead weight 2026-09-17 15:34:50 +02:00
risk-assessor.md refactor(agents): relocate example blocks to body (reviewer/planning agents) 2026-06-29 10:19:24 +02:00
scope-guardian.md refactor(agents): relocate example blocks to body (reviewer/planning agents) 2026-06-29 10:19:24 +02:00
security-researcher.md refactor(agents): relocate example blocks to body (researchers + gemini-bridge) 2026-06-29 10:14:35 +02:00
session-decomposer.md refactor(agents): relocate example blocks to body (reviewer/planning agents) 2026-06-29 10:19:24 +02:00
synthesis-agent.md chore(voyage): release v5.6.1 — one-line descriptions for reference/dormant agents (~700 tok trim) 2026-06-24 11:52:59 +02:00
task-finder.md refactor(agents): relocate example blocks to body (retrieval agents) 2026-06-29 10:16:48 +02:00
test-strategist.md refactor(agents): relocate example blocks to body (retrieval agents) 2026-06-29 10:16:48 +02:00