repo-mailbox/scripts/route.sh
Kjell Tore Guttormsen dd8f3ce042 feat(route)!: Fable rows are a hand-written override, never a route.sh outcome
The operator removed the "Fable can only be suggested after a failed Opus
5/xhigh session" policy on 2026-08-06 - it stood in the way too often. Until
now that policy was implemented as code: --opus-xhigh-failed was the ONLY
way route.sh's SELECTION chain could reach rows 5-6 (Fable), so the rubric
still enforced a policy the operator had already dropped.

Given a choice between (A) deleting the branches and the flag outright, so
the calculator's output range closes at row 4, or (C) keeping a path that
renders a Fable row on explicit instruction with rule=operator-override, the
operator chose A (AskUserQuestion) - CLAUDE.md is explicit that the rubric
stays the only deterministic lookup and a Fable choice is now always a
deviation from it, recorded in STATE as an override rather than produced
here. board.sh is untouched and still parses "Fable 5/high"/"Fable 5/xhigh"
written by hand into a board line (route-selftest.sh section 6 now pins that
half directly, since route.sh can no longer produce the strings itself).

TDD: every affected selftest check was rewritten to fail against the
unmodified route.sh first (72->68/69, confirmed red), then route.sh was
edited to match. --opus-xhigh-failed is gone outright - passing it now exits
2 like any other unknown argument, not silently accepted as a no-op.

route-selftest.sh: 73 -> 69 checks (three checks tested command shapes
row_advisor() can no longer produce; the two row-5/row-6 reachability checks
in section 1 collapsed into one "the flag is gone" check). coord 191, board
142 unaffected. Suite total 406 -> 402.

Also folds in a standalone fix already pushed this session: section 14 was
gating the wrong CLI fact (whether "claude --advisor fable" is rejected,
which row_advisor() never depends on) instead of the one it actually rests
on (whether opus/sonnet can advise a FABLE main model). Re-pointed and
verified against the installed CC 2.1.226.

Verified before committing: grepped every repo under ~/repos for a route
line carrying --opus-xhigh-failed (none - one repo has it in prose only,
not in its <!-- route: --> comment) and ran board.sh --plan/--brief over the
real tree to confirm no repo's command line broke. Sent a follow-up
coord-send to catalog superseding an earlier now-stale "406" stat-line
correction with the current 402.

skills/route/SKILL.md: usage block, "last-session record" framing, and the
closing --opus-xhigh-failed paragraph rewritten to match. CLAUDE.md, README
and CHANGELOG updated (checks 73->69, badge 406->402, new 0.21.0 entry).
Version bumped 0.20.3 -> 0.21.0 across plugin.json, package.json and all
three skill frontmatters (breaking CLI removal at 0.x -> minor, per v0.20.0
precedent).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ap1WKHCDPcSfpjo4ds2dQX
2026-08-09 22:03:05 +02:00

329 lines
16 KiB
Bash
Executable file

#!/bin/bash
# route.sh - score four traits of the NEXT task, get the model and effort to
# run it with. A pure calculator: reads nothing, writes nothing, prints one
# block of key=value lines on stdout. The session pastes the result into
# STATE.md; this script never touches a file.
#
# WHY THIS EXISTS. The board line's next-cost field had a reader (board.sh)
# and no writer, so its value was typed by hand every session and drifted into
# several competing spellings. Cleaning the data could not fix that.
# A writer with a CLOSED output range can: this script can only ever emit one
# of four strings, so a fifth cannot enter circulation.
#
# THE ROW TABLE IS THE POLICY, and it is the operator's rubric verbatim -
# moved here so there is one copy rather than one per repo. It has six rows;
# this calculator only ever computes four of them:
#
# 1 Sonnet 5/high reading, summarizing, docs, mechanical refactor
# 2 Sonnet 5/xhigh TDD cycle, known-root-cause bugfix, one-file change
# 3 Opus 5/high multi-file feature, architecture choice, hard debugging
# 4 Opus 5/xhigh long autonomous run, big refactor, cross-repo migration
# 5 Fable 5/high deliberate choice for big-picture/review/planning work
# 6 Fable 5/xhigh same, open-ended or longest horizon
#
# ROWS 5-6 ARE NEVER COMPUTED HERE. Until 2026-08-06 they fired only from an
# explicit --opus-xhigh-failed flag; that policy is REMOVED (operator
# decision), and nothing replaces it as a rubric outcome. A Fable choice is
# now always a deliberate deviation from this calculator - CLAUDE.md is
# explicit that the rubric stays the only deterministic lookup and a
# departure from it is recorded in STATE as an OVERRIDE, never produced here.
# Write "Fable 5/high" or "Fable 5/xhigh" into the board line by hand when
# that is the right call; board.sh still parses both (route-selftest.sh
# section 6 pins that half). The one fact worth carrying: a Fable session
# runs without an advisor (row_advisor() below, and gated against the
# installed claude by route-selftest.sh section 14) - informational, never a
# gate on reaching the row, since there is no longer a gate to reach.
#
# Cheapest first, so the rubric's "always name one row cheaper as the quota
# fallback" is row minus one, floored at row 1, correct by construction.
#
# THE TRAITS describe the task, never how it feels. "Hard", "complex" and
# "important" are deliberately absent: they are unfalsifiable and collapse to
# a hunch, which is the thing being replaced.
#
# --path known | partial | undetermined
# Is the solution route described, or must it be found?
# --verification strong | weak | none
# Will tests, types or a compiler catch the error?
# --reversibility cheap | costly | one-way
# --scope local | multi-file | cross-cutting
# --rationale required free text - WHY these four scores
#
# All five are required. None has a default, and that is load-bearing:
# verification carries the most signal and is the trait most often left out,
# and a default would be indistinguishable from a real score when the log is
# read back to find out whether the ROUTING was wrong or the SCORING was.
#
# SELECTION - first match wins, most expensive first. Rows 5-6 do not appear:
# they are never a trait-derived outcome (see above).
# row 4 reversibility=one-way OR scope=cross-cutting
# row 3 path=partial|undetermined OR reversibility=costly OR scope=multi-file
# row 2 verification=weak|none
# row 1 otherwise
#
# Escalation is ASYMMETRIC on purpose: any single trait escalates, while
# reaching row 1 needs all four at the cheap end. Underkill costs one session,
# overkill costs quota every session - but a wrong architecture decision in a
# published plugin costs more than either.
#
# THE ADVISOR is emitted into the command as '--advisor opus' - a second,
# stronger model consulted at key moments during the session. It is added on a
# NEED, never unconditionally: an always-on advisor is the global advisorModel
# setting, which burns quota on every session in every repo and is the thing
# this rule exists to replace. Two independent needs qualify:
#
# rows 1-2 ALWAYS. The main model is Sonnet, so opus is a capability LIFT
# rather than a peer - opus judgement at sonnet cost. This is what
# makes the FALLBACK safe to take: every fallback is one row
# cheaper, and the cheapest rows are the Sonnet ones.
# rows 3-4 only at reversibility=costly|one-way. The main model is already
# Opus, so the advisor buys peer review, worth paying for when a
# mistake is not cheap to undo.
#
# Rows 5-6 never reach this logic at all - the calculator cannot select them
# (see above). Informational only: were the operator to hand-write a Fable
# command, it would carry no advisor either way, since the CLI rejects every
# advisor for a Fable main model.
#
# The two triggers barely overlap: costly forces row 3 and one-way forces row
# 4, so a Sonnet row always has reversibility=cheap. verification=none is
# deliberately not a third trigger - beyond the stakes rule it would only add
# mistakes that are cheap to reverse, docs sessions among them.
#
# Applied per ROW, so 'fallback-command' carries its own correct answer rather
# than the winning row's.
#
# WHERE IT DISAGREES WITH THE RUBRIC'S EXAMPLES. The rows are task-type labels;
# the traits are a different classification over the same six outcomes. They
# part company by one row on the two cheapest rows - documentation scores
# known/none/cheap/local and lands on row 2 where the rubric's prose says row
# 1; a TDD cycle scores known/strong/cheap/local and lands on row 1 where the
# prose says row 2. This is left alone deliberately. Chasing the example
# phrases would mean re-implementing description-matching, which is the guess
# the traits exist to replace. It is a policy, not a theorem: --rationale is
# how a misscore is found afterwards.
#
# THE SPECIFICATION CHECK THAT COMES FREE. path=undetermined with no planned
# design phase means the TASK DESCRIPTION is underspecified - not that the
# model should be upgraded. Rewrite the next step; upgrading the model to
# compensate for a vague spec is the most expensive form of procrastination
# available.
#
# Usage:
# route.sh --path <v> --verification <v> --reversibility <v> --scope <v>
# --rationale <text>
#
# route.sh ... --last-model <Sonnet 5|Opus 5|Fable 5>
# --last-effort <low|medium|high|xhigh|max>
# --last-completed <yes|no> --last-corrections <n>
#
# THE LAST-SESSION RECORD (the four --last-* fields, all or none) is what makes
# any of this falsifiable. It records how the session that just ran actually
# went, so the policy can later be judged against outcomes instead of against
# how sensible it reads. --last-corrections is the cheap proxy: systematically
# high counts on row 1 mean the cheap row is too easy to reach, systematically
# zero on row 4 means escalation fires too readily.
#
# All four fields are closed sets or numbers, and required together, so the
# record reads back later as evidence rather than a guess - a partial record
# would emit an empty value indistinguishable from a real measurement. Every
# one of them must be MEASURED by the caller: --last-effort comes from CLAUDE_EFFORT,
# which Claude Code exports into every tool-use context as the session's current
# effort level. It is deliberately NOT defaulted from that variable here - a
# calculator that reads its own environment stops being deterministic from its
# arguments, and section 6 of the selftest depends on that determinism. Reading
# the value off the previous board line would measure what was PRESCRIBED rather
# than what was RUN; so, less obviously, does asking the operator, who reads it
# off the startup command they typed. Omit the record rather than guess - a
# guessed value reads back as a measurement.
#
# It is pure telemetry and never changes what the calculator outputs.
# Inferring "the model failed" from "the session did not finish" would fire on
# context exhaustion and on operator interrupts, which say nothing about the
# model - so the record stays descriptive, never a trigger.
#
# Exit 0 on a decision, 2 on any bad or missing argument. ASCII only,
# bash 3.2 safe.
set -u
export LC_ALL=C
PATH_T=""; VERIF=""; REVERS=""; SCOPE=""; RATIONALE=""; RAT_SET=0
L_MODEL=""; L_EFFORT=""; L_DONE=""; L_CORR=""; L_SET=0
die() { echo "route: $1" >&2; exit 2; }
need() { [ $# -ge 2 ] || die "$1 requires a value"; }
while [ $# -gt 0 ]; do
case "$1" in
# bash 3.2: `shift 2` past the end of $# is a no-op -> would loop forever.
--path) need "$@"; PATH_T="$2"; shift 2 ;;
--verification) need "$@"; VERIF="$2"; shift 2 ;;
--reversibility) need "$@"; REVERS="$2"; shift 2 ;;
--scope) need "$@"; SCOPE="$2"; shift 2 ;;
--rationale) need "$@"; RATIONALE="$2"; RAT_SET=1; shift 2 ;;
--last-model) need "$@"; L_MODEL="$2"; L_SET=1; shift 2 ;;
--last-effort) need "$@"; L_EFFORT="$2"; L_SET=1; shift 2 ;;
--last-completed) need "$@"; L_DONE="$2"; L_SET=1; shift 2 ;;
--last-corrections) need "$@"; L_CORR="$2"; L_SET=1; shift 2 ;;
-h|--help) grep '^#' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;;
*) die "unknown argument: $1" ;;
esac
done
# Unknown values are rejected rather than tolerated. A silently accepted
# typo would route on three traits and look exactly like a scored decision.
case "$PATH_T" in
known|partial|undetermined) ;;
"") die "--path is required (known|partial|undetermined)" ;;
*) die "--path: unknown value '$PATH_T' (known|partial|undetermined)" ;;
esac
case "$VERIF" in
strong|weak|none) ;;
"") die "--verification is required (strong|weak|none)" ;;
*) die "--verification: unknown value '$VERIF' (strong|weak|none)" ;;
esac
case "$REVERS" in
cheap|costly|one-way) ;;
"") die "--reversibility is required (cheap|costly|one-way)" ;;
*) die "--reversibility: unknown value '$REVERS' (cheap|costly|one-way)" ;;
esac
case "$SCOPE" in
local|multi-file|cross-cutting) ;;
"") die "--scope is required (local|multi-file|cross-cutting)" ;;
*) die "--scope: unknown value '$SCOPE' (local|multi-file|cross-cutting)" ;;
esac
[ "$RAT_SET" -eq 1 ] || die "--rationale is required (why these four scores)"
[ -n "$RATIONALE" ] || die "--rationale must not be empty"
# The last-session record is all four fields or none at all. A partial record
# emits an empty value that reads back later exactly like a measured one, and
# the whole purpose of the record is to be readable evidence months from now.
if [ "$L_SET" -eq 1 ]; then
[ -n "$L_MODEL" ] || die "--last-model is required with a last-session record"
[ -n "$L_EFFORT" ] || die "--last-effort is required with a last-session record"
[ -n "$L_DONE" ] || die "--last-completed is required with a last-session record"
[ -n "$L_CORR" ] || die "--last-corrections is required with a last-session record"
# Model and effort are closed sets here, not free text. The next session READS
# this record back as evidence months from now, so a drifted spelling
# ("opus 5" for "Opus 5") rebuilds the reader-versus-writer drift this whole
# script exists to remove, one field over.
case "$L_MODEL" in
"Sonnet 5"|"Opus 5"|"Fable 5") ;;
*) die "--last-model: '$L_MODEL' is not a row-table model (Sonnet 5|Opus 5|Fable 5)" ;;
esac
case "$L_EFFORT" in
low|medium|high|xhigh|max) ;;
*) die "--last-effort: '$L_EFFORT' is not a verified effort level (low|medium|high|xhigh|max)" ;;
esac
case "$L_DONE" in
yes|no) ;;
*) die "--last-completed: unknown value '$L_DONE' (yes|no)" ;;
esac
case "$L_CORR" in
""|*[!0-9]*) die "--last-corrections must be a whole number, got '$L_CORR'" ;;
esac
fi
# --- Selection: first match wins, most expensive first ---------------------
# Rows 5-6 (Fable) never appear: they are a hand-written operator override,
# never a trait-derived outcome (see the header note).
if [ "$REVERS" = "one-way" ]; then
ROW=4; RULE="reversibility=one-way"
elif [ "$SCOPE" = "cross-cutting" ]; then
ROW=4; RULE="scope=cross-cutting"
elif [ "$PATH_T" != "known" ]; then
ROW=3; RULE="path=$PATH_T"
elif [ "$REVERS" = "costly" ]; then
ROW=3; RULE="reversibility=costly"
elif [ "$SCOPE" = "multi-file" ]; then
ROW=3; RULE="scope=multi-file"
elif [ "$VERIF" != "strong" ]; then
ROW=2; RULE="verification=$VERIF"
else
ROW=1; RULE="no escalating trait (all four at the cheap end)"
fi
# --- The row table: one decision, two spellings ---------------------------
# The rubric name goes in the board line so repos compare by eye; the CLI alias
# goes in the command the operator pastes. Emitting both from one table is the
# point - two hand-maintained spellings of one decision is how they disagree.
# Aliases are gated against the installed claude by route-selftest.sh section
# 11, never assumed here. Rows 5-6 have no entry: $ROW can never be 5 or 6
# (see SELECTION above), so a case arm for them would be dead code.
row_name() {
case "$1" in
1) echo "Sonnet 5/high" ;; 2) echo "Sonnet 5/xhigh" ;;
3) echo "Opus 5/high" ;; 4) echo "Opus 5/xhigh" ;;
esac
}
row_base_cmd() {
case "$1" in
1) echo "claude --model sonnet --effort high" ;;
2) echo "claude --model sonnet --effort xhigh" ;;
3) echo "claude --model opus --effort high" ;;
4) echo "claude --model opus --effort xhigh" ;;
esac
}
# THE ADVISOR is a second, stronger model consulted mid-task. It costs real
# tokens per session, so it fires on a NEED and nowhere else - an unconditional
# advisor is just the global advisorModel setting, which is the thing this
# replaces. Two independent needs qualify, and they are almost disjoint:
#
# rows 1-2 (Sonnet) ALWAYS. opus is a capability LIFT here, not a peer:
# opus judgement at sonnet cost. This half is what makes
# the fallback-command safe, since every fallback is one
# row cheaper and the cheapest rows are Sonnet.
# rows 3-4 (Opus) only at costly|one-way stakes, where the advisor is a
# peer review and being wrong is not cheap to undo.
#
# Rows 5-6 (Fable) never reach this function - $ROW can only be 1-4 (see
# SELECTION above). Informational only: the only advisor this script ever
# emits is opus (pinned by selftest 14's "the only advisor value ever emitted
# is opus"), and opus is refused as under-capable for a fable main model -
# measured against the installed claude, still true at CC 2.1.226 - so a
# hand-written Fable command carries no advisor either way.
#
# costly forces row 3 and one-way forces row 4, so a Sonnet row always has
# reversibility=cheap - the stakes rule can never reach rows 1-2, and the model
# rule never reaches rows 3-4. verification=none is deliberately NOT a trigger:
# beyond the stakes rule it would only add cheap-to-reverse mistakes, and it
# would put an advisor on every docs session (known/none/cheap/local).
#
# Applied per ROW rather than once, because the fallback is a real command the
# operator pastes under quota pressure and must carry its own correct answer.
row_advisor() {
case "$1" in
1|2) echo " --advisor opus" ;;
3|4) case "$2" in costly|one-way) echo " --advisor opus" ;; *) echo "" ;; esac ;;
*) echo "" ;;
esac
}
row_cmd() { printf '%s%s\n' "$(row_base_cmd "$1")" "$(row_advisor "$1" "$REVERS")"; }
FB=$((ROW - 1)); [ "$FB" -ge 1 ] || FB=1
# The rationale is free text from a session and lands inside an HTML comment on
# ONE line. A newline would split the line and hand the remainder to board.sh's
# NESTE extractor as the repo's next step; a literal '-->' would close the
# comment early and spill the rest into the rendered STATE.md. Same
# line-oriented sanitizing the send side applies to its own fields.
RAT_CLEAN="$(printf '%s' "$RATIONALE" | tr '\r\n' ' ' | tr -d '\000-\037' \
| sed 's/-->/-- >/g')"
# rationale goes LAST in the line: it is the only field that may contain a
# ';', so anything after it would be unparseable.
echo "row=$ROW"
echo "rule=$RULE"
echo "next-cost=$(row_name "$ROW")"
echo "command=$(row_cmd "$ROW")"
echo "fallback=$(row_name "$FB")"
echo "fallback-command=$(row_cmd "$FB")"
echo "route-line=<!-- route: path=$PATH_T; verification=$VERIF; reversibility=$REVERS; scope=$SCOPE; rationale=$RAT_CLEAN -->"
# No sanitizing needed on the record: all four fields are validated against
# closed sets above, so none of them can carry a newline or a '-->'.
if [ "$L_SET" -eq 1 ]; then
echo "route-last=<!-- route-last: model=$L_MODEL; effort=$L_EFFORT; completed=$L_DONE; corrections=$L_CORR -->"
fi
exit 0