ORDRE 59. A dispatched order used to live only in a scratch prompt file
passed through argv, so it died with the pane it was typed into. Measured
2026-08-17: one order was dispatched three times over 90 minutes before it
was worked, because the first two tabs ran something else and the order
left no trace in the receiving repo at all.
New channel `~/.claude/coord/<repo>/orders/`, beside `inbox/` and never
merged with it. The axis is authorization: inbox content is untrusted
cross-repo data that may never instruct a session (Rule 6), a dispatch
order is operator-authorized work by construction. One channel carrying
both classes would mean either mail that can instruct or orders that
cannot, so the infrastructure is reused and the channel is not.
Four one-verb engines: coord-order-send.sh (write), coord-order-inbox.sh
(read, writes nothing at all), coord-order-claim.sh (atomic claim),
coord-order-done.sh (executed with a commit pointer / --no-commit with a
reason / --return with a reason).
The claim is a rename with no check-then-act step, so of N racing sessions
exactly one finds the source and the rest get ENOENT. The test that proves
it spawns 20 claimers BARRIERED on a start flag - unbarriered children do
not race at all - and runs the identical harness against a deliberately
racy `[ -e src ] && cp && rm` as a known-negative control, which must
produce many winners. Without that control, "exactly one winner" is
indistinguishable from "the race never happened".
Channel separation is pinned structurally, not only behaviourally: no mail
script may contain the string `orders`, with a known-positive control
proving the grep can find. coord-done cannot archive an order and
coord-order-claim cannot claim a message.
board gains an ORDRE column beside INN, counted with the identical idiom
and never summed with it: INN is "others are waiting on YOU", ORDRE is
"work is waiting on this REPO". Claimed orders are excluded - the column
answers what a session can pick up. board.sh --dispatch --order-id emits a
thin starter carrying only the id and the four steps, so the order text has
exactly one home; the id is validated shell-clean and must be pending in
the target's queue.
SessionStart injects the queue as its own block below the mailbox block.
Two channels, two blocks, mail first: it carries Rule 7, and the queue
order is mail -> orders -> STATE's NESTE.
Also folds in dde392d (board prefix-match fix), which landed after the
0.26.0 bump and before any tag. v0.26.0 was never tagged, so 0.27.0 is the
release that carries all of it.
Suites: coord 220, board 237, route 69, orders 97, guard 40; npm test 11/11.
Antakelse 4 (atomic claim) and antakelse 6 (morning --plan-file --dry-run
reports 1 of 1 for the thin starter) both measured, not assumed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0134iB7ipXGgEpv9imYoVmr2
209 lines
11 KiB
Markdown
209 lines
11 KiB
Markdown
---
|
|
name: route
|
|
description: >-
|
|
Decide which model and reasoning effort the NEXT session should run with, by
|
|
scoring four traits of the next task and running them through the rubric row
|
|
table (`route.sh`). Use at session end, whenever STATE.md's next step is
|
|
written or rewritten, and whenever the operator asks what to launch the next
|
|
session with: "what model should I use next", "which effort level", "close
|
|
the session", "wrap up", "update STATE", "what should the next session run
|
|
with", "is opus overkill here", "give me the startup command". Also triggers
|
|
on Norwegian phrasings: "hvilken modell neste økt", "hvilken effort", "avslutt
|
|
sesjonen", "oppdater STATE", "hva skal neste økt kjøre med", "er opus
|
|
overkill", "gi meg oppstartskommandoen", "modell og effort". Trigger even when
|
|
the operator names no model and no tool — choosing the model for the next
|
|
session IS this skill. Not for choosing which REPO gets the next session:
|
|
that is the `board` skill.
|
|
version: "0.27.0"
|
|
---
|
|
|
|
# route — what the next session should run with
|
|
|
|
The operator's closing line requires two things every single session: the model
|
|
and effort for the next session, and a pasteable startup command. This skill
|
|
produces both from evidence instead of from a hunch, and leaves the scoring
|
|
behind in STATE.md so a wrong call can be found later.
|
|
|
|
**The division of labour is the whole design.** Scoring the four traits is
|
|
judgement and belongs to you. Turning scores into a model is a lookup, and
|
|
`route.sh` does it deterministically — same scores, same answer, every time,
|
|
at zero token cost. Never "reason your way" to a model. If you find yourself
|
|
weighing whether the task feels hard enough for Opus, you have skipped the
|
|
scoring step and are doing the thing this skill replaces.
|
|
|
|
## The engine
|
|
|
|
ROUTE="${CLAUDE_PLUGIN_ROOT}/scripts/route.sh"
|
|
|
|
"$ROUTE" --path <known|partial|undetermined> \
|
|
--verification <strong|weak|none> \
|
|
--reversibility <cheap|costly|one-way> \
|
|
--scope <local|multi-file|cross-cutting> \
|
|
--rationale "why these four scores" \
|
|
[--last-model <name> --last-effort <level> \
|
|
--last-completed <yes|no> --last-corrections <n>]
|
|
|
|
There is no deployed copy anywhere else and no fallback path — one entry point
|
|
is deliberate. Exit 2 means a bad or missing argument; read stderr and fix the
|
|
call rather than dropping the flag. `route.sh --help` carries the full row table
|
|
and the reasoning behind it.
|
|
|
|
The script writes nothing. It prints `row`, `rule`, `next-cost`, `command`,
|
|
`fallback`, `fallback-command`, `route-line` and (when the record is given)
|
|
`route-last`. **You** paste those into STATE.md.
|
|
|
|
## Scoring the four traits
|
|
|
|
Score the task the NEXT session will do — the one in the `👉 NESTE` block —
|
|
not the one that just finished. Read the next step as written before scoring;
|
|
if you cannot score it from what is written, that is a finding (see below).
|
|
|
|
| Trait | Score it by asking |
|
|
|---|---|
|
|
| `path` | Is the solution route already described? `known` = the steps are written down or the pattern exists in this repo. `partial` = the shape is clear, one real decision is open. `undetermined` = it has to be found first. |
|
|
| `verification` | Will a machine catch the error? `strong` = tests, types or a compiler fail on it. `weak` = a smoke test or manual run would probably surface it. `none` = prose, API shape, a security judgement — a wrong answer just sits there. |
|
|
| `reversibility` | `cheap` = a commit away. `costly` = touches published state, needs a migration or a follow-up release. `one-way` = a pushed tag, a public interface, a deletion. |
|
|
| `scope` | `local` = one file or one function. `multi-file` = several files, one repo. `cross-cutting` = many subsystems, or more than one repo. |
|
|
|
|
`verification` carries the most signal and is the trait most often skipped.
|
|
Strong verification means a cheap model's mistakes get caught and corrected —
|
|
cheap model plus tight feedback beats an expensive model without it. When
|
|
nothing verifies the output, model quality is the only defence left.
|
|
|
|
Three rules that keep the scoring honest:
|
|
|
|
- **Score the task, never the feeling.** "Hard", "complex" and "important" are
|
|
not traits here. They are unfalsifiable, and they always resolve upward.
|
|
- **`--rationale` is required and is the point.** It is where a misscore is
|
|
caught weeks later, when the recommendation turns out to have been wrong. One
|
|
sentence naming the evidence: "the pattern exists in handlers/, but the error
|
|
handling is undecided" — not "medium difficulty".
|
|
- **When torn between two scores, take the more expensive one.** Escalation is
|
|
asymmetric by design: any one trait escalates, and row 1 needs all four at
|
|
the cheap end.
|
|
|
|
## The specification check you get for free
|
|
|
|
If the next step is scored `path=undetermined` and no design phase is planned,
|
|
**the task description is underspecified — the model is not too small.** Say so,
|
|
and rewrite the next step until it can be scored. Upgrading the model to
|
|
compensate for a vague specification is the most expensive form of
|
|
procrastination available, and it hides the real defect.
|
|
|
|
This check is worth more than the tokens the routing saves. Do not skip it by
|
|
scoring `partial` to keep things moving.
|
|
|
|
## Fable rows are never this calculator's output
|
|
|
|
`route.sh` only ever emits rows 1-4. Rows 5-6 (Fable) fired from an explicit
|
|
`--opus-xhigh-failed` flag until 2026-08-06, when the operator removed that
|
|
policy; nothing replaced it as a trait-derived outcome. Choosing Fable is now
|
|
always a deliberate deviation from the rubric — CLAUDE.md is explicit that the
|
|
rubric stays the only deterministic lookup and a departure from it is recorded
|
|
in STATE as an **override**, never as something this skill produces. If Fable
|
|
is the right call for the next step's *form* (big-picture, review, planning),
|
|
write the board line and the `rule` by hand — `board.sh` still parses
|
|
"Fable 5/high" and "Fable 5/xhigh" — and say so plainly in the rationale rather
|
|
than scoring the four traits to land there. One fact worth carrying into that
|
|
override: a Fable session runs without an advisor.
|
|
|
|
## The last-session record
|
|
|
|
Write it whenever all four fields are actually known, from what happened in
|
|
the session that is ending — never from what STATE.md prescribed:
|
|
|
|
- `--last-model` / `--last-effort` — what this session actually ran with. Both
|
|
are closed sets (`Sonnet 5|Opus 5|Fable 5`, and the verified effort levels),
|
|
because the next session compares these values rather than just displaying
|
|
them.
|
|
- `--last-completed yes|no` — did this session finish the next step the previous
|
|
STATE.md set out? Answer about that step, not about the session in general.
|
|
- `--last-corrections <n>` — how many rounds of rework it took. This is the
|
|
cheap proxy for whether the routing was right.
|
|
|
|
**Measure the effort, never infer it.** Read what this session actually
|
|
resolved:
|
|
|
|
```bash
|
|
echo "$CLAUDE_EFFORT"
|
|
```
|
|
|
|
`CLAUDE_EFFORT` is Claude Code's own *current* effort level, exported into every
|
|
tool-use context — which is why a Bash call can read it. Pass it verbatim as
|
|
`--last-effort`. The model is not in the environment (there is no
|
|
`CLAUDE_MODEL`); take it from what this session knows itself to be running as.
|
|
|
|
Two sources are **wrong on purpose**, and both fail the same way. The previous
|
|
board line holds what was *prescribed*, not what was *run* — the two come apart
|
|
exactly when the record would be most interesting. Asking the operator launders
|
|
that same prescription through a human, who is reading it off the startup
|
|
command they typed rather than off the running process. Confirming a measured
|
|
value with them is fine; sourcing it from them is not.
|
|
|
|
So the record no longer waits on anyone: all four fields are knowable from
|
|
inside the session that is ending. Still omit it entirely — all four or none —
|
|
if any one of them is genuinely unknown. A guessed value is worse than a missing
|
|
one, because it reads back later as a measurement.
|
|
|
|
**This file must never declare an `effort:` frontmatter field.** Skill
|
|
frontmatter overrides the session effort while the skill is active, so the
|
|
reading above would report *this skill's* effort instead of the session's — a
|
|
measurement measuring itself, with nothing in the output to show it happened.
|
|
Pinned by `route-selftest.sh` section 13.
|
|
|
|
One honest limit: `CLAUDE_EFFORT` is the *current* level, so if the operator
|
|
changed it mid-session with `/effort`, "the effort this session ran with" is not
|
|
a single value. Record the level the work was actually done at and say so.
|
|
|
|
## Writing it into STATE.md
|
|
|
|
Three single-line HTML comments sit directly under the `👉 NESTE` heading, in
|
|
this order:
|
|
|
|
<!-- board: status=in-progress; blocked-on=-; next-cost=Opus 5/high -->
|
|
<!-- route: path=partial; verification=strong; ...; rationale=... -->
|
|
<!-- route-last: model=Opus 5; effort=xhigh; completed=yes; corrections=1 -->
|
|
|
|
Splice the emitted `next-cost` value into the existing board line — leave
|
|
`status` and `blocked-on` alone, they answer a different question and this skill
|
|
knows nothing about them.
|
|
|
|
**They must stay single-line comments.** `board.sh` reads the first line under
|
|
the heading that is not blank, not a heading, and does not *start* with `<!--`,
|
|
and shows it as that repo's next step across every repo. A YAML block or a
|
|
comment broken across lines therefore replaces the repo's next step on the board
|
|
with `next_task:`. Measured, not assumed — `route-selftest.sh` section 7 pins it.
|
|
|
|
## Reporting it
|
|
|
|
Give the operator the two closing-line items and nothing more:
|
|
|
|
- **Modell neste økt:** the `next-cost` value, plus the `fallback` one row
|
|
cheaper for quota pressure. Name the `rule` that fired — "path=partial" — so
|
|
the call is auditable rather than asserted.
|
|
- **Oppstartskommando:** the `command` string, in its own code block, with
|
|
`/exit` named. **Never prefix it with `cd`:** one repo per terminal tab, so
|
|
the working directory is already right. If the next step belongs in a
|
|
different repo, say so in plain words — that is a different tab, not a `cd`.
|
|
|
|
**Paste `command` verbatim, `--advisor opus` included.** The calculator decides
|
|
the advisor per row, and it is not decoration: on a Sonnet row it is what lifts
|
|
the session to Opus judgement at Sonnet cost, which is what makes the cheaper
|
|
`fallback-command` safe to take under quota pressure. Dropping it because it
|
|
looks like noise silently removes that. Equally, never *add* it to a command
|
|
that came back without one — an unconditional advisor is the global
|
|
`advisorModel` setting, which costs quota in every session in every repo and is
|
|
the failure mode this rule replaces. `route.sh --help` carries the full rule.
|
|
|
|
If `command` and `fallback-command` are the same as the current session's model,
|
|
say `/clear` is enough instead — but only if no newly installed plugin or skill
|
|
needs a fresh process to be picked up.
|
|
|
|
**`--advisor` is part of that comparison, not an afterthought.** It is a launch
|
|
flag, so `/clear` reuses the process and keeps whatever advisor the session
|
|
started with. If `command` carries `--advisor opus` and this session was not
|
|
launched with it, `/clear` is *not* enough — the operator needs `/exit` and the
|
|
full command, or the advisor silently never appears.
|
|
|
|
Do not paste the whole output block. One row, the rule that produced it, the
|
|
command.
|