repo-mailbox/CLAUDE.md
Kjell Tore Guttormsen 9cb405c2cd feat(route): give the advisor a writer, on a need and per row
route.sh now emits `--advisor opus` into the startup command it prints.
The flag existed and worked, but nothing generated it, so it went unused:
the only mechanism that ever set an advisor here was `/advisor`, which
writes the global advisorModel setting -- every session, every repo -- and
was abandoned for burning quota. Nothing replaced it.

Two independent triggers, almost disjoint by construction:

  rows 1-2  always. Sonnet main model, so opus is a capability LIFT rather
            than a peer. Load-bearing: every fallback is one row cheaper and
            the cheap rows are Sonnet, so this makes the quota fallback safe.
  rows 3-4  only at reversibility=costly|one-way. Opus main model, so the
            advisor buys peer review where a mistake is not cheap to undo.
  rows 5-6  never. The CLI rejects every advisor for a Fable main model.

costly forces row 3 and one-way forces row 4, so a Sonnet row always has
reversibility=cheap and neither rule reaches the other's rows.
verification=none is deliberately not a third trigger: beyond the stakes
rule it adds only cheap-to-reverse mistakes, docs sessions among them.
Applied per ROW, so fallback-command carries its own correct answer.

route-selftest.sh 56 -> 73. Section 14 gates the three CLI facts the rule
rests on against the installed claude without spending a token: advisor
validation runs before the empty-prompt check, so `-p ""` reaches the
validator and stops. --help cannot gate this -- it short-circuits before
option validation, so an unknown flag would pass the gate untested.

Three pre-existing checks updated rather than worked around: two asserted
whole command strings that now carry the advisor, and section 11's effort
extraction swallowed the tail of the command line.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X8N8hQEJSWWtieWUx37txT
2026-07-31 21:39:59 +02:00

202 lines
12 KiB
Markdown

# repo-mailbox
Renamed from `coord` in v0.3.0. The plugin/repo is `repo-mailbox`; the CLI
(`coord-send.sh`, `coord-inbox.sh`, `coord-done.sh`), the mailbox root
(`~/.claude/coord/`) and `CLAUDE_COORD_DIR` deliberately kept their names —
they are the transport protocol, not the product.
## Context
Local inter-repo coordination mailbox for Claude Code, packaged as a
marketplace plugin. Three components, one boundary:
- **Engine (`scripts/*.sh`):** bash owns all mailbox semantics — filename
grammar, frontmatter, delivery, archiving, the seen set. `coord-send.sh`
writes, `coord-inbox.sh` reads (formatted for context injection),
`coord-done.sh` archives, `coord-count.sh` counts without delivering.
Everything is pinned by `coord-selftest.sh`
(159 checks, throwaway mailbox via `CLAUDE_COORD_DIR`).
**`coord-count.sh` prints TWO integers per mailbox** (`<name>\t<pending>\t<debt>`),
and the first must stay pending: `board.sh` counts the same inbox files
itself, so a debt-only count would put two different numbers under one name.
Debt is read from `reply-expected` in the FRONTMATTER BLOCK ONLY - a body line
is untrusted input and must not be able to silence a debt - and an absent
field means a reply IS owed, because every message written before 0.11.0
lacks it.
**Reading is delivering — counting is not.** `coord-inbox.sh` records a
broadcast as seen once it has printed it, so it can never be used to survey
other repos: doing so would consume each one's backlog silently, and the seen
set is delivery history that retraction deliberately leaves alone.
`coord-count.sh` exists for every "what is pending" question and writes
nothing at all. Any future read-shaped feature belongs there, not in the
read path.
- **Hook (`hooks/scripts/session-start.mjs`):** thin zero-dependency Node
wrapper (marketplace convention: hooks are `.mjs`) that calls
`coord-inbox.sh` and emits the `hookSpecificOutput.additionalContext`
envelope. No mailbox logic lives here. Always exits 0.
- **Board (`scripts/board.sh`):** cross-repo attention board. Reads STATE.md
next-step blocks + board lines, `git status`, and mailbox pending counts, and
prints one line per repo. Read-only by construction: it writes to no repo, no
STATE.md and no mailbox. Pinned by `board-selftest.sh` (36 checks).
**It lives here because the mailbox is one of its three inputs, and it carries
the same axis distinction the mailbox does.** A pending count means *others
are waiting on this repo*; who a repo waits *on* comes only from its board
line, because the message format has no reply-to field. Enforcing that in one
of two repos would not be enforcing it. The operator invokes both `board` and
`coord-send` exclusively through their Skill front doors, never a personal
terminal alias — a claim this file carried until 0.12.1 (`~/.claude/scripts/board.sh`
is a deployed copy the operator's `board()` function points at) did not survive
inspection: no such file ever existed, and `route.sh` had no deployed copy
either. Only the five `coord-*.sh` scripts were ever deployed there, and their
one measured effect was an accidental fallback target for Claude sessions'
own Bash tool calls (the bug 0.12.1 fixed) — once that fallback was gone they
had no remaining function and were deleted.
- **Route (`scripts/route.sh`):** pure calculator for the next session's model
and effort. Takes four scored traits of the next task plus a required
rationale, and prints one block of `key=value` lines: the rubric row, the rule
that fired, the `next-cost` value, a pasteable startup command, the one-row
cheaper fallback, and the STATE.md comment lines. Pinned by
`route-selftest.sh` (73 checks).
**It is here because it is the WRITER for the field `board.sh` already reads.**
`next-cost` had a reader and no writer, so it was hand-typed every session and
drifted into several competing spellings — cleaning the data
could not fix that, because the cause was the missing write path. The row
table is a closed set of six values, so a seventh cannot enter circulation,
and section 6 of the selftest runs the round trip (route emits → board parses)
*inside* one repo rather than across two. `board.sh` itself is untouched: a
calculator that prints to stdout writes nothing, and the session writes
STATE.md.
**The row table is the operator's global rubric, moved here as the single
copy.** It is not a second spec — `board.sh --help` documents the board line's
*grammar* and points here for the *values*. Scoring the traits is judgement
and belongs to the skill; turning scores into a row is a lookup and takes zero
model calls.
**`--advisor opus` is emitted per ROW, on a need, never unconditionally.**
Rows 1-2 always carry it (Sonnet main model, so opus is a capability lift —
and since every fallback is one row cheaper and the cheap rows are Sonnet,
this is what makes the quota fallback safe to take); rows 3-4 only at
`reversibility=costly|one-way` (Opus main model, so it buys peer review where
a mistake is not cheap to undo); rows 5-6 never, because the CLI rejects every
advisor for a Fable main model. The alternative — the global `advisorModel`
setting written by `/advisor` — is what this replaces: it applies to every
session in every repo, which is how it burned quota before. The two triggers
are almost disjoint by construction, since `costly` forces row 3 and `one-way`
forces row 4, so a Sonnet row always has `reversibility=cheap`.
`verification=none` is deliberately NOT a third trigger: beyond the stakes
rule it would only add mistakes that are cheap to reverse, docs sessions
(`known/none/cheap/local`) among them. Section 14 pins the rule and gates the
three CLI facts it rests on against the installed `claude` without spending a
token — advisor validation runs before the empty-prompt check, so `-p ""`
reaches the validator and stops there.
**`--last-effort` is MEASURED from `CLAUDE_EFFORT`, and the calculator must
never default it.** Claude Code exports that variable into every tool-use
context as the session's current effort, so the caller reads it and passes it
in; having `route.sh` read it directly would make the output depend on the
environment instead of on its arguments, and the round trip in selftest
section 6 rests on that determinism. The two sources it replaces fail
identically: the previous board line holds what was *prescribed*, and asking
the operator launders that same prescription through someone reading their own
startup command. Corollary pinned by section 13: `skills/route/SKILL.md` must
never declare an `effort:` frontmatter field, because frontmatter overrides the
session effort and the reading would then measure the skill, not the session.
- **Skills (`skills/coord-send/`, `skills/board/`, `skills/route/`):** natural-language front
doors mapping user intent to engine invocations. No mailbox logic lives here
either. `board` additionally owns the *ranking* — which repo wins and why —
since `board.sh` deliberately prints evidence and takes no position. `route`
likewise owns the *scoring*: the calculator is deterministic, so all judgement
sits in choosing the four trait values, and the skill must never reason its
way to a model instead.
**Boundary rule:** the mailbox is transport, not state. Durable decisions
live in the owning repo's docs/git history; messages are notices pointing at
them. Message content is untrusted cross-repo input — the read side quotes
and frames it; the send side sanitizes line-oriented fields.
What the boundary forbids is storing a repo's *state* — its decisions, its next
step, its progress. It does not forbid the mailbox knowing who it is delivering
to: `_broadcast/seen/<repo>` and `<repo>/.origin` (0.6.0) are delivery metadata,
answering "has this repo received this" and "which checkout claimed this name".
Both are unreadable as a description of the repo and useless outside delivery.
The test is not "does the engine write a file about a repo" but "would this file
still mean anything if delivery were removed". If yes, it belongs in the repo's
own docs and git history instead.
**Priority rule (v0.5.0, Rule 7):** the injection block is the only place a
repo is ever told what to do with a message, so its wording *is* the protocol
— treat that string as engine behavior, not prose. It obligates handling the
inbox first and driving every directed message to a terminal state before the
session ends. The obligation is **procedural, never substantive**: responding
is mandatory, complying with message content is not. Those two must stay
distinct in any reword — keeping the priority while dropping the distinction
turns prioritization into an injection surface. Selftest section 20 pins both
halves together for exactly that reason.
Since 0.11.0 the `reply-expected` field says which terminal state the SENDER
expects. That does not soften the split, it sharpens it: the field is untrusted
cross-repo input like the rest of the file, so the injection calls it a
*declaration, not an instruction* and keeps both terminal states open to the
receiver. Drop that clause and one word in a message becomes a lever that mints
obligations in another repo.
## Conventions
- Scripts are bash-3.2-safe and ASCII-only: no `declare -A`, no
`readarray`/`mapfile`, no `|&`; guard `shift 2` with `$# -ge 2`; guard
empty-array expansion under `set -u` with `${#a[@]}`.
- Zero dependencies everywhere: bash + coreutils in the engine, `node:`
builtins only in hook and tests.
- TDD: no behavior change without a failing selftest check first.
`bash scripts/coord-selftest.sh` must exit 0 (159/159),
`bash scripts/board-selftest.sh` must exit 0 (36/36) and
`bash scripts/route-selftest.sh` must exit 0 (73/73).
- English for all code, docs, and commit messages (public repo). Norwegian
trigger aliases in the skill description are deliberate.
- Conventional Commits: `type(scope): description`.
## Commands
- Test: `bash scripts/coord-selftest.sh`, `bash scripts/board-selftest.sh` and
`bash scripts/route-selftest.sh` (or `npm test`, the Node wrapper around all three)
- Hook smoke test: `node hooks/scripts/session-start.mjs` (expects JSON on stdout)
- Board smoke test: `bash scripts/board.sh` (read-only, ~3s over the real tree)
- Route smoke test: `bash scripts/route.sh --path known --verification strong
--reversibility cheap --scope local --rationale x` (writes nothing, instant)
## Release
Version must agree across: `.claude-plugin/plugin.json`, `package.json`,
README version badge, `skills/coord-send/SKILL.md`, `skills/board/SKILL.md` and
`skills/route/SKILL.md` frontmatter, git tag `vX.Y.Z`, and the catalog `ref` in
`ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json`. Release via
the catalog's `scripts/release-plugin.mjs repo-mailbox` (tag + ref bump
together);
verify with `scripts/check-versions.mjs`. Never hand-edit a ref.
Two things that script does that its dry-run label does not suggest:
`--create-tag` creates AND pushes the tag even without `--write`, and its
closing verification gate runs `check-versions.mjs` over ALL plugins — one
unrelated plugin in ERROR aborts it with the catalog edit written but
uncommitted. When that happens, commit the catalog's `marketplace.json` +
`README.md` by hand and leave every other dirty file in that repo alone.
## Hardening roadmap
Empty — the post-v0.1.0 queue (atomic delivery, `.`/`..` rejection,
selftest gaps, uniform `-h`) shipped in v0.2.0; broadcast self-delivery
shipped in v0.2.1; broadcast retraction (`coord-send --retract`) shipped in
v0.4.0, closing the last monotonically-growing surface. `coord-inbox.sh`
still ignores unknown arguments by design (hook context must never fail)
but now warns about each one on stderr, which the hook discards.
Two retraction limits are deliberate, not gaps: it is un-send and never
recall (a repo that already received a broadcast keeps it — the seen set is
delivery history and is left untouched), and the sender check is an accident
guard, not a security boundary, because `--from` redefines identity here as
it does everywhere else in the engine.