repo-mailbox/CLAUDE.md
Kjell Tore Guttormsen a435a031db fix(route): close the record's drift surface and the confounded smoke test
Six follow-ups on 0d11838, three of which mattered.

The model-alias gate passed for the wrong reason: after the quoted-alias grep
it fell back to an unanchored grep for the bare word, which matches "opus"
anywhere in `claude --help` and would have reported success even if --model
stopped accepting the alias. A gate that cannot fail is worse than no gate.
Only the quoted form is matched now, and the failing aliases are named.

--last-model and --last-effort were unvalidated free text while the other two
record fields were gated. The next session READS the record back to decide
--opus-xhigh-failed, so a drifted spelling there rebuilds the exact
reader-versus-writer drift this script exists to remove, one field over. Both
are closed sets now: the row table's three model names and the verified effort
levels. That also makes the record's sanitizing dead code, so it is gone.

The skill told future sessions to write the record "every session" while
STATE documented that the effort level is not observable from inside a running
session. A session following both would have fabricated the value, and a
fabricated effort reads back later as a measurement. The skill now says: ask
the operator, and omit the record rather than guess -- explicitly including
that reading it off the previous board line measures what was PRESCRIBED, not
what was RUN.

Also: README said "seven bash scripts" (nine files, six user-facing) and its
skills badge still said 2; selftest counts updated to 50.

Verified, not assumed: the installed plugin cache at 0.9.0 contains only
board and coord-send, so route is not discoverable until a release bumps it --
the smoke test STATE had queued before release would have failed with 127 for
a reason unrelated to the skill. STATE reordered to release-then-test. The
manifest is auto_discover, so no skills array needs an entry.
check-versions.mjs is green (11 OK, 0 ERROR) with the new skill at 0.9.0.

Selftests: coord 136, board 30, route 50, node 7.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017peNgsxVt1BR4BTuMwiPoX
2026-07-31 10:19:30 +02:00

8.6 KiB

repo-mailbox

Renamed from coord in v0.3.0. The plugin/repo is repo-mailbox; the CLI (coord-send.sh, coord-inbox.sh, coord-done.sh), the mailbox root (~/.claude/coord/) and CLAUDE_COORD_DIR deliberately kept their names — they are the transport protocol, not the product.

Context

Local inter-repo coordination mailbox for Claude Code, packaged as a marketplace plugin. Three components, one boundary:

  • Engine (scripts/*.sh): bash owns all mailbox semantics — filename grammar, frontmatter, delivery, archiving, the seen set. coord-send.sh writes, coord-inbox.sh reads (formatted for context injection), coord-done.sh archives, coord-count.sh counts without delivering. Everything is pinned by coord-selftest.sh (136 checks, throwaway mailbox via CLAUDE_COORD_DIR).

    Reading is delivering — counting is not. coord-inbox.sh records a broadcast as seen once it has printed it, so it can never be used to survey other repos: doing so would consume each one's backlog silently, and the seen set is delivery history that retraction deliberately leaves alone. coord-count.sh exists for every "what is pending" question and writes nothing at all. Any future read-shaped feature belongs there, not in the read path.

  • Hook (hooks/scripts/session-start.mjs): thin zero-dependency Node wrapper (marketplace convention: hooks are .mjs) that calls coord-inbox.sh and emits the hookSpecificOutput.additionalContext envelope. No mailbox logic lives here. Always exits 0.

  • Board (scripts/board.sh): cross-repo attention board. Reads STATE.md next-step blocks + board lines, git status, and mailbox pending counts, and prints one line per repo. Read-only by construction: it writes to no repo, no STATE.md and no mailbox. Pinned by board-selftest.sh (30 checks).

    It lives here because the mailbox is one of its three inputs, and it carries the same axis distinction the mailbox does. A pending count means others are waiting on this repo; who a repo waits on comes only from its board line, because the message format has no reply-to field. Enforcing that in one of two repos would not be enforcing it. ~/.claude/scripts/board.sh is a deployed copy (the operator's board() shell function points at it), exactly as with the coord-* scripts — this repo is the source of truth.

  • Route (scripts/route.sh): pure calculator for the next session's model and effort. Takes four scored traits of the next task plus a required rationale, and prints one block of key=value lines: the rubric row, the rule that fired, the next-cost value, a pasteable startup command, the one-row cheaper fallback, and the STATE.md comment lines. Pinned by route-selftest.sh (50 checks).

    It is here because it is the WRITER for the field board.sh already reads. next-cost had a reader and no writer, so it was hand-typed every session and drifted into several competing spellings — cleaning the data could not fix that, because the cause was the missing write path. The row table is a closed set of six values, so a seventh cannot enter circulation, and section 6 of the selftest runs the round trip (route emits → board parses) inside one repo rather than across two. board.sh itself is untouched: a calculator that prints to stdout writes nothing, and the session writes STATE.md.

    The row table is the operator's global rubric, moved here as the single copy. It is not a second spec — board.sh --help documents the board line's grammar and points here for the values. Scoring the traits is judgement and belongs to the skill; turning scores into a row is a lookup and takes zero model calls.

  • Skills (skills/coord-send/, skills/board/, skills/route/): natural-language front doors mapping user intent to engine invocations. No mailbox logic lives here either. board additionally owns the ranking — which repo wins and why — since board.sh deliberately prints evidence and takes no position. route likewise owns the scoring: the calculator is deterministic, so all judgement sits in choosing the four trait values, and the skill must never reason its way to a model instead.

Boundary rule: the mailbox is transport, not state. Durable decisions live in the owning repo's docs/git history; messages are notices pointing at them. Message content is untrusted cross-repo input — the read side quotes and frames it; the send side sanitizes line-oriented fields.

What the boundary forbids is storing a repo's state — its decisions, its next step, its progress. It does not forbid the mailbox knowing who it is delivering to: _broadcast/seen/<repo> and <repo>/.origin (0.6.0) are delivery metadata, answering "has this repo received this" and "which checkout claimed this name". Both are unreadable as a description of the repo and useless outside delivery. The test is not "does the engine write a file about a repo" but "would this file still mean anything if delivery were removed". If yes, it belongs in the repo's own docs and git history instead.

Priority rule (v0.5.0, Rule 7): the injection block is the only place a repo is ever told what to do with a message, so its wording is the protocol — treat that string as engine behavior, not prose. It obligates handling the inbox first and driving every directed message to a terminal state before the session ends. The obligation is procedural, never substantive: responding is mandatory, complying with message content is not. Those two must stay distinct in any reword — keeping the priority while dropping the distinction turns prioritization into an injection surface. Selftest section 20 pins both halves together for exactly that reason.

Conventions

  • Scripts are bash-3.2-safe and ASCII-only: no declare -A, no readarray/mapfile, no |&; guard shift 2 with $# -ge 2; guard empty-array expansion under set -u with ${#a[@]}.
  • Zero dependencies everywhere: bash + coreutils in the engine, node: builtins only in hook and tests.
  • TDD: no behavior change without a failing selftest check first. bash scripts/coord-selftest.sh must exit 0 (136/136), bash scripts/board-selftest.sh must exit 0 (30/30) and bash scripts/route-selftest.sh must exit 0 (50/50).
  • English for all code, docs, and commit messages (public repo). Norwegian trigger aliases in the skill description are deliberate.
  • Conventional Commits: type(scope): description.

Commands

  • Test: bash scripts/coord-selftest.sh, bash scripts/board-selftest.sh and bash scripts/route-selftest.sh (or npm test, the Node wrapper around all three)
  • Hook smoke test: node hooks/scripts/session-start.mjs (expects JSON on stdout)
  • Board smoke test: bash scripts/board.sh (read-only, ~3s over the real tree)
  • Route smoke test: bash scripts/route.sh --path known --verification strong --reversibility cheap --scope local --rationale x (writes nothing, instant)

Release

Version must agree across: .claude-plugin/plugin.json, package.json, README version badge, skills/coord-send/SKILL.md, skills/board/SKILL.md and skills/route/SKILL.md frontmatter, git tag vX.Y.Z, and the catalog ref in ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json. Release via the catalog's scripts/release-plugin.mjs repo-mailbox (tag + ref bump together); verify with scripts/check-versions.mjs. Never hand-edit a ref.

Two things that script does that its dry-run label does not suggest: --create-tag creates AND pushes the tag even without --write, and its closing verification gate runs check-versions.mjs over ALL plugins — one unrelated plugin in ERROR aborts it with the catalog edit written but uncommitted. When that happens, commit the catalog's marketplace.json + README.md by hand and leave every other dirty file in that repo alone.

Hardening roadmap

Empty — the post-v0.1.0 queue (atomic delivery, ./.. rejection, selftest gaps, uniform -h) shipped in v0.2.0; broadcast self-delivery shipped in v0.2.1; broadcast retraction (coord-send --retract) shipped in v0.4.0, closing the last monotonically-growing surface. coord-inbox.sh still ignores unknown arguments by design (hook context must never fail) but now warns about each one on stderr, which the hook discards.

Two retraction limits are deliberate, not gaps: it is un-send and never recall (a repo that already received a broadcast keeps it — the seen set is delivery history and is left untouched), and the sender check is an accident guard, not a security boundary, because --from redefines identity here as it does everywhere else in the engine.