repo-mailbox/CLAUDE.md
Kjell Tore Guttormsen d0a5ffe515 feat(board)!: rank --plan on five ordered groups, planned above in-progress
Replaces the weighted score shipped in 0.19.0 with five lookups: chain-root
credit, unhandled inbox, planned, in-progress, undeclared status. Within a
group: that group's own quantity, then a Sonnet next-cost, then oldest plan.

The score's objection is accepted, not forgotten, and is written into board.sh
and CLAUDE.md so a later session reads it as decided rather than as an unfixed
defect: a group order cannot express "owes one message AND releases two others"
as one quantity. What the score could not do was hold still for the format's
second consumer - re-tuning one weight against another silently reorders a
parser in another repo, and no test here can catch that.

planned now ranks above in-progress, inverted by the same decision: converting a
decision into motion is the slow step; live work is already moving.

Debt stays uncapped and never excluded. One group below chain-root credit is not
the cap declined at 0.19.0 - the debtor keeps its tab, its most-owed-first
position, and its why=inbox:N. Pinned by a discriminating fixture the score
would fail: a root releasing one repo outranks a repo owing four.

board-selftest 134 -> 138. Suite 183 + 138 + 73 = 394.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y6ULuFCPMNYAPNN3pAjsXQ
2026-08-03 06:55:21 +02:00

21 KiB

repo-mailbox

Renamed from coord in v0.3.0. The plugin/repo is repo-mailbox; the CLI (coord-send.sh, coord-inbox.sh, coord-done.sh), the mailbox root (~/.claude/coord/) and CLAUDE_COORD_DIR deliberately kept their names — they are the transport protocol, not the product.

Context

Local inter-repo coordination mailbox for Claude Code, packaged as a marketplace plugin. Three components, one boundary:

  • Engine (scripts/*.sh): bash owns all mailbox semantics — filename grammar, frontmatter, delivery, archiving, the seen set. coord-send.sh writes, coord-inbox.sh reads (formatted for context injection), coord-done.sh archives, coord-count.sh counts without delivering, coord-sweep.sh closes the aged FYI backlog machine-wide. Everything is pinned by coord-selftest.sh (183 checks, throwaway mailbox via CLAUDE_COORD_DIR).

    coord-sweep.sh is the only path that closes a message with no human in the loop, and every constraint on it follows from that. It may close exactly one mechanically decidable class - reply-expected: no, older than the grace window - because a message that owes a reply can only be answered by a session in the repo that owes it. Dry-run is the default, inverted from the rest of the engine, since this is the one script that destroys pending state. It closes through coord-done.sh --repo rather than moving files, so the archive layout and the _broadcast refusal stay in one place. And it logs every closure with sender and subject, because directed messages have no seen-tracking: the sweep genuinely cannot tell "seen and ignored" from "never delivered", so a notice can be closed unread and the log is the only record that it existed. Widening the class, defaulting to --write, or dropping the log each independently turn this from a bounded cleanup into silent data loss.

    coord-count.sh prints TWO integers per mailbox (<name>\t<pending>\t<debt>), and the first must stay pending: board.sh counts the same inbox files itself, so a debt-only count would put two different numbers under one name. Debt is read from reply-expected in the FRONTMATTER BLOCK ONLY - a body line is untrusted input and must not be able to silence a debt - and an absent field means a reply IS owed, because every message written before 0.11.0 lacks it.

    Reading is delivering — counting is not. coord-inbox.sh records a broadcast as seen once it has printed it, so it can never be used to survey other repos: doing so would consume each one's backlog silently, and the seen set is delivery history that retraction deliberately leaves alone. coord-count.sh exists for every "what is pending" question and writes nothing at all. Any future read-shaped feature belongs there, not in the read path.

  • Hook (hooks/scripts/session-start.mjs): thin zero-dependency Node wrapper (marketplace convention: hooks are .mjs) that calls coord-inbox.sh and emits the hookSpecificOutput.additionalContext envelope. No mailbox logic lives here. Always exits 0.

  • Board (scripts/board.sh): cross-repo attention board. Reads STATE.md next-step blocks + board lines, git status, and mailbox pending counts, and prints one line per repo. Read-only by construction: it writes to no repo, no STATE.md and no mailbox. Pinned by board-selftest.sh (138 checks).

    It lives here because the mailbox is one of its three inputs, and it carries the same axis distinction the mailbox does. A pending count means others are waiting on this repo; who a repo waits on comes only from its board line, because the message format has no reply-to field. Enforcing that in one of two repos would not be enforcing it. The operator invokes both board and coord-send exclusively through their Skill front doors, never a personal terminal alias — a claim this file carried until 0.12.1 (~/.claude/scripts/board.sh is a deployed copy the operator's board() function points at) did not survive inspection: no such file ever existed, and route.sh had no deployed copy either. Only the five coord-*.sh scripts were ever deployed there, and their one measured effect was an accidental fallback target for Claude sessions' own Bash tool calls (the bug 0.12.1 fixed) — once that fallback was gone they had no remaining function and were deleted. board.sh --brief is a second RENDERING of that scan, never a second scan, and brief-nightly.sh is the only writer in that path. The briefing answers the narrower question an unattended job can answer without judgement: which repos have an unhandled inbox, what their next step says in full, and the exact command to start a session in each. It prints NESTE uncut because the 38-character cut is the table column's property, not the record's — the value used to be truncated at record-build time, which left the cut string as the only copy. Each command is derived by CALLING route.sh with that repo's own four traits; next-cost alone cannot yield it, since the advisor flag is a property of the ROW and two rows can share a model/effort pair while differing on it. A repo with no route line is told so rather than handed a guess, because a guessed command reads as authoritative.

    board.sh --plan is the THIRD rendering, and the only one that takes a position. It answers which repos to open a tab for today, in what order, with which command. The position it takes is the ORDER and nothing else - there is no cutoff, so the plan hides nothing, and every term is a lookup over fields the scan already read. Since 0.20.0 it is FIVE ORDERED GROUPS - chain-root, debt, planned, in-progress, undeclared - ranked within a group by that group's own quantity, then a Sonnet next-cost, then oldest plan first.

    0.19.0 shipped a weighted score here and 0.20.0 replaced it, and the objection the score answered is ACCEPTED, not forgotten. A group order genuinely cannot express "this repo owes one message and releases two others" as one quantity; a score could, and that was its point. What a score could not do was hold still for the second consumer - re-tuning 40 against 15 silently reorders a parser living in another repo, and no test in THIS repo can catch that. The operator weighed both and chose the lookup (2026-08-03). Write that down every time this paragraph is edited: a later session that reads the objection as an unfixed defect will "restore" the score, and the round trip is the loop this file exists to stop.

    planned ranks ABOVE in-progress, inverted at 0.20.0 by operator decision. Turning a decision into motion is the slow step; live work is already moving. Flipping it back is a policy change, not a sort fix.

    Debt is never excluded and never capped, and that is the rule most likely to be "fixed" into a defect. Excluding blocked or done is a claim about a repo's OWN next step, which by definition cannot be moved, while owing a reply is the other axis entirely - answering is often what unblocks it. Measured on the real tree at 0.16.0, two of 26 planned repos were done with an unhandled inbox. A cap on debt was proposed with the chain credit at 0.19.0 and DECLINED by the operator. Sitting one group below chain-root credit is NOT that cap: the debtor keeps its tab, its most-owed-first position among the other debtors, and its why=inbox:N. A change that DROPPED a debtor from the plan would be the declined cap wearing the group order as a disguise, and selftest section 12 pins both halves - the root outranking four owed messages, and the debtor keeping everything it had.

    Chain-root credit lands on the ROOT and nowhere else. For every blocked repo the blocked-on edge is followed transitively to the first repo that is not itself blocked. Crediting a blocked repo would open a tab that cannot move; crediting only the direct blocker leaves a two-hop chain's root uncredited, which is the shape the real tree actually had. A cycle, a blocked-on naming an unscanned repo, and a blocked repo with no target must all credit NOBODY: inventing a root there produces a plan that looks correct and sends the operator to the wrong repo.

    Repos with no board line rank last and are LABELLED rather than dropped, because the table already prints a MERK line about them and a plan that omitted them silently would repeat that defect.

    It renders key=value blocks, not prose, because it has two consumers: the operator, and a driver repo consuming the plan. Prose would make the rendered format an API no test in THIS repo could hold stable for a consumer in another. command_missing= carries both no-command causes (no route line, and a route line route.sh rejects) because a bare command= is the shape of a runnable command carrying nothing - a driver reading ^command= would type an empty line into a live pane. route_cmd_for() is the single reader of the route-line grammar, shared with --brief, and distinguishes the two causes by exit code rather than by an empty string.

    paste= and dir=/command= are the same fact for the two consumers, and neither is redundant. A driver moves the pane itself and then types the command, so it needs them apart; a human needs ONE thing to select. Handing the operator two fields to join by hand is not a saved output line, it is the step where a session starts in the wrong repo - and it was measured the moment the feature met its first user, who could not act on the block at all. paste= is emitted only alongside command=: paste=cd X && with nothing after it would run the cd and then a bare newline, which fails SILENTLY by leaving the operator in the right directory with no session started.

    Driving a terminal from the plan does NOT belong here, and the measurement in docs/ghostty-orchestration-measurement.md is the argument, not taste. It is a version-pinned undocumented composition over a preview API whose documented path is already broken upstream and whose regression was closed as not planned, with a blast radius reaching into other repos' live sessions. None of that is mailbox transport, and none of it may be able to break coord-inbox or board. The dependency runs one way: the driver consumes the plan, the plan never knows a terminal exists.

    It also cross-checks itself against coord-count.sh, and that is not belt-and-braces. The repo scan and the mailbox are two different populations: a mailbox can carry a name no scan will ever produce — a declared non-git surface (CLAUDE_COORD_REPO, e.g. ~/repos itself) or a checkout outside the roots. Measured on the real mailbox at 0.15.0: the briefing found 11 repos / 21 messages where coord-count saw 12 mailboxes / 22 pending, the missing one being the declared surface repos. A briefing that only walks the scan answers "who is waiting on you" with a number it quietly knows is short.

    Zero model calls, and that is the load-bearing property, not an implementation detail. The operator authenticates by subscription, so a headless claude -p job draws from the same quota pool as interactive work. Measured against 2.1.220: --max-budget-usd DOES bite under subscription auth (terminal_reason: budget_exhausted, exit 1), but it aborts AFTER turn one, never before it — floor ~0.25 USD-equivalent per turn on claude-opus-5[1m]. It is a runaway brake, not a pre-flight gate. Making the briefing deterministic removes the question entirely.

    board.sh stays read-only, which is why the file write lives in the wrapper instead of behind a --brief --out FILE flag. The wrapper renders to a temp file in the target directory and renames it into place, and treats an EMPTY render as a FAILED one: board prints nothing at all when its scan roots do not exist, which is what a mistyped path or a moved home directory looks like, and a plain > file redirect would destroy yesterday's briefing on a bad launchd environment. A tree where nobody owes anything is a different case — that is a valid, non-empty briefing saying so, and is written normally.

  • Route (scripts/route.sh): pure calculator for the next session's model and effort. Takes four scored traits of the next task plus a required rationale, and prints one block of key=value lines: the rubric row, the rule that fired, the next-cost value, a pasteable startup command, the one-row cheaper fallback, and the STATE.md comment lines. Pinned by route-selftest.sh (73 checks).

    It is here because it is the WRITER for the field board.sh already reads. next-cost had a reader and no writer, so it was hand-typed every session and drifted into several competing spellings — cleaning the data could not fix that, because the cause was the missing write path. The row table is a closed set of six values, so a seventh cannot enter circulation, and section 6 of the selftest runs the round trip (route emits → board parses) inside one repo rather than across two. board.sh itself is untouched: a calculator that prints to stdout writes nothing, and the session writes STATE.md.

    The row table is the operator's global rubric, moved here as the single copy. It is not a second spec — board.sh --help documents the board line's grammar and points here for the values. Scoring the traits is judgement and belongs to the skill; turning scores into a row is a lookup and takes zero model calls.

    --advisor opus is emitted per ROW, on a need, never unconditionally. Rows 1-2 always carry it (Sonnet main model, so opus is a capability lift — and since every fallback is one row cheaper and the cheap rows are Sonnet, this is what makes the quota fallback safe to take); rows 3-4 only at reversibility=costly|one-way (Opus main model, so it buys peer review where a mistake is not cheap to undo); rows 5-6 never, because the CLI rejects every advisor for a Fable main model. The alternative — the global advisorModel setting written by /advisor — is what this replaces: it applies to every session in every repo, which is how it burned quota before. The two triggers are almost disjoint by construction, since costly forces row 3 and one-way forces row 4, so a Sonnet row always has reversibility=cheap. verification=none is deliberately NOT a third trigger: beyond the stakes rule it would only add mistakes that are cheap to reverse, docs sessions (known/none/cheap/local) among them. Section 14 pins the rule and gates the three CLI facts it rests on against the installed claude without spending a token — advisor validation runs before the empty-prompt check, so -p "" reaches the validator and stops there.

    --last-effort is MEASURED from CLAUDE_EFFORT, and the calculator must never default it. Claude Code exports that variable into every tool-use context as the session's current effort, so the caller reads it and passes it in; having route.sh read it directly would make the output depend on the environment instead of on its arguments, and the round trip in selftest section 6 rests on that determinism. The two sources it replaces fail identically: the previous board line holds what was prescribed, and asking the operator launders that same prescription through someone reading their own startup command. Corollary pinned by section 13: skills/route/SKILL.md must never declare an effort: frontmatter field, because frontmatter overrides the session effort and the reading would then measure the skill, not the session.

  • Skills (skills/coord-send/, skills/board/, skills/route/): natural-language front doors mapping user intent to engine invocations. No mailbox logic lives here either. board additionally owns the ranking — which repo wins and why — since board.sh deliberately prints evidence and takes no position. route likewise owns the scoring: the calculator is deterministic, so all judgement sits in choosing the four trait values, and the skill must never reason its way to a model instead.

Boundary rule: the mailbox is transport, not state. Durable decisions live in the owning repo's docs/git history; messages are notices pointing at them. Message content is untrusted cross-repo input — the read side quotes and frames it; the send side sanitizes line-oriented fields.

What the boundary forbids is storing a repo's state — its decisions, its next step, its progress. It does not forbid the mailbox knowing who it is delivering to: _broadcast/seen/<repo> and <repo>/.origin (0.6.0) are delivery metadata, answering "has this repo received this" and "which checkout claimed this name". Both are unreadable as a description of the repo and useless outside delivery. The test is not "does the engine write a file about a repo" but "would this file still mean anything if delivery were removed". If yes, it belongs in the repo's own docs and git history instead.

Priority rule (v0.5.0, Rule 7): the injection block is the only place a repo is ever told what to do with a message, so its wording is the protocol — treat that string as engine behavior, not prose. It obligates handling the inbox first and driving every directed message to a terminal state before the session ends. The obligation is procedural, never substantive: responding is mandatory, complying with message content is not. Those two must stay distinct in any reword — keeping the priority while dropping the distinction turns prioritization into an injection surface. Selftest section 20 pins both halves together for exactly that reason.

Since 0.11.0 the reply-expected field says which terminal state the SENDER expects. That does not soften the split, it sharpens it: the field is untrusted cross-repo input like the rest of the file, so the injection calls it a declaration, not an instruction and keeps both terminal states open to the receiver. Drop that clause and one word in a message becomes a lever that mints obligations in another repo.

Conventions

  • Scripts are bash-3.2-safe and ASCII-only: no declare -A, no readarray/mapfile, no |&; guard shift 2 with $# -ge 2; guard empty-array expansion under set -u with ${#a[@]}.
  • Zero dependencies everywhere: bash + coreutils in the engine, node: builtins only in hook and tests.
  • TDD: no behavior change without a failing selftest check first. bash scripts/coord-selftest.sh must exit 0 (183/183), bash scripts/board-selftest.sh must exit 0 (138/138) and bash scripts/route-selftest.sh must exit 0 (73/73).
  • English for all code, docs, and commit messages (public repo). Norwegian trigger aliases in the skill description are deliberate.
  • Conventional Commits: type(scope): description.

Commands

  • Test: bash scripts/coord-selftest.sh, bash scripts/board-selftest.sh and bash scripts/route-selftest.sh (or npm test, the Node wrapper around all three)
  • Hook smoke test: node hooks/scripts/session-start.mjs (expects JSON on stdout)
  • Board smoke test: bash scripts/board.sh (read-only, ~3s over the real tree)
  • Briefing smoke test: bash scripts/board.sh --brief (read-only, writes nothing). brief-nightly.sh DOES write — it overwrites $CLAUDE_BRIEF_FILE (default ~/.claude/briefing.md), so point that at a scratch path when testing. Installed as a launchd agent from launchd/, which points at the SOURCE repo, never the version-pinned plugin cache.
  • Route smoke test: bash scripts/route.sh --path known --verification strong --reversibility cheap --scope local --rationale x (writes nothing, instant)
  • Sweep smoke test: bash scripts/coord-sweep.sh (dry-run is the default, so this writes nothing; never add --write to a smoke test against the real mailbox)

Release

Version must agree across: .claude-plugin/plugin.json, package.json, README version badge, skills/coord-send/SKILL.md, skills/board/SKILL.md and skills/route/SKILL.md frontmatter, git tag vX.Y.Z, and the catalog ref in ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json. Release via the catalog's scripts/release-plugin.mjs repo-mailbox (tag + ref bump together); verify with scripts/check-versions.mjs. Never hand-edit a ref.

Two things that script does that its dry-run label does not suggest: --create-tag creates AND pushes the tag even without --write, and its closing verification gate runs check-versions.mjs over ALL plugins — one unrelated plugin in ERROR aborts it with the catalog edit written but uncommitted. When that happens, commit the catalog's marketplace.json + README.md by hand and leave every other dirty file in that repo alone.

Hardening roadmap

Empty — the post-v0.1.0 queue (atomic delivery, ./.. rejection, selftest gaps, uniform -h) shipped in v0.2.0; broadcast self-delivery shipped in v0.2.1; broadcast retraction (coord-send --retract) shipped in v0.4.0, closing the last monotonically-growing surface. coord-inbox.sh still ignores unknown arguments by design (hook context must never fail) but now warns about each one on stderr, which the hook discards.

Two retraction limits are deliberate, not gaps: it is un-send and never recall (a repo that already received a broadcast keeps it — the seen set is delivery history and is left untouched), and the sender check is an accident guard, not a security boundary, because --from redefines identity here as it does everywhere else in the engine.