Found in review of the previous commit, before catalog tags 0.22.0: once n_owe counts OWED repos rather than raw pending, its ==0 branch could fire while a repo still held FYI-only mail, making "Ingen repo har uhaandtert innboks" false at the exact moment it printed. Fixed to state only the debt claim, and to name any FYI-only mailboxes found instead of letting their existence go unmentioned. No fixture in the shared test tree ever reached n_owe==0 (it always carries a debtor), so this needed its own isolated-root fixture to pin. Still part of the 0.22.0 release -- amends that changelog entry rather than bumping again, since no tag exists yet. board-selftest.sh: 150 -> 152 checks. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGWMPskXBsTjMrrQ2GofFx
27 KiB
repo-mailbox
Renamed from coord in v0.3.0. The plugin/repo is repo-mailbox; the CLI
(coord-send.sh, coord-inbox.sh, coord-done.sh), the mailbox root
(~/.claude/coord/) and CLAUDE_COORD_DIR deliberately kept their names —
they are the transport protocol, not the product.
Context
Local inter-repo coordination mailbox for Claude Code, packaged as a marketplace plugin. Three components, one boundary:
-
Engine (
scripts/*.sh): bash owns all mailbox semantics — filename grammar, frontmatter, delivery, archiving, the seen set.coord-send.shwrites,coord-inbox.shreads (formatted for context injection),coord-done.sharchives,coord-count.shcounts without delivering,coord-sweep.shcloses the aged FYI backlog machine-wide. Everything is pinned bycoord-selftest.sh(191 checks, throwaway mailbox viaCLAUDE_COORD_DIR).coord-sweep.shis the only path that closes a message with no human in the loop, and every constraint on it follows from that. It may close exactly one mechanically decidable class -reply-expected: no, older than the grace window - because a message that owes a reply can only be answered by a session in the repo that owes it. Dry-run is the default, inverted from the rest of the engine, since this is the one script that destroys pending state. It closes throughcoord-done.sh --reporather than moving files, so the archive layout and the_broadcastrefusal stay in one place. And it logs every closure with sender and subject, because directed messages have no seen-tracking: the sweep genuinely cannot tell "seen and ignored" from "never delivered", so a notice can be closed unread and the log is the only record that it existed. Widening the class, defaulting to--write, or dropping the log each independently turn this from a bounded cleanup into silent data loss.coord-count.shprints TWO integers per mailbox (<name>\t<pending>\t<debt>), and the first must stay pending:board.shcounts the same inbox files itself, so a debt-only count would put two different numbers under one name. Debt is read fromreply-expectedin the FRONTMATTER BLOCK ONLY - a body line is untrusted input and must not be able to silence a debt - and an absent field means a reply IS owed, because every message written before 0.11.0 lacks it.Reading is delivering — counting is not.
coord-inbox.shrecords a broadcast as seen once it has printed it, so it can never be used to survey other repos: doing so would consume each one's backlog silently, and the seen set is delivery history that retraction deliberately leaves alone.coord-count.shexists for every "what is pending" question and writes nothing at all. Any future read-shaped feature belongs there, not in the read path. -
Hook (
hooks/scripts/session-start.mjs): thin zero-dependency Node wrapper (marketplace convention: hooks are.mjs) that callscoord-inbox.shand emits thehookSpecificOutput.additionalContextenvelope. No mailbox logic lives here. Always exits 0. -
Board (
scripts/board.sh): cross-repo attention board. Reads STATE.md next-step blocks + board lines,git status, and mailbox pending counts, and prints one line per repo. Read-only by construction: it writes to no repo, no STATE.md and no mailbox. Pinned byboard-selftest.sh(152 checks).It lives here because the mailbox is one of its three inputs, and it carries the same axis distinction the mailbox does. A pending count means others are waiting on this repo; who a repo waits on comes only from its board line, because the message format has no reply-to field. Enforcing that in one of two repos would not be enforcing it. The operator invokes both
boardandcoord-sendexclusively through their Skill front doors, never a personal terminal alias — a claim this file carried until 0.12.1 (~/.claude/scripts/board.shis a deployed copy the operator'sboard()function points at) did not survive inspection: no such file ever existed, androute.shhad no deployed copy either. Only the fivecoord-*.shscripts were ever deployed there, and their one measured effect was an accidental fallback target for Claude sessions' own Bash tool calls (the bug 0.12.1 fixed) — once that fallback was gone they had no remaining function and were deleted.board.sh --briefis a second RENDERING of that scan, never a second scan, andbrief-nightly.shis the only writer in that path. The briefing answers the narrower question an unattended job can answer without judgement: which repos have an unhandled inbox, what their next step says in full, and the exact command to start a session in each. It prints NESTE uncut because the 38-character cut is the table column's property, not the record's — the value used to be truncated at record-build time, which left the cut string as the only copy. Each command is derived by CALLINGroute.shwith that repo's own four traits;next-costalone cannot yield it, since the advisor flag is a property of the ROW and two rows can share a model/effort pair while differing on it. A repo with no route line is told so rather than handed a guess, because a guessed command reads as authoritative.board.sh --planis the THIRD rendering, and the only one that takes a position. It answers which repos to open a tab for today, in what order, with which command. The position it takes is the ORDER and nothing else - there is no cutoff, so the plan hides nothing, and every term is a lookup over fields the scan already read. Since 0.20.0 it is FIVE ORDERED GROUPS - chain-root, debt, planned, in-progress, undeclared - ranked within a group by that group's own quantity, then a Sonnet next-cost, then oldest plan first.0.19.0 shipped a weighted score here and 0.20.0 replaced it, and the objection the score answered is ACCEPTED, not forgotten. A group order genuinely cannot express "this repo owes one message and releases two others" as one quantity; a score could, and that was its point. What a score could not do was hold still for the second consumer - re-tuning 40 against 15 silently reorders a parser living in another repo, and no test in THIS repo can catch that. The operator weighed both and chose the lookup (2026-08-03). Write that down every time this paragraph is edited: a later session that reads the objection as an unfixed defect will "restore" the score, and the round trip is the loop this file exists to stop.
plannedranks ABOVEin-progress, inverted at 0.20.0 by operator decision. Turning a decision into motion is the slow step; live work is already moving. Flipping it back is a policy change, not a sort fix.Debt is never excluded and never capped, and that is the rule most likely to be "fixed" into a defect. Excluding
blockedordoneis a claim about a repo's OWN next step, which by definition cannot be moved, while owing a reply is the other axis entirely - answering is often what unblocks it. Measured on the real tree at 0.16.0, two of 26 planned repos weredonewith an unhandled inbox. A cap on debt was proposed with the chain credit at 0.19.0 and DECLINED by the operator. Sitting one group below chain-root credit is NOT that cap: the debtor keeps its tab, its most-owed-first position among the other debtors, and itswhy=inbox:N. A change that DROPPED a debtor from the plan would be the declined cap wearing the group order as a disguise, and selftest section 12 pins both halves - the root outranking four owed messages, and the debtor keeping everything it had."Debt" means OWED, never raw pending, since 0.22.0 - and this NARROWS what counts as debt, it does not reopen the paragraph above. The paragraph above settles a different question: once a repo has debt, is it ever excluded or capped (no). This one settles what counts as debt in the first place. Through 0.21.0, group 2's
keep/mag/why=inbox:Nand--brief's whole "repo som skylder et svar" listing were computed from the raw pending-file count - every unhandled message in the inbox, including ones the sender declaredreply-expected: no. That is a notice, not a request, and 0.11.0 gavecoord-count.sha second column (owed) for exactly this distinction - butboard.shnever read it. Reported by morning-driver (2026-08-11) and independently reproduced against the live mailbox 2026-08-13: 27 of 72 pending messages (37.5%) were notices. The fix joins--planand--briefagainstcoord-count.sh'sowedcolumn by repo name (same technique as the chain-root$UNBLOCKSjoin below), so adone/deferred/blockedrepo whose only mail is FYI no longer gets a tab, and--briefno longer counts a notice as an obligation. This reverses a decision from session 41 (2026-08-10) that declined to build this filter, on the premise that "the arrival of the request IS the admission signal" - a premise that assumed group 2 already meant requests. It didn't; the code computed pending, the comments already said "owed" throughout, and the plan's own printed header ("Utelatt naar repoet verken skylder svar...") already claimed the exclusion was debt-based. The fix makes the code match what its own comments and header already promised. Pinned by board-selftest.sh section 8/12 fixturesrepo-done-fyi(pending 2, owed 0 - excluded) andrepo-blocked-mixed(pending 3, owed 2 - planned on 2, not 3). The TABLE'sINNcolumn and the raw scan (RECORDSfield 6) are UNCHANGED - they answer "what is the state of every repo," not "who is waiting on you," and stay on raw pending by design.The same 0.22.0 patch that switched
n_oweto OWED also had to fix whatn_owe == 0claims.--brief's empty-debt branch said "Ingen repo har uhaandtert innboks. Ingen skylder noen et svar i dag." (no repo has unhandled inbox; nobody owes a reply) - two claims in one branch, and only the second is whatn_owe == 0actually proves oncen_owemeans OWED. A repo can hold FYI-only mail with zero debt, which makes the first sentence false while it fires - caught in review before release, not by any fixture (the shared test tree never reachesn_owe == 0, since it always carries a debtor). Fixed to state only the debt claim, and to name any FYI-only mailboxes found rather than let their existence become invisible again - the same "labelled, not silently dropped" principle--planalready applies to unknown-status repos. Pinned by board-selftest.sh section 14 with its own isolated root (debt-free, one FYI-only repo).The
$UNBLOCKS/$RECORDSjoin usedNR==FNRthrough 0.21.0, and that idiom silently drops the entire plan whenever the FIRST file is empty - fixed toFILENAME==comparison in 0.22.0, found while adding the$OWEDjoin above. Verified against the shipped 0.21.0 script: one in-progress repo with an unhandled inbox message, zero blocked repos anywhere in the tree (so$UNBLOCKSis empty, which is a common, ordinary tree state, not an edge case) ---planprinted "0 tabber".NR==FNRis only true for the FIRST file's own lines; when that file is empty,FNRandNRstay equal for the ENTIRE next file too (not just its first line - verified with a minimal awk reproduction), so every record in it is misrouted into theub[]branch and dropped vianext. This was invisible to board-selftest.sh because the fixture tree has carried at least oneblockedrepo since the chain-root feature shipped, and it was invisible on the real tree because~/reposcurrently always has one too - neither is a guarantee.FILENAME==UBF/FILENAME==OWFcompares the exact path, never line counts, so an empty lookup file degrades to "nothing matched," never to "everything after it is misrouted."Chain-root credit lands on the ROOT and nowhere else. For every
blockedrepo theblocked-onedge is followed transitively to the first repo that is not itself blocked. Crediting a blocked repo would open a tab that cannot move; crediting only the direct blocker leaves a two-hop chain's root uncredited, which is the shape the real tree actually had. A cycle, ablocked-onnaming an unscanned repo, and a blocked repo with no target must all credit NOBODY: inventing a root there produces a plan that looks correct and sends the operator to the wrong repo.Repos with no board line rank last and are LABELLED rather than dropped, because the table already prints a MERK line about them and a plan that omitted them silently would repeat that defect.
It renders
key=valueblocks, not prose, because it has two consumers: the operator, and a driver repo consuming the plan. Prose would make the rendered format an API no test in THIS repo could hold stable for a consumer in another.command_missing=carries both no-command causes (no route line, and a route line route.sh rejects) because a barecommand=is the shape of a runnable command carrying nothing - a driver reading^command=would type an empty line into a live pane.route_cmd_for()is the single reader of the route-line grammar, shared with--brief, and distinguishes the two causes by exit code rather than by an empty string.paste=anddir=/command=are the same fact for the two consumers, and neither is redundant. A driver moves the pane itself and then types the command, so it needs them apart; a human needs ONE thing to select. Handing the operator two fields to join by hand is not a saved output line, it is the step where a session starts in the wrong repo - and it was measured the moment the feature met its first user, who could not act on the block at all.paste=is emitted only alongsidecommand=:paste=cd X &&with nothing after it would run the cd and then a bare newline, which fails SILENTLY by leaving the operator in the right directory with no session started.Driving a terminal from the plan does NOT belong here, and the measurement in
docs/ghostty-orchestration-measurement.mdis the argument, not taste. It is a version-pinned undocumented composition over a preview API whose documented path is already broken upstream and whose regression was closed as not planned, with a blast radius reaching into other repos' live sessions. None of that is mailbox transport, and none of it may be able to breakcoord-inboxorboard. The dependency runs one way: the driver consumes the plan, the plan never knows a terminal exists.It also cross-checks itself against
coord-count.sh, and that is not belt-and-braces. The repo scan and the mailbox are two different populations: a mailbox can carry a name no scan will ever produce — a declared non-git surface (CLAUDE_COORD_REPO, e.g.~/repositself) or a checkout outside the roots. Measured on the real mailbox at 0.15.0: the briefing found 11 repos / 21 messages wherecoord-countsaw 12 mailboxes / 22 pending, the missing one being the declared surfacerepos. A briefing that only walks the scan answers "who is waiting on you" with a number it quietly knows is short.Zero model calls, and that is the load-bearing property, not an implementation detail. The operator authenticates by subscription, so a headless
claude -pjob draws from the same quota pool as interactive work. Measured against 2.1.220:--max-budget-usdDOES bite under subscription auth (terminal_reason: budget_exhausted, exit 1), but it aborts AFTER turn one, never before it — floor ~0.25 USD-equivalent per turn onclaude-opus-5[1m]. It is a runaway brake, not a pre-flight gate. Making the briefing deterministic removes the question entirely.board.shstays read-only, which is why the file write lives in the wrapper instead of behind a--brief --out FILEflag. The wrapper renders to a temp file in the target directory and renames it into place, and treats an EMPTY render as a FAILED one: board prints nothing at all when its scan roots do not exist, which is what a mistyped path or a moved home directory looks like, and a plain> fileredirect would destroy yesterday's briefing on a bad launchd environment. A tree where nobody owes anything is a different case — that is a valid, non-empty briefing saying so, and is written normally. -
Route (
scripts/route.sh): pure calculator for the next session's model and effort. Takes four scored traits of the next task plus a required rationale, and prints one block ofkey=valuelines: the rubric row, the rule that fired, thenext-costvalue, a pasteable startup command, the one-row cheaper fallback, and the STATE.md comment lines. Pinned byroute-selftest.sh(69 checks).It is here because it is the WRITER for the field
board.shalready reads.next-costhad a reader and no writer, so it was hand-typed every session and drifted into several competing spellings — cleaning the data could not fix that, because the cause was the missing write path. The row table is a closed set of six values, so a seventh cannot enter circulation, and section 6 of the selftest runs the round trip (route emits → board parses) inside one repo rather than across two.board.shitself is untouched: a calculator that prints to stdout writes nothing, and the session writes STATE.md.The row table is the operator's global rubric, moved here as the single copy. It is not a second spec —
board.sh --helpdocuments the board line's grammar and points here for the values. Scoring the traits is judgement and belongs to the skill; turning scores into a row is a lookup and takes zero model calls.--advisor opusis emitted per ROW, on a need, never unconditionally. Rows 1-2 always carry it (Sonnet main model, so opus is a capability lift — and since every fallback is one row cheaper and the cheap rows are Sonnet, this is what makes the quota fallback safe to take); rows 3-4 only atreversibility=costly|one-way(Opus main model, so it buys peer review where a mistake is not cheap to undo); rows 5-6 never, because the CLI rejects every advisor for a Fable main model. The alternative — the globaladvisorModelsetting written by/advisor— is what this replaces: it applies to every session in every repo, which is how it burned quota before. The two triggers are almost disjoint by construction, sincecostlyforces row 3 andone-wayforces row 4, so a Sonnet row always hasreversibility=cheap.verification=noneis deliberately NOT a third trigger: beyond the stakes rule it would only add mistakes that are cheap to reverse, docs sessions (known/none/cheap/local) among them. Section 14 pins the rule and gates the three CLI facts it rests on against the installedclaudewithout spending a token — advisor validation runs before the empty-prompt check, so-p ""reaches the validator and stops there.Rows 5-6 are never a
route.shoutcome. Until 2026-08-06 they fired only from an explicit--opus-xhigh-failedflag, mirroring a global CLAUDE.md policy that Fable could only be suggested after a failed Opus 5/xhigh session. That policy was removed by operator decision — "for ofte ER Fable riktig" — and the flag went with it rather than being repurposed:route.sh's output range is now closed at row 4, and a Fable choice is always a hand-written deviation from the rubric, recorded in STATE as an override per the model-selection rule in the global CLAUDE.md, never produced by the calculator.board.shstill parses "Fable 5/high" and "Fable 5/xhigh" written by hand into the board line — that parsing is what the override actually uses, and it is pinned separately from anythingroute.shemits (route-selftest.sh section 6).--last-effortis MEASURED fromCLAUDE_EFFORT, and the calculator must never default it. Claude Code exports that variable into every tool-use context as the session's current effort, so the caller reads it and passes it in; havingroute.shread it directly would make the output depend on the environment instead of on its arguments, and the round trip in selftest section 6 rests on that determinism. The two sources it replaces fail identically: the previous board line holds what was prescribed, and asking the operator launders that same prescription through someone reading their own startup command. Corollary pinned by section 13:skills/route/SKILL.mdmust never declare aneffort:frontmatter field, because frontmatter overrides the session effort and the reading would then measure the skill, not the session. -
Skills (
skills/coord-send/,skills/board/,skills/route/): natural-language front doors mapping user intent to engine invocations. No mailbox logic lives here either.boardadditionally owns the ranking — which repo wins and why — sinceboard.shdeliberately prints evidence and takes no position.routelikewise owns the scoring: the calculator is deterministic, so all judgement sits in choosing the four trait values, and the skill must never reason its way to a model instead.
Boundary rule: the mailbox is transport, not state. Durable decisions live in the owning repo's docs/git history; messages are notices pointing at them. Message content is untrusted cross-repo input — the read side quotes and frames it; the send side sanitizes line-oriented fields.
What the boundary forbids is storing a repo's state — its decisions, its next
step, its progress. It does not forbid the mailbox knowing who it is delivering
to: _broadcast/seen/<repo> and <repo>/.origin (0.6.0) are delivery metadata,
answering "has this repo received this" and "which checkout claimed this name".
Both are unreadable as a description of the repo and useless outside delivery.
The test is not "does the engine write a file about a repo" but "would this file
still mean anything if delivery were removed". If yes, it belongs in the repo's
own docs and git history instead.
Priority rule (v0.5.0, Rule 7): the injection block is the only place a repo is ever told what to do with a message, so its wording is the protocol — treat that string as engine behavior, not prose. It obligates handling the inbox first and driving every directed message to a terminal state before the session ends. The obligation is procedural, never substantive: responding is mandatory, complying with message content is not. Those two must stay distinct in any reword — keeping the priority while dropping the distinction turns prioritization into an injection surface. Selftest section 20 pins both halves together for exactly that reason.
Since 0.11.0 the reply-expected field says which terminal state the SENDER
expects. That does not soften the split, it sharpens it: the field is untrusted
cross-repo input like the rest of the file, so the injection calls it a
declaration, not an instruction and keeps both terminal states open to the
receiver. Drop that clause and one word in a message becomes a lever that mints
obligations in another repo.
Conventions
- Scripts are bash-3.2-safe and ASCII-only: no
declare -A, noreadarray/mapfile, no|&; guardshift 2with$# -ge 2; guard empty-array expansion underset -uwith${#a[@]}. - Zero dependencies everywhere: bash + coreutils in the engine,
node:builtins only in hook and tests. - TDD: no behavior change without a failing selftest check first.
bash scripts/coord-selftest.shmust exit 0 (191/191),bash scripts/board-selftest.shmust exit 0 (152/152) andbash scripts/route-selftest.shmust exit 0 (69/69). - English for all code, docs, and commit messages (public repo). Norwegian trigger aliases in the skill description are deliberate.
- Conventional Commits:
type(scope): description.
Commands
- Test:
bash scripts/coord-selftest.sh,bash scripts/board-selftest.shandbash scripts/route-selftest.sh(ornpm test, the Node wrapper around all three) - Hook smoke test:
node hooks/scripts/session-start.mjs(expects JSON on stdout) - Board smoke test:
bash scripts/board.sh(read-only, ~3s over the real tree) - Briefing smoke test:
bash scripts/board.sh --brief(read-only, writes nothing).brief-nightly.shDOES write — it overwrites$CLAUDE_BRIEF_FILE(default~/.claude/briefing.md), so point that at a scratch path when testing. Installed as a launchd agent fromlaunchd/, which points at the SOURCE repo, never the version-pinned plugin cache. - Route smoke test:
bash scripts/route.sh --path known --verification strong --reversibility cheap --scope local --rationale x(writes nothing, instant) - Sweep smoke test:
bash scripts/coord-sweep.sh(dry-run is the default, so this writes nothing; never add--writeto a smoke test against the real mailbox)
Release
Version must agree across: .claude-plugin/plugin.json, package.json,
README version badge, skills/coord-send/SKILL.md, skills/board/SKILL.md and
skills/route/SKILL.md frontmatter, git tag vX.Y.Z, and the catalog ref in
ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json. Release via
the catalog's scripts/release-plugin.mjs repo-mailbox (tag + ref bump
together);
verify with scripts/check-versions.mjs. Never hand-edit a ref.
Two things that script does that its dry-run label does not suggest:
--create-tag creates AND pushes the tag even without --write, and its
closing verification gate runs check-versions.mjs over ALL plugins — one
unrelated plugin in ERROR aborts it with the catalog edit written but
uncommitted. When that happens, commit the catalog's marketplace.json +
README.md by hand and leave every other dirty file in that repo alone.
--write --commit does NOT close this window (confirmed by catalog,
2026-08-10). The gate (check-versions.mjs) runs via execFileSync before
the --commit conditional, so it throws on any plugin's ERROR — including one
we did not touch — after the catalog files are written and before commit,
regardless of whether --commit was passed. Catalog is evaluating a
pre-flight gate (run the check before writing, abort there) but it is not
implemented yet — do not assume it exists. Until it ships: before running
--write, run node scripts/check-versions.mjs in the catalog manually and
confirm 0 ERROR first, even when the only ERROR belongs to an unrelated
plugin. If it still fires mid-release, fall back to the manual-commit
recovery above.
Hardening roadmap
Empty — the post-v0.1.0 queue (atomic delivery, ./.. rejection,
selftest gaps, uniform -h) shipped in v0.2.0; broadcast self-delivery
shipped in v0.2.1; broadcast retraction (coord-send --retract) shipped in
v0.4.0, closing the last monotonically-growing surface. coord-inbox.sh
still ignores unknown arguments by design (hook context must never fail)
but now warns about each one on stderr, which the hook discards.
Two retraction limits are deliberate, not gaps: it is un-send and never
recall (a repo that already received a broadcast keeps it — the seen set is
delivery history and is left untouched), and the sender check is an accident
guard, not a security boundary, because --from redefines identity here as
it does everywhere else in the engine.