repo-mailbox/CLAUDE.md
Kjell Tore Guttormsen 61aebad748 fix(board): --brief's zero-debt branch no longer claims zero pending mail
Found in review of the previous commit, before catalog tags 0.22.0: once
n_owe counts OWED repos rather than raw pending, its ==0 branch could fire
while a repo still held FYI-only mail, making "Ingen repo har uhaandtert
innboks" false at the exact moment it printed. Fixed to state only the
debt claim, and to name any FYI-only mailboxes found instead of letting
their existence go unmentioned. No fixture in the shared test tree ever
reached n_owe==0 (it always carries a debtor), so this needed its own
isolated-root fixture to pin.

Still part of the 0.22.0 release -- amends that changelog entry rather
than bumping again, since no tag exists yet.

board-selftest.sh: 150 -> 152 checks.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGWMPskXBsTjMrrQ2GofFx
2026-08-13 21:07:39 +02:00

27 KiB

repo-mailbox

Renamed from coord in v0.3.0. The plugin/repo is repo-mailbox; the CLI (coord-send.sh, coord-inbox.sh, coord-done.sh), the mailbox root (~/.claude/coord/) and CLAUDE_COORD_DIR deliberately kept their names — they are the transport protocol, not the product.

Context

Local inter-repo coordination mailbox for Claude Code, packaged as a marketplace plugin. Three components, one boundary:

  • Engine (scripts/*.sh): bash owns all mailbox semantics — filename grammar, frontmatter, delivery, archiving, the seen set. coord-send.sh writes, coord-inbox.sh reads (formatted for context injection), coord-done.sh archives, coord-count.sh counts without delivering, coord-sweep.sh closes the aged FYI backlog machine-wide. Everything is pinned by coord-selftest.sh (191 checks, throwaway mailbox via CLAUDE_COORD_DIR).

    coord-sweep.sh is the only path that closes a message with no human in the loop, and every constraint on it follows from that. It may close exactly one mechanically decidable class - reply-expected: no, older than the grace window - because a message that owes a reply can only be answered by a session in the repo that owes it. Dry-run is the default, inverted from the rest of the engine, since this is the one script that destroys pending state. It closes through coord-done.sh --repo rather than moving files, so the archive layout and the _broadcast refusal stay in one place. And it logs every closure with sender and subject, because directed messages have no seen-tracking: the sweep genuinely cannot tell "seen and ignored" from "never delivered", so a notice can be closed unread and the log is the only record that it existed. Widening the class, defaulting to --write, or dropping the log each independently turn this from a bounded cleanup into silent data loss.

    coord-count.sh prints TWO integers per mailbox (<name>\t<pending>\t<debt>), and the first must stay pending: board.sh counts the same inbox files itself, so a debt-only count would put two different numbers under one name. Debt is read from reply-expected in the FRONTMATTER BLOCK ONLY - a body line is untrusted input and must not be able to silence a debt - and an absent field means a reply IS owed, because every message written before 0.11.0 lacks it.

    Reading is delivering — counting is not. coord-inbox.sh records a broadcast as seen once it has printed it, so it can never be used to survey other repos: doing so would consume each one's backlog silently, and the seen set is delivery history that retraction deliberately leaves alone. coord-count.sh exists for every "what is pending" question and writes nothing at all. Any future read-shaped feature belongs there, not in the read path.

  • Hook (hooks/scripts/session-start.mjs): thin zero-dependency Node wrapper (marketplace convention: hooks are .mjs) that calls coord-inbox.sh and emits the hookSpecificOutput.additionalContext envelope. No mailbox logic lives here. Always exits 0.

  • Board (scripts/board.sh): cross-repo attention board. Reads STATE.md next-step blocks + board lines, git status, and mailbox pending counts, and prints one line per repo. Read-only by construction: it writes to no repo, no STATE.md and no mailbox. Pinned by board-selftest.sh (152 checks).

    It lives here because the mailbox is one of its three inputs, and it carries the same axis distinction the mailbox does. A pending count means others are waiting on this repo; who a repo waits on comes only from its board line, because the message format has no reply-to field. Enforcing that in one of two repos would not be enforcing it. The operator invokes both board and coord-send exclusively through their Skill front doors, never a personal terminal alias — a claim this file carried until 0.12.1 (~/.claude/scripts/board.sh is a deployed copy the operator's board() function points at) did not survive inspection: no such file ever existed, and route.sh had no deployed copy either. Only the five coord-*.sh scripts were ever deployed there, and their one measured effect was an accidental fallback target for Claude sessions' own Bash tool calls (the bug 0.12.1 fixed) — once that fallback was gone they had no remaining function and were deleted. board.sh --brief is a second RENDERING of that scan, never a second scan, and brief-nightly.sh is the only writer in that path. The briefing answers the narrower question an unattended job can answer without judgement: which repos have an unhandled inbox, what their next step says in full, and the exact command to start a session in each. It prints NESTE uncut because the 38-character cut is the table column's property, not the record's — the value used to be truncated at record-build time, which left the cut string as the only copy. Each command is derived by CALLING route.sh with that repo's own four traits; next-cost alone cannot yield it, since the advisor flag is a property of the ROW and two rows can share a model/effort pair while differing on it. A repo with no route line is told so rather than handed a guess, because a guessed command reads as authoritative.

    board.sh --plan is the THIRD rendering, and the only one that takes a position. It answers which repos to open a tab for today, in what order, with which command. The position it takes is the ORDER and nothing else - there is no cutoff, so the plan hides nothing, and every term is a lookup over fields the scan already read. Since 0.20.0 it is FIVE ORDERED GROUPS - chain-root, debt, planned, in-progress, undeclared - ranked within a group by that group's own quantity, then a Sonnet next-cost, then oldest plan first.

    0.19.0 shipped a weighted score here and 0.20.0 replaced it, and the objection the score answered is ACCEPTED, not forgotten. A group order genuinely cannot express "this repo owes one message and releases two others" as one quantity; a score could, and that was its point. What a score could not do was hold still for the second consumer - re-tuning 40 against 15 silently reorders a parser living in another repo, and no test in THIS repo can catch that. The operator weighed both and chose the lookup (2026-08-03). Write that down every time this paragraph is edited: a later session that reads the objection as an unfixed defect will "restore" the score, and the round trip is the loop this file exists to stop.

    planned ranks ABOVE in-progress, inverted at 0.20.0 by operator decision. Turning a decision into motion is the slow step; live work is already moving. Flipping it back is a policy change, not a sort fix.

    Debt is never excluded and never capped, and that is the rule most likely to be "fixed" into a defect. Excluding blocked or done is a claim about a repo's OWN next step, which by definition cannot be moved, while owing a reply is the other axis entirely - answering is often what unblocks it. Measured on the real tree at 0.16.0, two of 26 planned repos were done with an unhandled inbox. A cap on debt was proposed with the chain credit at 0.19.0 and DECLINED by the operator. Sitting one group below chain-root credit is NOT that cap: the debtor keeps its tab, its most-owed-first position among the other debtors, and its why=inbox:N. A change that DROPPED a debtor from the plan would be the declined cap wearing the group order as a disguise, and selftest section 12 pins both halves - the root outranking four owed messages, and the debtor keeping everything it had.

    "Debt" means OWED, never raw pending, since 0.22.0 - and this NARROWS what counts as debt, it does not reopen the paragraph above. The paragraph above settles a different question: once a repo has debt, is it ever excluded or capped (no). This one settles what counts as debt in the first place. Through 0.21.0, group 2's keep/mag/why=inbox:N and --brief's whole "repo som skylder et svar" listing were computed from the raw pending-file count - every unhandled message in the inbox, including ones the sender declared reply-expected: no. That is a notice, not a request, and 0.11.0 gave coord-count.sh a second column (owed) for exactly this distinction - but board.sh never read it. Reported by morning-driver (2026-08-11) and independently reproduced against the live mailbox 2026-08-13: 27 of 72 pending messages (37.5%) were notices. The fix joins --plan and --brief against coord-count.sh's owed column by repo name (same technique as the chain-root $UNBLOCKS join below), so a done/deferred/blocked repo whose only mail is FYI no longer gets a tab, and --brief no longer counts a notice as an obligation. This reverses a decision from session 41 (2026-08-10) that declined to build this filter, on the premise that "the arrival of the request IS the admission signal" - a premise that assumed group 2 already meant requests. It didn't; the code computed pending, the comments already said "owed" throughout, and the plan's own printed header ("Utelatt naar repoet verken skylder svar...") already claimed the exclusion was debt-based. The fix makes the code match what its own comments and header already promised. Pinned by board-selftest.sh section 8/12 fixtures repo-done-fyi (pending 2, owed 0 - excluded) and repo-blocked-mixed (pending 3, owed 2 - planned on 2, not 3). The TABLE's INN column and the raw scan (RECORDS field 6) are UNCHANGED - they answer "what is the state of every repo," not "who is waiting on you," and stay on raw pending by design.

    The same 0.22.0 patch that switched n_owe to OWED also had to fix what n_owe == 0 claims. --brief's empty-debt branch said "Ingen repo har uhaandtert innboks. Ingen skylder noen et svar i dag." (no repo has unhandled inbox; nobody owes a reply) - two claims in one branch, and only the second is what n_owe == 0 actually proves once n_owe means OWED. A repo can hold FYI-only mail with zero debt, which makes the first sentence false while it fires - caught in review before release, not by any fixture (the shared test tree never reaches n_owe == 0, since it always carries a debtor). Fixed to state only the debt claim, and to name any FYI-only mailboxes found rather than let their existence become invisible again - the same "labelled, not silently dropped" principle --plan already applies to unknown-status repos. Pinned by board-selftest.sh section 14 with its own isolated root (debt-free, one FYI-only repo).

    The $UNBLOCKS/$RECORDS join used NR==FNR through 0.21.0, and that idiom silently drops the entire plan whenever the FIRST file is empty - fixed to FILENAME== comparison in 0.22.0, found while adding the $OWED join above. Verified against the shipped 0.21.0 script: one in-progress repo with an unhandled inbox message, zero blocked repos anywhere in the tree (so $UNBLOCKS is empty, which is a common, ordinary tree state, not an edge case) - --plan printed "0 tabber". NR==FNR is only true for the FIRST file's own lines; when that file is empty, FNR and NR stay equal for the ENTIRE next file too (not just its first line - verified with a minimal awk reproduction), so every record in it is misrouted into the ub[] branch and dropped via next. This was invisible to board-selftest.sh because the fixture tree has carried at least one blocked repo since the chain-root feature shipped, and it was invisible on the real tree because ~/repos currently always has one too - neither is a guarantee. FILENAME==UBF/FILENAME==OWF compares the exact path, never line counts, so an empty lookup file degrades to "nothing matched," never to "everything after it is misrouted."

    Chain-root credit lands on the ROOT and nowhere else. For every blocked repo the blocked-on edge is followed transitively to the first repo that is not itself blocked. Crediting a blocked repo would open a tab that cannot move; crediting only the direct blocker leaves a two-hop chain's root uncredited, which is the shape the real tree actually had. A cycle, a blocked-on naming an unscanned repo, and a blocked repo with no target must all credit NOBODY: inventing a root there produces a plan that looks correct and sends the operator to the wrong repo.

    Repos with no board line rank last and are LABELLED rather than dropped, because the table already prints a MERK line about them and a plan that omitted them silently would repeat that defect.

    It renders key=value blocks, not prose, because it has two consumers: the operator, and a driver repo consuming the plan. Prose would make the rendered format an API no test in THIS repo could hold stable for a consumer in another. command_missing= carries both no-command causes (no route line, and a route line route.sh rejects) because a bare command= is the shape of a runnable command carrying nothing - a driver reading ^command= would type an empty line into a live pane. route_cmd_for() is the single reader of the route-line grammar, shared with --brief, and distinguishes the two causes by exit code rather than by an empty string.

    paste= and dir=/command= are the same fact for the two consumers, and neither is redundant. A driver moves the pane itself and then types the command, so it needs them apart; a human needs ONE thing to select. Handing the operator two fields to join by hand is not a saved output line, it is the step where a session starts in the wrong repo - and it was measured the moment the feature met its first user, who could not act on the block at all. paste= is emitted only alongside command=: paste=cd X && with nothing after it would run the cd and then a bare newline, which fails SILENTLY by leaving the operator in the right directory with no session started.

    Driving a terminal from the plan does NOT belong here, and the measurement in docs/ghostty-orchestration-measurement.md is the argument, not taste. It is a version-pinned undocumented composition over a preview API whose documented path is already broken upstream and whose regression was closed as not planned, with a blast radius reaching into other repos' live sessions. None of that is mailbox transport, and none of it may be able to break coord-inbox or board. The dependency runs one way: the driver consumes the plan, the plan never knows a terminal exists.

    It also cross-checks itself against coord-count.sh, and that is not belt-and-braces. The repo scan and the mailbox are two different populations: a mailbox can carry a name no scan will ever produce — a declared non-git surface (CLAUDE_COORD_REPO, e.g. ~/repos itself) or a checkout outside the roots. Measured on the real mailbox at 0.15.0: the briefing found 11 repos / 21 messages where coord-count saw 12 mailboxes / 22 pending, the missing one being the declared surface repos. A briefing that only walks the scan answers "who is waiting on you" with a number it quietly knows is short.

    Zero model calls, and that is the load-bearing property, not an implementation detail. The operator authenticates by subscription, so a headless claude -p job draws from the same quota pool as interactive work. Measured against 2.1.220: --max-budget-usd DOES bite under subscription auth (terminal_reason: budget_exhausted, exit 1), but it aborts AFTER turn one, never before it — floor ~0.25 USD-equivalent per turn on claude-opus-5[1m]. It is a runaway brake, not a pre-flight gate. Making the briefing deterministic removes the question entirely.

    board.sh stays read-only, which is why the file write lives in the wrapper instead of behind a --brief --out FILE flag. The wrapper renders to a temp file in the target directory and renames it into place, and treats an EMPTY render as a FAILED one: board prints nothing at all when its scan roots do not exist, which is what a mistyped path or a moved home directory looks like, and a plain > file redirect would destroy yesterday's briefing on a bad launchd environment. A tree where nobody owes anything is a different case — that is a valid, non-empty briefing saying so, and is written normally.

  • Route (scripts/route.sh): pure calculator for the next session's model and effort. Takes four scored traits of the next task plus a required rationale, and prints one block of key=value lines: the rubric row, the rule that fired, the next-cost value, a pasteable startup command, the one-row cheaper fallback, and the STATE.md comment lines. Pinned by route-selftest.sh (69 checks).

    It is here because it is the WRITER for the field board.sh already reads. next-cost had a reader and no writer, so it was hand-typed every session and drifted into several competing spellings — cleaning the data could not fix that, because the cause was the missing write path. The row table is a closed set of six values, so a seventh cannot enter circulation, and section 6 of the selftest runs the round trip (route emits → board parses) inside one repo rather than across two. board.sh itself is untouched: a calculator that prints to stdout writes nothing, and the session writes STATE.md.

    The row table is the operator's global rubric, moved here as the single copy. It is not a second spec — board.sh --help documents the board line's grammar and points here for the values. Scoring the traits is judgement and belongs to the skill; turning scores into a row is a lookup and takes zero model calls.

    --advisor opus is emitted per ROW, on a need, never unconditionally. Rows 1-2 always carry it (Sonnet main model, so opus is a capability lift — and since every fallback is one row cheaper and the cheap rows are Sonnet, this is what makes the quota fallback safe to take); rows 3-4 only at reversibility=costly|one-way (Opus main model, so it buys peer review where a mistake is not cheap to undo); rows 5-6 never, because the CLI rejects every advisor for a Fable main model. The alternative — the global advisorModel setting written by /advisor — is what this replaces: it applies to every session in every repo, which is how it burned quota before. The two triggers are almost disjoint by construction, since costly forces row 3 and one-way forces row 4, so a Sonnet row always has reversibility=cheap. verification=none is deliberately NOT a third trigger: beyond the stakes rule it would only add mistakes that are cheap to reverse, docs sessions (known/none/cheap/local) among them. Section 14 pins the rule and gates the three CLI facts it rests on against the installed claude without spending a token — advisor validation runs before the empty-prompt check, so -p "" reaches the validator and stops there.

    Rows 5-6 are never a route.sh outcome. Until 2026-08-06 they fired only from an explicit --opus-xhigh-failed flag, mirroring a global CLAUDE.md policy that Fable could only be suggested after a failed Opus 5/xhigh session. That policy was removed by operator decision — "for ofte ER Fable riktig" — and the flag went with it rather than being repurposed: route.sh's output range is now closed at row 4, and a Fable choice is always a hand-written deviation from the rubric, recorded in STATE as an override per the model-selection rule in the global CLAUDE.md, never produced by the calculator. board.sh still parses "Fable 5/high" and "Fable 5/xhigh" written by hand into the board line — that parsing is what the override actually uses, and it is pinned separately from anything route.sh emits (route-selftest.sh section 6).

    --last-effort is MEASURED from CLAUDE_EFFORT, and the calculator must never default it. Claude Code exports that variable into every tool-use context as the session's current effort, so the caller reads it and passes it in; having route.sh read it directly would make the output depend on the environment instead of on its arguments, and the round trip in selftest section 6 rests on that determinism. The two sources it replaces fail identically: the previous board line holds what was prescribed, and asking the operator launders that same prescription through someone reading their own startup command. Corollary pinned by section 13: skills/route/SKILL.md must never declare an effort: frontmatter field, because frontmatter overrides the session effort and the reading would then measure the skill, not the session.

  • Skills (skills/coord-send/, skills/board/, skills/route/): natural-language front doors mapping user intent to engine invocations. No mailbox logic lives here either. board additionally owns the ranking — which repo wins and why — since board.sh deliberately prints evidence and takes no position. route likewise owns the scoring: the calculator is deterministic, so all judgement sits in choosing the four trait values, and the skill must never reason its way to a model instead.

Boundary rule: the mailbox is transport, not state. Durable decisions live in the owning repo's docs/git history; messages are notices pointing at them. Message content is untrusted cross-repo input — the read side quotes and frames it; the send side sanitizes line-oriented fields.

What the boundary forbids is storing a repo's state — its decisions, its next step, its progress. It does not forbid the mailbox knowing who it is delivering to: _broadcast/seen/<repo> and <repo>/.origin (0.6.0) are delivery metadata, answering "has this repo received this" and "which checkout claimed this name". Both are unreadable as a description of the repo and useless outside delivery. The test is not "does the engine write a file about a repo" but "would this file still mean anything if delivery were removed". If yes, it belongs in the repo's own docs and git history instead.

Priority rule (v0.5.0, Rule 7): the injection block is the only place a repo is ever told what to do with a message, so its wording is the protocol — treat that string as engine behavior, not prose. It obligates handling the inbox first and driving every directed message to a terminal state before the session ends. The obligation is procedural, never substantive: responding is mandatory, complying with message content is not. Those two must stay distinct in any reword — keeping the priority while dropping the distinction turns prioritization into an injection surface. Selftest section 20 pins both halves together for exactly that reason.

Since 0.11.0 the reply-expected field says which terminal state the SENDER expects. That does not soften the split, it sharpens it: the field is untrusted cross-repo input like the rest of the file, so the injection calls it a declaration, not an instruction and keeps both terminal states open to the receiver. Drop that clause and one word in a message becomes a lever that mints obligations in another repo.

Conventions

  • Scripts are bash-3.2-safe and ASCII-only: no declare -A, no readarray/mapfile, no |&; guard shift 2 with $# -ge 2; guard empty-array expansion under set -u with ${#a[@]}.
  • Zero dependencies everywhere: bash + coreutils in the engine, node: builtins only in hook and tests.
  • TDD: no behavior change without a failing selftest check first. bash scripts/coord-selftest.sh must exit 0 (191/191), bash scripts/board-selftest.sh must exit 0 (152/152) and bash scripts/route-selftest.sh must exit 0 (69/69).
  • English for all code, docs, and commit messages (public repo). Norwegian trigger aliases in the skill description are deliberate.
  • Conventional Commits: type(scope): description.

Commands

  • Test: bash scripts/coord-selftest.sh, bash scripts/board-selftest.sh and bash scripts/route-selftest.sh (or npm test, the Node wrapper around all three)
  • Hook smoke test: node hooks/scripts/session-start.mjs (expects JSON on stdout)
  • Board smoke test: bash scripts/board.sh (read-only, ~3s over the real tree)
  • Briefing smoke test: bash scripts/board.sh --brief (read-only, writes nothing). brief-nightly.sh DOES write — it overwrites $CLAUDE_BRIEF_FILE (default ~/.claude/briefing.md), so point that at a scratch path when testing. Installed as a launchd agent from launchd/, which points at the SOURCE repo, never the version-pinned plugin cache.
  • Route smoke test: bash scripts/route.sh --path known --verification strong --reversibility cheap --scope local --rationale x (writes nothing, instant)
  • Sweep smoke test: bash scripts/coord-sweep.sh (dry-run is the default, so this writes nothing; never add --write to a smoke test against the real mailbox)

Release

Version must agree across: .claude-plugin/plugin.json, package.json, README version badge, skills/coord-send/SKILL.md, skills/board/SKILL.md and skills/route/SKILL.md frontmatter, git tag vX.Y.Z, and the catalog ref in ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json. Release via the catalog's scripts/release-plugin.mjs repo-mailbox (tag + ref bump together); verify with scripts/check-versions.mjs. Never hand-edit a ref.

Two things that script does that its dry-run label does not suggest: --create-tag creates AND pushes the tag even without --write, and its closing verification gate runs check-versions.mjs over ALL plugins — one unrelated plugin in ERROR aborts it with the catalog edit written but uncommitted. When that happens, commit the catalog's marketplace.json + README.md by hand and leave every other dirty file in that repo alone.

--write --commit does NOT close this window (confirmed by catalog, 2026-08-10). The gate (check-versions.mjs) runs via execFileSync before the --commit conditional, so it throws on any plugin's ERROR — including one we did not touch — after the catalog files are written and before commit, regardless of whether --commit was passed. Catalog is evaluating a pre-flight gate (run the check before writing, abort there) but it is not implemented yet — do not assume it exists. Until it ships: before running --write, run node scripts/check-versions.mjs in the catalog manually and confirm 0 ERROR first, even when the only ERROR belongs to an unrelated plugin. If it still fires mid-release, fall back to the manual-commit recovery above.

Hardening roadmap

Empty — the post-v0.1.0 queue (atomic delivery, ./.. rejection, selftest gaps, uniform -h) shipped in v0.2.0; broadcast self-delivery shipped in v0.2.1; broadcast retraction (coord-send --retract) shipped in v0.4.0, closing the last monotonically-growing surface. coord-inbox.sh still ignores unknown arguments by design (hook context must never fail) but now warns about each one on stderr, which the hook discards.

Two retraction limits are deliberate, not gaps: it is un-send and never recall (a repo that already received a broadcast keeps it — the seen set is delivery history and is left untouched), and the sender check is an accident guard, not a security boundary, because --from redefines identity here as it does everywhere else in the engine.