repo-mailbox/CLAUDE.md
Kjell Tore Guttormsen 44fb34ab72 fix(hooks): raise state-line-guard MAX_LINES from 60 to 120
Operator decision 2026-08-14: the STATE.md convention's line limit moved
from ~60 to ~120 (global CLAUDE.md already updated). Re-bases every
selftest fixture and boundary value that encoded 60 as a literal,
including section 8's ratchet fixtures, so they still exercise the
ratchet rather than degenerating into a flat gate at the new threshold.
Re-verified the real-tree justification at 120: 13 files over the limit,
one at 1496 (was 23 over 60, one at 1405 — left as historical record).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194eV8b6BXNv6aKLovP8TP6
2026-08-14 21:35:31 +02:00

32 KiB

repo-mailbox

Renamed from coord in v0.3.0. The plugin/repo is repo-mailbox; the CLI (coord-send.sh, coord-inbox.sh, coord-done.sh), the mailbox root (~/.claude/coord/) and CLAUDE_COORD_DIR deliberately kept their names — they are the transport protocol, not the product.

Context

Local inter-repo coordination mailbox for Claude Code, packaged as a marketplace plugin. Three components, one boundary:

  • Engine (scripts/*.sh): bash owns all mailbox semantics — filename grammar, frontmatter, delivery, archiving, the seen set. coord-send.sh writes, coord-inbox.sh reads (formatted for context injection), coord-done.sh archives, coord-count.sh counts without delivering, coord-sweep.sh closes the aged FYI backlog machine-wide. Everything is pinned by coord-selftest.sh (191 checks, throwaway mailbox via CLAUDE_COORD_DIR).

    coord-sweep.sh is the only path that closes a message with no human in the loop, and every constraint on it follows from that. It may close exactly one mechanically decidable class - reply-expected: no, older than the grace window - because a message that owes a reply can only be answered by a session in the repo that owes it. Dry-run is the default, inverted from the rest of the engine, since this is the one script that destroys pending state. It closes through coord-done.sh --repo rather than moving files, so the archive layout and the _broadcast refusal stay in one place. And it logs every closure with sender and subject, because directed messages have no seen-tracking: the sweep genuinely cannot tell "seen and ignored" from "never delivered", so a notice can be closed unread and the log is the only record that it existed. Widening the class, defaulting to --write, or dropping the log each independently turn this from a bounded cleanup into silent data loss.

    coord-count.sh prints TWO integers per mailbox (<name>\t<pending>\t<debt>), and the first must stay pending: board.sh counts the same inbox files itself, so a debt-only count would put two different numbers under one name. Debt is read from reply-expected in the FRONTMATTER BLOCK ONLY - a body line is untrusted input and must not be able to silence a debt - and an absent field means a reply IS owed, because every message written before 0.11.0 lacks it.

    Reading is delivering — counting is not. coord-inbox.sh records a broadcast as seen once it has printed it, so it can never be used to survey other repos: doing so would consume each one's backlog silently, and the seen set is delivery history that retraction deliberately leaves alone. coord-count.sh exists for every "what is pending" question and writes nothing at all. Any future read-shaped feature belongs there, not in the read path.

  • Hook (hooks/scripts/session-start.mjs): thin zero-dependency Node wrapper (marketplace convention: hooks are .mjs) that calls coord-inbox.sh and emits the hookSpecificOutput.additionalContext envelope. No mailbox logic lives here. Always exits 0.

  • Hook (hooks/scripts/pre-state-line-guard.mjs): a PreToolUse hook on Write|Edit that enforces the STATE.md convention's maks ~120 linjer (global CLAUDE.md; raised from ~60 by operator decision 2026-08-14 — see the dated paragraph below) mechanically. It exists because the prose limit alone failed: a real STATE.md drifted to 155-156 lines before an /insights sweep of 160 sessions noticed, and one trim pass on it increased the line count instead of shrinking it. org-ops dispatched the work order (20260814T144553Z) asking for a PostToolUse hook — that was the wrong event, and the fix is not cosmetic: PostToolUse fires only after the tool has already written the file (confirmed against the official hooks docs, 2026-08-14 — "Can block? No", stderr is shown to the model but the write already landed), so it cannot stop an oversized STATE.md from landing, only nag about it afterward. PreToolUse is the only event that can deny the call before the file is touched, which is what "enforces" has to mean here. Denial is stderr + exit 2, matching llm-security's pre-write-pathguard.mjs — the only other PreToolUse Write|Edit guard in this marketplace — rather than the hookSpecificOutput.permissionDecision JSON form; both block, and matching the sibling convention keeps one idiom for "block a write" instead of two. For Write the projected content is the call's own content; for Edit it is the CURRENT on-disk file (read fresh, since PreToolUse fires before the edit is applied) with old_string replaced by new_string — every occurrence when replace_all is set, otherwise only the first, mirroring what the real Edit tool does. Getting replace_all wrong in either direction is not a hypothetical: a hook that only ever replaced the first occurrence would silently pass a bulk edit that balloons the file, so state-line-guard-selftest.sh (21 checks) pins a fixture where only counting every replace_all occurrence produces the correct denial. Anything the hook cannot project with confidence — a missing file, an old_string that is not present, fields of the wrong type — is left to the real tool, which reports a clearer error than a guess here would; the guard only ever touches files named exactly STATE.md, at any depth, matching the same basename rule the global session-start hook's nearest-STATE-wins search already uses.

    It is a RATCHET against the file's current size, not a flat gate at 60 — found by advisor review before the tag landed, not by the selftest, which had no fixture for it. The first cut compared the projected line count only against MAX_LINES, never against what the file already was, so trimming an oversized STATE.md from, say, 156 to 100 lines — still over 60, but strictly smaller — was denied exactly like growing it would have been. Verified empirically against the real tree (2026-08-14): wc -l ~/repos/*/STATE.md ~/repos/*/*/STATE.md | awk '$1 > 60' found 23 files already over 60 lines, one at 1405. Shipped as a flat gate, this hook would have made most of the machine's STATE.md files un-editable except by a single write landing at <=60 in one shot — backwards for a guard whose whole point is making the trim the /insights finding asked for actually possible. The fix reads the file's current line count for BOTH tool types (previously only Edit read the file at all) and denies only when the projection is over MAX_LINES and larger than that current count: a compliant file still cannot grow past the limit, a brand-new file still cannot be created oversized (current defaults to 0), but an already-oversized file can always be edited toward compliance, one write at a time, without ever making it worse. Section 8 of the selftest pins all four cases: shrink-while-still-over-limit allows, same-size-rewrite allows, grow-an- already-oversized-file still denies, and create-new-oversized-file still denies.

    MAX_LINES raised 60 -> 120, operator decision 2026-08-14 (evening), reported via coord by .claude after the global CLAUDE.md prose was already updated. The ratchet mechanics above are unchanged - only the constant moved, plus every selftest fixture and boundary value that encoded 60 as a literal. Re-verified empirically against the real tree at the new threshold (2026-08-14): wc -l ~/repos/*/STATE.md ~/repos/*/*/STATE.md | awk '$1 > 120' found 13 files already over 120 lines, one at 1496 - the historical 23-files-over-60/one-at-1405 figures above describe the tree as it was at the moment the ratchet bug was found, not the current threshold, and are left as-is rather than rewritten. session-start.mjs's 160-line injection window still covers the new 120-line limit with room to spare, so no change was needed there.

  • Board (scripts/board.sh): cross-repo attention board. Reads STATE.md next-step blocks + board lines, git status, and mailbox pending counts, and prints one line per repo. Read-only by construction: it writes to no repo, no STATE.md and no mailbox. Pinned by board-selftest.sh (152 checks).

    It lives here because the mailbox is one of its three inputs, and it carries the same axis distinction the mailbox does. A pending count means others are waiting on this repo; who a repo waits on comes only from its board line, because the message format has no reply-to field. Enforcing that in one of two repos would not be enforcing it. The operator invokes both board and coord-send exclusively through their Skill front doors, never a personal terminal alias — a claim this file carried until 0.12.1 (~/.claude/scripts/board.sh is a deployed copy the operator's board() function points at) did not survive inspection: no such file ever existed, and route.sh had no deployed copy either. Only the five coord-*.sh scripts were ever deployed there, and their one measured effect was an accidental fallback target for Claude sessions' own Bash tool calls (the bug 0.12.1 fixed) — once that fallback was gone they had no remaining function and were deleted. board.sh --brief is a second RENDERING of that scan, never a second scan, and brief-nightly.sh is the only writer in that path. The briefing answers the narrower question an unattended job can answer without judgement: which repos have an unhandled inbox, what their next step says in full, and the exact command to start a session in each. It prints NESTE uncut because the 38-character cut is the table column's property, not the record's — the value used to be truncated at record-build time, which left the cut string as the only copy. Each command is derived by CALLING route.sh with that repo's own four traits; next-cost alone cannot yield it, since the advisor flag is a property of the ROW and two rows can share a model/effort pair while differing on it. A repo with no route line is told so rather than handed a guess, because a guessed command reads as authoritative.

    board.sh --plan is the THIRD rendering, and the only one that takes a position. It answers which repos to open a tab for today, in what order, with which command. The position it takes is the ORDER and nothing else - there is no cutoff, so the plan hides nothing, and every term is a lookup over fields the scan already read. Since 0.20.0 it is FIVE ORDERED GROUPS - chain-root, debt, planned, in-progress, undeclared - ranked within a group by that group's own quantity, then a Sonnet next-cost, then oldest plan first.

    0.19.0 shipped a weighted score here and 0.20.0 replaced it, and the objection the score answered is ACCEPTED, not forgotten. A group order genuinely cannot express "this repo owes one message and releases two others" as one quantity; a score could, and that was its point. What a score could not do was hold still for the second consumer - re-tuning 40 against 15 silently reorders a parser living in another repo, and no test in THIS repo can catch that. The operator weighed both and chose the lookup (2026-08-03). Write that down every time this paragraph is edited: a later session that reads the objection as an unfixed defect will "restore" the score, and the round trip is the loop this file exists to stop.

    planned ranks ABOVE in-progress, inverted at 0.20.0 by operator decision. Turning a decision into motion is the slow step; live work is already moving. Flipping it back is a policy change, not a sort fix.

    Debt is never excluded and never capped, and that is the rule most likely to be "fixed" into a defect. Excluding blocked or done is a claim about a repo's OWN next step, which by definition cannot be moved, while owing a reply is the other axis entirely - answering is often what unblocks it. Measured on the real tree at 0.16.0, two of 26 planned repos were done with an unhandled inbox. A cap on debt was proposed with the chain credit at 0.19.0 and DECLINED by the operator. Sitting one group below chain-root credit is NOT that cap: the debtor keeps its tab, its most-owed-first position among the other debtors, and its why=inbox:N. A change that DROPPED a debtor from the plan would be the declined cap wearing the group order as a disguise, and selftest section 12 pins both halves - the root outranking four owed messages, and the debtor keeping everything it had.

    "Debt" means OWED, never raw pending, since 0.22.0 - and this NARROWS what counts as debt, it does not reopen the paragraph above. The paragraph above settles a different question: once a repo has debt, is it ever excluded or capped (no). This one settles what counts as debt in the first place. Through 0.21.0, group 2's keep/mag/why=inbox:N and --brief's whole "repo som skylder et svar" listing were computed from the raw pending-file count - every unhandled message in the inbox, including ones the sender declared reply-expected: no. That is a notice, not a request, and 0.11.0 gave coord-count.sh a second column (owed) for exactly this distinction - but board.sh never read it. Reported by morning-driver (2026-08-11) and independently reproduced against the live mailbox 2026-08-13: 27 of 72 pending messages (37.5%) were notices. The fix joins --plan and --brief against coord-count.sh's owed column by repo name (same technique as the chain-root $UNBLOCKS join below), so a done/deferred/blocked repo whose only mail is FYI no longer gets a tab, and --brief no longer counts a notice as an obligation. This reverses a decision from session 41 (2026-08-10) that declined to build this filter, on the premise that "the arrival of the request IS the admission signal" - a premise that assumed group 2 already meant requests. It didn't; the code computed pending, the comments already said "owed" throughout, and the plan's own printed header ("Utelatt naar repoet verken skylder svar...") already claimed the exclusion was debt-based. The fix makes the code match what its own comments and header already promised. Pinned by board-selftest.sh section 8/12 fixtures repo-done-fyi (pending 2, owed 0 - excluded) and repo-blocked-mixed (pending 3, owed 2 - planned on 2, not 3). The TABLE's INN column and the raw scan (RECORDS field 6) are UNCHANGED - they answer "what is the state of every repo," not "who is waiting on you," and stay on raw pending by design.

    The same 0.22.0 patch that switched n_owe to OWED also had to fix what n_owe == 0 claims. --brief's empty-debt branch said "Ingen repo har uhaandtert innboks. Ingen skylder noen et svar i dag." (no repo has unhandled inbox; nobody owes a reply) - two claims in one branch, and only the second is what n_owe == 0 actually proves once n_owe means OWED. A repo can hold FYI-only mail with zero debt, which makes the first sentence false while it fires - caught in review before release, not by any fixture (the shared test tree never reaches n_owe == 0, since it always carries a debtor). Fixed to state only the debt claim, and to name any FYI-only mailboxes found rather than let their existence become invisible again - the same "labelled, not silently dropped" principle --plan already applies to unknown-status repos. Pinned by board-selftest.sh section 14 with its own isolated root (debt-free, one FYI-only repo).

    The $UNBLOCKS/$RECORDS join used NR==FNR through 0.21.0, and that idiom silently drops the entire plan whenever the FIRST file is empty - fixed to FILENAME== comparison in 0.22.0, found while adding the $OWED join above. Verified against the shipped 0.21.0 script: one in-progress repo with an unhandled inbox message, zero blocked repos anywhere in the tree (so $UNBLOCKS is empty, which is a common, ordinary tree state, not an edge case) - --plan printed "0 tabber". NR==FNR is only true for the FIRST file's own lines; when that file is empty, FNR and NR stay equal for the ENTIRE next file too (not just its first line - verified with a minimal awk reproduction), so every record in it is misrouted into the ub[] branch and dropped via next. This was invisible to board-selftest.sh because the fixture tree has carried at least one blocked repo since the chain-root feature shipped, and it was invisible on the real tree because ~/repos currently always has one too - neither is a guarantee. FILENAME==UBF/FILENAME==OWF compares the exact path, never line counts, so an empty lookup file degrades to "nothing matched," never to "everything after it is misrouted."

    Chain-root credit lands on the ROOT and nowhere else. For every blocked repo the blocked-on edge is followed transitively to the first repo that is not itself blocked. Crediting a blocked repo would open a tab that cannot move; crediting only the direct blocker leaves a two-hop chain's root uncredited, which is the shape the real tree actually had. A cycle, a blocked-on naming an unscanned repo, and a blocked repo with no target must all credit NOBODY: inventing a root there produces a plan that looks correct and sends the operator to the wrong repo.

    Repos with no board line rank last and are LABELLED rather than dropped, because the table already prints a MERK line about them and a plan that omitted them silently would repeat that defect.

    It renders key=value blocks, not prose, because it has two consumers: the operator, and a driver repo consuming the plan. Prose would make the rendered format an API no test in THIS repo could hold stable for a consumer in another. command_missing= carries both no-command causes (no route line, and a route line route.sh rejects) because a bare command= is the shape of a runnable command carrying nothing - a driver reading ^command= would type an empty line into a live pane. route_cmd_for() is the single reader of the route-line grammar, shared with --brief, and distinguishes the two causes by exit code rather than by an empty string.

    paste= and dir=/command= are the same fact for the two consumers, and neither is redundant. A driver moves the pane itself and then types the command, so it needs them apart; a human needs ONE thing to select. Handing the operator two fields to join by hand is not a saved output line, it is the step where a session starts in the wrong repo - and it was measured the moment the feature met its first user, who could not act on the block at all. paste= is emitted only alongside command=: paste=cd X && with nothing after it would run the cd and then a bare newline, which fails SILENTLY by leaving the operator in the right directory with no session started.

    Driving a terminal from the plan does NOT belong here, and the measurement in docs/ghostty-orchestration-measurement.md is the argument, not taste. It is a version-pinned undocumented composition over a preview API whose documented path is already broken upstream and whose regression was closed as not planned, with a blast radius reaching into other repos' live sessions. None of that is mailbox transport, and none of it may be able to break coord-inbox or board. The dependency runs one way: the driver consumes the plan, the plan never knows a terminal exists.

    It also cross-checks itself against coord-count.sh, and that is not belt-and-braces. The repo scan and the mailbox are two different populations: a mailbox can carry a name no scan will ever produce — a declared non-git surface (CLAUDE_COORD_REPO, e.g. ~/repos itself) or a checkout outside the roots. Measured on the real mailbox at 0.15.0: the briefing found 11 repos / 21 messages where coord-count saw 12 mailboxes / 22 pending, the missing one being the declared surface repos. A briefing that only walks the scan answers "who is waiting on you" with a number it quietly knows is short.

    Zero model calls, and that is the load-bearing property, not an implementation detail. The operator authenticates by subscription, so a headless claude -p job draws from the same quota pool as interactive work. Measured against 2.1.220: --max-budget-usd DOES bite under subscription auth (terminal_reason: budget_exhausted, exit 1), but it aborts AFTER turn one, never before it — floor ~0.25 USD-equivalent per turn on claude-opus-5[1m]. It is a runaway brake, not a pre-flight gate. Making the briefing deterministic removes the question entirely.

    board.sh stays read-only, which is why the file write lives in the wrapper instead of behind a --brief --out FILE flag. The wrapper renders to a temp file in the target directory and renames it into place, and treats an EMPTY render as a FAILED one: board prints nothing at all when its scan roots do not exist, which is what a mistyped path or a moved home directory looks like, and a plain > file redirect would destroy yesterday's briefing on a bad launchd environment. A tree where nobody owes anything is a different case — that is a valid, non-empty briefing saying so, and is written normally.

  • Route (scripts/route.sh): pure calculator for the next session's model and effort. Takes four scored traits of the next task plus a required rationale, and prints one block of key=value lines: the rubric row, the rule that fired, the next-cost value, a pasteable startup command, the one-row cheaper fallback, and the STATE.md comment lines. Pinned by route-selftest.sh (69 checks).

    It is here because it is the WRITER for the field board.sh already reads. next-cost had a reader and no writer, so it was hand-typed every session and drifted into several competing spellings — cleaning the data could not fix that, because the cause was the missing write path. The row table is a closed set of six values, so a seventh cannot enter circulation, and section 6 of the selftest runs the round trip (route emits → board parses) inside one repo rather than across two. board.sh itself is untouched: a calculator that prints to stdout writes nothing, and the session writes STATE.md.

    The row table is the operator's global rubric, moved here as the single copy. It is not a second spec — board.sh --help documents the board line's grammar and points here for the values. Scoring the traits is judgement and belongs to the skill; turning scores into a row is a lookup and takes zero model calls.

    --advisor opus is emitted per ROW, on a need, never unconditionally. Rows 1-2 always carry it (Sonnet main model, so opus is a capability lift — and since every fallback is one row cheaper and the cheap rows are Sonnet, this is what makes the quota fallback safe to take); rows 3-4 only at reversibility=costly|one-way (Opus main model, so it buys peer review where a mistake is not cheap to undo); rows 5-6 never, because the CLI rejects every advisor for a Fable main model. The alternative — the global advisorModel setting written by /advisor — is what this replaces: it applies to every session in every repo, which is how it burned quota before. The two triggers are almost disjoint by construction, since costly forces row 3 and one-way forces row 4, so a Sonnet row always has reversibility=cheap. verification=none is deliberately NOT a third trigger: beyond the stakes rule it would only add mistakes that are cheap to reverse, docs sessions (known/none/cheap/local) among them. Section 14 pins the rule and gates the three CLI facts it rests on against the installed claude without spending a token — advisor validation runs before the empty-prompt check, so -p "" reaches the validator and stops there.

    Rows 5-6 are never a route.sh outcome. Until 2026-08-06 they fired only from an explicit --opus-xhigh-failed flag, mirroring a global CLAUDE.md policy that Fable could only be suggested after a failed Opus 5/xhigh session. That policy was removed by operator decision — "for ofte ER Fable riktig" — and the flag went with it rather than being repurposed: route.sh's output range is now closed at row 4, and a Fable choice is always a hand-written deviation from the rubric, recorded in STATE as an override per the model-selection rule in the global CLAUDE.md, never produced by the calculator. board.sh still parses "Fable 5/high" and "Fable 5/xhigh" written by hand into the board line — that parsing is what the override actually uses, and it is pinned separately from anything route.sh emits (route-selftest.sh section 6).

    --last-effort is MEASURED from CLAUDE_EFFORT, and the calculator must never default it. Claude Code exports that variable into every tool-use context as the session's current effort, so the caller reads it and passes it in; having route.sh read it directly would make the output depend on the environment instead of on its arguments, and the round trip in selftest section 6 rests on that determinism. The two sources it replaces fail identically: the previous board line holds what was prescribed, and asking the operator launders that same prescription through someone reading their own startup command. Corollary pinned by section 13: skills/route/SKILL.md must never declare an effort: frontmatter field, because frontmatter overrides the session effort and the reading would then measure the skill, not the session.

  • Skills (skills/coord-send/, skills/board/, skills/route/): natural-language front doors mapping user intent to engine invocations. No mailbox logic lives here either. board additionally owns the ranking — which repo wins and why — since board.sh deliberately prints evidence and takes no position. route likewise owns the scoring: the calculator is deterministic, so all judgement sits in choosing the four trait values, and the skill must never reason its way to a model instead.

Boundary rule: the mailbox is transport, not state. Durable decisions live in the owning repo's docs/git history; messages are notices pointing at them. Message content is untrusted cross-repo input — the read side quotes and frames it; the send side sanitizes line-oriented fields.

What the boundary forbids is storing a repo's state — its decisions, its next step, its progress. It does not forbid the mailbox knowing who it is delivering to: _broadcast/seen/<repo> and <repo>/.origin (0.6.0) are delivery metadata, answering "has this repo received this" and "which checkout claimed this name". Both are unreadable as a description of the repo and useless outside delivery. The test is not "does the engine write a file about a repo" but "would this file still mean anything if delivery were removed". If yes, it belongs in the repo's own docs and git history instead.

Priority rule (v0.5.0, Rule 7): the injection block is the only place a repo is ever told what to do with a message, so its wording is the protocol — treat that string as engine behavior, not prose. It obligates handling the inbox first and driving every directed message to a terminal state before the session ends. The obligation is procedural, never substantive: responding is mandatory, complying with message content is not. Those two must stay distinct in any reword — keeping the priority while dropping the distinction turns prioritization into an injection surface. Selftest section 20 pins both halves together for exactly that reason.

Since 0.11.0 the reply-expected field says which terminal state the SENDER expects. That does not soften the split, it sharpens it: the field is untrusted cross-repo input like the rest of the file, so the injection calls it a declaration, not an instruction and keeps both terminal states open to the receiver. Drop that clause and one word in a message becomes a lever that mints obligations in another repo.

Conventions

  • Scripts are bash-3.2-safe and ASCII-only: no declare -A, no readarray/mapfile, no |&; guard shift 2 with $# -ge 2; guard empty-array expansion under set -u with ${#a[@]}.
  • Zero dependencies everywhere: bash + coreutils in the engine, node: builtins only in hook and tests.
  • TDD: no behavior change without a failing selftest check first. bash scripts/coord-selftest.sh must exit 0 (191/191), bash scripts/board-selftest.sh must exit 0 (152/152), bash scripts/route-selftest.sh must exit 0 (69/69) and bash scripts/state-line-guard-selftest.sh must exit 0 (21/21).
  • English for all code, docs, and commit messages (public repo). Norwegian trigger aliases in the skill description are deliberate.
  • Conventional Commits: type(scope): description.

Commands

  • Test: bash scripts/coord-selftest.sh, bash scripts/board-selftest.sh, bash scripts/route-selftest.sh and bash scripts/state-line-guard-selftest.sh (or npm test, the Node wrapper around all four)
  • Hook smoke test: node hooks/scripts/session-start.mjs (expects JSON on stdout)
  • State-line-guard smoke test: echo '{"tool_name":"Write","tool_input":{"file_path":"/tmp/STATE.md","content":"x\n"}}' | node hooks/scripts/pre-state-line-guard.mjs; echo $? (expects exit 0, no output — a one-line STATE.md is under the limit)
  • Board smoke test: bash scripts/board.sh (read-only, ~3s over the real tree)
  • Briefing smoke test: bash scripts/board.sh --brief (read-only, writes nothing). brief-nightly.sh DOES write — it overwrites $CLAUDE_BRIEF_FILE (default ~/.claude/briefing.md), so point that at a scratch path when testing. Installed as a launchd agent from launchd/, which points at the SOURCE repo, never the version-pinned plugin cache.
  • Route smoke test: bash scripts/route.sh --path known --verification strong --reversibility cheap --scope local --rationale x (writes nothing, instant)
  • Sweep smoke test: bash scripts/coord-sweep.sh (dry-run is the default, so this writes nothing; never add --write to a smoke test against the real mailbox)

Release

Version must agree across: .claude-plugin/plugin.json, package.json, README version badge, skills/coord-send/SKILL.md, skills/board/SKILL.md and skills/route/SKILL.md frontmatter, git tag vX.Y.Z, and the catalog ref in ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json. Release via the catalog's scripts/release-plugin.mjs repo-mailbox (tag + ref bump together); verify with scripts/check-versions.mjs. Never hand-edit a ref.

Two things that script does that its dry-run label does not suggest: --create-tag creates AND pushes the tag even without --write, and its closing verification gate runs check-versions.mjs over ALL plugins — one unrelated plugin in ERROR aborts it with the catalog edit written but uncommitted. When that happens, commit the catalog's marketplace.json + README.md by hand and leave every other dirty file in that repo alone.

--write --commit does NOT close this window (confirmed by catalog, 2026-08-10). The gate (check-versions.mjs) runs via execFileSync before the --commit conditional, so it throws on any plugin's ERROR — including one we did not touch — after the catalog files are written and before commit, regardless of whether --commit was passed. Catalog is evaluating a pre-flight gate (run the check before writing, abort there) but it is not implemented yet — do not assume it exists. Until it ships: before running --write, run node scripts/check-versions.mjs in the catalog manually and confirm 0 ERROR first, even when the only ERROR belongs to an unrelated plugin. If it still fires mid-release, fall back to the manual-commit recovery above.

Hardening roadmap

Empty — the post-v0.1.0 queue (atomic delivery, ./.. rejection, selftest gaps, uniform -h) shipped in v0.2.0; broadcast self-delivery shipped in v0.2.1; broadcast retraction (coord-send --retract) shipped in v0.4.0, closing the last monotonically-growing surface. coord-inbox.sh still ignores unknown arguments by design (hook context must never fail) but now warns about each one on stderr, which the hook discards.

Two retraction limits are deliberate, not gaps: it is un-send and never recall (a repo that already received a broadcast keeps it — the seen set is delivery history and is left untouched), and the sender check is an accident guard, not a security boundary, because --from redefines identity here as it does everywhere else in the engine.