# repo-mailbox Renamed from `coord` in v0.3.0. The plugin/repo is `repo-mailbox`; the CLI (`coord-send.sh`, `coord-inbox.sh`, `coord-done.sh`), the mailbox root (`~/.claude/coord/`) and `CLAUDE_COORD_DIR` deliberately kept their names — they are the transport protocol, not the product. ## Context Local inter-repo coordination mailbox for Claude Code, packaged as a marketplace plugin. Three components, one boundary: - **Engine (`scripts/*.sh`):** bash owns all mailbox semantics — filename grammar, frontmatter, delivery, archiving, the seen set. `coord-send.sh` writes, `coord-inbox.sh` reads (formatted for context injection), `coord-done.sh` archives, `coord-count.sh` counts without delivering, `coord-sweep.sh` closes the aged FYI backlog machine-wide. Everything is pinned by `coord-selftest.sh` (220 checks, throwaway mailbox via `CLAUDE_COORD_DIR`). **`ktg-plugin-marketplace` is a RETIRED `--to` address (operator decision 2026-08-15), rejected rather than redirected.** It is a polyrepo directory, not a git repo, so `basename(git toplevel)` can never resolve to it and no session was ever able to hold that identity naturally - mail for it belongs to `catalog` instead. A silent redirect was considered and declined: it delivers mail somewhere the sender does not believe it landed, which is the same misdelivery defect this closes a second time (2 messages sat undelivered 2 days on this exact misaddressing before `catalog`'s H4 count caught 6 more). Rejection fails loud at the sender, at the moment the mistake is made. Only `--to` is retired, not `--from` - the defect was mail *arriving* there, never mail claiming to *originate* there. **`coord-send --reply-to` asserts "marked handled" against GROUND TRUTH, and the exit code alone is NOT that ground truth.** The line used to print unconditionally with `coord-done`'s output discarded (`>/dev/null 2>&1`), so the one line a session relies on to close a reply debt was false at the moment it was printed — measured with a stub `coord-done` exiting 1: original still in the inbox, no `archive/`, exit 0, "marked handled". Every reply this repo sent had to be verified by hand afterwards, which is what a false success in the TRANSPORT costs. The check is `exit 0` **and** the original no longer being at `$COORD/$FROM/inbox/$REPLYTO`, because `coord-done` exits 0 when it archives nothing (an unknown name is idempotently fine by its own contract), so a nonzero-exit test still certifies a message that never moved. The path is recomputed rather than reusing `$REPLY_ORIG`, which resolves to the inbox OR the archive — replying to an already-archived original moves nothing and must not warn. Failure is exit **1**, a new status: the reply WAS delivered and re-sending would duplicate it, so 2 stays the nothing-was-written status it has always been. Selftest section 34 pins all four cases (fails outright / exits 0 without moving / real happy path / archive-path reply). **`coord-sweep.sh` is the only path that closes a message with no human in the loop, and every constraint on it follows from that.** It may close exactly one mechanically decidable class - `reply-expected: no`, older than the grace window - because a message that owes a reply can only be answered by a session in the repo that owes it. Dry-run is the default, inverted from the rest of the engine, since this is the one script that destroys pending state. It closes through `coord-done.sh --repo` rather than moving files, so the archive layout and the `_broadcast` refusal stay in one place. And it logs every closure with sender and subject, because directed messages have no seen-tracking: the sweep genuinely cannot tell "seen and ignored" from "never delivered", so a notice can be closed unread and the log is the only record that it existed. Widening the class, defaulting to `--write`, or dropping the log each independently turn this from a bounded cleanup into silent data loss. **`coord-count.sh` prints TWO integers per mailbox** (`\t\t`), and the first must stay pending: `board.sh` counts the same inbox files itself, so a debt-only count would put two different numbers under one name. Debt is read from `reply-expected` in the FRONTMATTER BLOCK ONLY - a body line is untrusted input and must not be able to silence a debt - and an absent field means a reply IS owed, because every message written before 0.11.0 lacks it. **Reading is delivering — counting is not.** `coord-inbox.sh` records a broadcast as seen once it has printed it, so it can never be used to survey other repos: doing so would consume each one's backlog silently, and the seen set is delivery history that retraction deliberately leaves alone. `coord-count.sh` exists for every "what is pending" question and writes nothing at all. Any future read-shaped feature belongs there, not in the read path. - **Hook (`hooks/scripts/session-start.mjs`):** thin zero-dependency Node wrapper (marketplace convention: hooks are `.mjs`) that calls `coord-inbox.sh` and emits the `hookSpecificOutput.additionalContext` envelope. No mailbox logic lives here. Always exits 0. - **Hook (`hooks/scripts/pre-state-line-guard.mjs`):** a `PreToolUse` hook on `Write|Edit` that enforces the STATE.md convention's `maks ~120 linjer` (global CLAUDE.md; raised from `~60` by operator decision 2026-08-14 — see the dated paragraph below) mechanically. It exists because the prose limit alone failed: a real STATE.md drifted to 155-156 lines before an /insights sweep of 160 sessions noticed, and one trim pass on it *increased* the line count instead of shrinking it. org-ops dispatched the work order (20260814T144553Z) asking for a `PostToolUse` hook — that was the wrong event, and the fix is not cosmetic: `PostToolUse` fires only after the tool has already written the file (confirmed against the official hooks docs, 2026-08-14 — "Can block? No", stderr is shown to the model but the write already landed), so it cannot stop an oversized STATE.md from landing, only nag about it afterward. `PreToolUse` is the only event that can deny the call before the file is touched, which is what "enforces" has to mean here. Denial is stderr + `exit 2`, matching `llm-security`'s `pre-write-pathguard.mjs` — the only other `PreToolUse` `Write|Edit` guard in this marketplace — rather than the `hookSpecificOutput.permissionDecision` JSON form; both block, and matching the sibling convention keeps one idiom for "block a write" instead of two. For `Write` the projected content is the call's own `content`; for `Edit` it is the CURRENT on-disk file (read fresh, since `PreToolUse` fires before the edit is applied) with `old_string` replaced by `new_string` — every occurrence when `replace_all` is set, otherwise only the first, mirroring what the real Edit tool does. Getting `replace_all` wrong in either direction is not a hypothetical: a hook that only ever replaced the first occurrence would silently pass a bulk edit that balloons the file, so `state-line-guard-selftest.sh` (21 checks) pins a fixture where only counting every `replace_all` occurrence produces the correct denial. Anything the hook cannot project with confidence — a missing file, an `old_string` that is not present, fields of the wrong type — is left to the real tool, which reports a clearer error than a guess here would; the guard only ever touches files named exactly `STATE.md`, at any depth, matching the same basename rule the global session-start hook's nearest-STATE-wins search already uses. **It is a RATCHET against the file's current size, not a flat gate at 60 — found by advisor review before the tag landed, not by the selftest, which had no fixture for it.** The first cut compared the projected line count only against `MAX_LINES`, never against what the file already was, so trimming an oversized STATE.md from, say, 156 to 100 lines — still over 60, but strictly smaller — was denied exactly like growing it would have been. Verified empirically against the real tree (2026-08-14): `wc -l ~/repos/*/STATE.md ~/repos/*/*/STATE.md | awk '$1 > 60'` found 23 files already over 60 lines, one at 1405. Shipped as a flat gate, this hook would have made most of the machine's STATE.md files un-editable except by a single write landing at `<=60` in one shot — backwards for a guard whose whole point is making the trim the /insights finding asked for actually possible. The fix reads the file's current line count for BOTH tool types (previously only `Edit` read the file at all) and denies only when the projection is over `MAX_LINES` **and** larger than that current count: a compliant file still cannot grow past the limit, a brand-new file still cannot be created oversized (current defaults to 0), but an already-oversized file can always be edited toward compliance, one write at a time, without ever making it worse. Section 8 of the selftest pins all four cases: shrink-while-still-over-limit allows, same-size-rewrite allows, grow-an- already-oversized-file still denies, and create-new-oversized-file still denies. **`MAX_LINES` raised 60 -> 120, operator decision 2026-08-14 (evening), reported via coord by `.claude` after the global CLAUDE.md prose was already updated.** The ratchet mechanics above are unchanged - only the constant moved, plus every selftest fixture and boundary value that encoded 60 as a literal. Re-verified empirically against the real tree at the new threshold (2026-08-14): `wc -l ~/repos/*/STATE.md ~/repos/*/*/STATE.md | awk '$1 > 120'` found 13 files already over 120 lines, one at 1496 - the historical 23-files-over-60/one-at-1405 figures above describe the tree as it was at the moment the ratchet bug was found, not the current threshold, and are left as-is rather than rewritten. `session-start.mjs`'s 160-line injection window still covers the new 120-line limit with room to spare, so no change was needed there. **The Edit path's `current.replace(oldStr, newStr)` was a dollar-pattern injection bug, found and fixed 2026-08-15.** Passing `newStr` as a STRING makes JavaScript interpret `$`-sequences inside it ($&, `` $` ``, `$'`, `$$`, `$n`) as special replacement patterns, even though `oldStr` (the search side) is a plain string, not a RegExp. A `new_string` documenting old backtick-substitution style (`` $`cmd` ``) - exactly the prose a STATE.md's shell-conventions section writes routinely - triggers it. Measured against the real bug (`.claude/STATE.md`): a 5-line addition on a 112-line file projected to 219 lines and was wrongly denied. Direction is always fail-CLOSED (never fail-open: it can only over-block, never under-block a real oversize), but it made exactly the kind of STATE.md that documents shell conventions hard to edit. Fix: replace with a function, `current.replace(oldStr, () => newStr)` - a function result is never pattern-substituted, so this covers every `$`-sequence at once, not a `` $` ``-specific escape. The `replace_all` branch (`split`/`join`) was never affected - `join` does not interpret its argument as a pattern. Pinned by state-line-guard-selftest.sh section 9 (`$\`` as the real repro, `$&` as a second sequence proving the fix is general). **Since ORDRE 42 (operator, 2026-08-16) it carries a SECOND invariant: the projected content may not claim `status=done` in its board line while the repo holds commits the branch's upstream does not have.** Measured that day: two sessions had their push refused by the UFW rate limit on port 22, said so honestly in the coord inbox, and wrote `status=done` regardless - board line green, one commit unpushed, published surface 404. `done` meant "the session finished" where every reader takes it to mean "the work landed", and since `done` drops a repo from the board plan, `morning --say ` could not reach either of them: one defect hid the other. **The order recommended a session-end hook and that direction does not exist in the form it assumes - measured against the official hooks docs, not reasoned.** `Stop` fires "once per turn", not once when the session ends, with no signal marking the last turn; its exit 2 "prevents Claude from stopping, continues the conversation", so a repo that genuinely cannot push (the very rate limit that caused the incident) would get a session that will not end. `SessionEnd` is the once-per-session event and cannot block at all ("Can block? No" - exit 2 "shows stderr to user only"), which is the after-the-fact nagging the order explicitly refused. Warn-on-write plus deny-at-session-end inherits the broken half and buys nothing. So the deny sits on the write, where the false claim is actually made. **The false-positive trap is real but bounded, and the deny is escapable by telling the truth.** STATE.md is written BEFORE the session's final commit, so a session that batches its pushes does hold unpushed commits at that moment - but the global git rule already requires a push immediately after every commit, and the real tree bears that out (2026-08-16: 43 of 44 repos carrying a STATE.md had nothing unpushed; the one exception was `status=blocked` and honest). `status=blocked` and `status=in-progress` stay writable in the same single edit, so a session that cannot push is never wedged - only stopped from claiming otherwise. **No ratchet here, unlike the line limit above, and the asymmetry is the reason.** An oversized file needs many writes to come back under the limit, so denying the intermediate steps would make trimming impossible; a false `done` is corrected by changing one token in the write already being made. A "deny only the transition into done" variant was rejected outright: the common shape is a repo that ended `done` last session and writes `done` again this session, which such a rule waves straight through. It **fails OPEN** on every git uncertainty - no upstream, detached HEAD, missing remote-tracking ref, not a repo, git slow or absent - because 8 of those 44 repos have no upstream at all (one already `status=done`), and a confident denial resting on a measurement that never happened is the worse error. The board line is selected with `board.sh`'s own anchor (`^` instead - the same `[^;>]*` + trim shape `next-cost` already used for its own reason (spec-conformant values contain spaces and capitals) - so the classifier and `route.sh`'s case statement see the value un-truncated and the exact-match either accepts it or correctly calls it MALFORMED / `command_missing=`. `blocked-on`'s `[A-Za-z0-9._-]*` capture is a different, wider class feeding a different mechanism (matched against scanned repo names for chain-root credit, not a closed vocabulary) and was checked, not touched - grepping `scripts/board.sh` for `[a-z-]*` after the fix returns nothing. Selftest fixtures `repo-status-prefix`, `repo-status-case`, `repo-trait-prefix` and `repo-revers-prefix` pin all four cases, plus the known-positive controls that a genuinely absent board line still reads `?` and a genuinely valid route line still derives a real command. - **Route (`scripts/route.sh`):** pure calculator for the next session's model and effort. Takes four scored traits of the next task plus a required rationale, and prints one block of `key=value` lines: the rubric row, the rule that fired, the `next-cost` value, a pasteable startup command, the one-row cheaper fallback, and the STATE.md comment lines. Pinned by `route-selftest.sh` (69 checks). **It is here because it is the WRITER for the field `board.sh` already reads.** `next-cost` had a reader and no writer, so it was hand-typed every session and drifted into several competing spellings — cleaning the data could not fix that, because the cause was the missing write path. The row table is a closed set of six values, so a seventh cannot enter circulation, and section 6 of the selftest runs the round trip (route emits → board parses) *inside* one repo rather than across two. `board.sh` itself is untouched: a calculator that prints to stdout writes nothing, and the session writes STATE.md. **The row table is the operator's global rubric, moved here as the single copy.** It is not a second spec — `board.sh --help` documents the board line's *grammar* and points here for the *values*. Scoring the traits is judgement and belongs to the skill; turning scores into a row is a lookup and takes zero model calls. **`--advisor opus` is emitted per ROW, on a need, never unconditionally.** Rows 1-2 always carry it (Sonnet main model, so opus is a capability lift — and since every fallback is one row cheaper and the cheap rows are Sonnet, this is what makes the quota fallback safe to take); rows 3-4 only at `reversibility=costly|one-way` (Opus main model, so it buys peer review where a mistake is not cheap to undo); rows 5-6 never, because the CLI rejects every advisor for a Fable main model. The alternative — the global `advisorModel` setting written by `/advisor` — is what this replaces: it applies to every session in every repo, which is how it burned quota before. The two triggers are almost disjoint by construction, since `costly` forces row 3 and `one-way` forces row 4, so a Sonnet row always has `reversibility=cheap`. `verification=none` is deliberately NOT a third trigger: beyond the stakes rule it would only add mistakes that are cheap to reverse, docs sessions (`known/none/cheap/local`) among them. Section 14 pins the rule and gates the three CLI facts it rests on against the installed `claude` without spending a token — advisor validation runs before the empty-prompt check, so `-p ""` reaches the validator and stops there. **Rows 5-6 are never a `route.sh` outcome.** Until 2026-08-06 they fired only from an explicit `--opus-xhigh-failed` flag, mirroring a global CLAUDE.md policy that Fable could only be *suggested* after a failed Opus 5/xhigh session. That policy was removed by operator decision — "for ofte ER Fable riktig" — and the flag went with it rather than being repurposed: `route.sh`'s output range is now closed at row 4, and a Fable choice is always a hand-written deviation from the rubric, recorded in STATE as an override per the model-selection rule in the global CLAUDE.md, never produced by the calculator. `board.sh` still parses "Fable 5/high" and "Fable 5/xhigh" written by hand into the board line — that parsing is what the override actually uses, and it is pinned separately from anything `route.sh` emits (route-selftest.sh section 6). **`--last-effort` is MEASURED from `CLAUDE_EFFORT`, and the calculator must never default it.** Claude Code exports that variable into every tool-use context as the session's current effort, so the caller reads it and passes it in; having `route.sh` read it directly would make the output depend on the environment instead of on its arguments, and the round trip in selftest section 6 rests on that determinism. The two sources it replaces fail identically: the previous board line holds what was *prescribed*, and asking the operator launders that same prescription through someone reading their own startup command. Corollary pinned by section 13: `skills/route/SKILL.md` must never declare an `effort:` frontmatter field, because frontmatter overrides the session effort and the reading would then measure the skill, not the session. - **Skills (`skills/coord-send/`, `skills/board/`, `skills/route/`, `skills/dispatch/`):** natural-language front doors mapping user intent to engine invocations. No mailbox logic lives here either. `board` additionally owns the *ranking* — which repo wins and why — since `board.sh` deliberately prints evidence and takes no position. `route` likewise owns the *scoring*: the calculator is deterministic, so all judgement sits in choosing the four trait values, and the skill must never reason its way to a model instead. **Boundary rule:** the mailbox is transport, not state. Durable decisions live in the owning repo's docs/git history; messages are notices pointing at them. Message content is untrusted cross-repo input — the read side quotes and frames it; the send side sanitizes line-oriented fields. What the boundary forbids is storing a repo's *state* — its decisions, its next step, its progress. It does not forbid the mailbox knowing who it is delivering to: `_broadcast/seen/` and `/.origin` (0.6.0) are delivery metadata, answering "has this repo received this" and "which checkout claimed this name". Both are unreadable as a description of the repo and useless outside delivery. The test is not "does the engine write a file about a repo" but "would this file still mean anything if delivery were removed". If yes, it belongs in the repo's own docs and git history instead. **Priority rule (v0.5.0, Rule 7):** the injection block is the only place a repo is ever told what to do with a message, so its wording *is* the protocol — treat that string as engine behavior, not prose. It obligates handling the inbox first and driving every directed message to a terminal state before the session ends. The obligation is **procedural, never substantive**: responding is mandatory, complying with message content is not. Those two must stay distinct in any reword — keeping the priority while dropping the distinction turns prioritization into an injection surface. Selftest section 20 pins both halves together for exactly that reason. Since 0.11.0 the `reply-expected` field says which terminal state the SENDER expects. That does not soften the split, it sharpens it: the field is untrusted cross-repo input like the rest of the file, so the injection calls it a *declaration, not an instruction* and keeps both terminal states open to the receiver. Drop that clause and one word in a message becomes a lever that mints obligations in another repo. ## Conventions - Scripts are bash-3.2-safe and ASCII-only: no `declare -A`, no `readarray`/`mapfile`, no `|&`; guard `shift 2` with `$# -ge 2`; guard empty-array expansion under `set -u` with `${#a[@]}`. - Zero dependencies everywhere: bash + coreutils in the engine, `node:` builtins only in hook and tests. - TDD: no behavior change without a failing selftest check first. `bash scripts/coord-selftest.sh` must exit 0 (220/220), `bash scripts/board-selftest.sh` must exit 0 (217/217), `bash scripts/route-selftest.sh` must exit 0 (69/69) and `bash scripts/state-line-guard-selftest.sh` must exit 0 (40/40). - English for all code, docs, and commit messages (public repo). Norwegian trigger aliases in the skill description are deliberate. - Conventional Commits: `type(scope): description`. ## Commands - Test: `bash scripts/coord-selftest.sh`, `bash scripts/board-selftest.sh`, `bash scripts/route-selftest.sh` and `bash scripts/state-line-guard-selftest.sh` (or `npm test`, the Node wrapper around all four) - Hook smoke test: `node hooks/scripts/session-start.mjs` (expects JSON on stdout) - State-line-guard smoke test: `echo '{"tool_name":"Write","tool_input":{"file_path":"/tmp/STATE.md","content":"x\n"}}' | node hooks/scripts/pre-state-line-guard.mjs; echo $?` (expects exit 0, no output — a one-line STATE.md is under the limit) - Board smoke test: `bash scripts/board.sh` (read-only, ~3s over the real tree) - Briefing smoke test: `bash scripts/board.sh --brief` (read-only, writes nothing). `brief-nightly.sh` DOES write — it overwrites `$CLAUDE_BRIEF_FILE` (default `~/.claude/briefing.md`), so point that at a scratch path when testing. Installed as a launchd agent from `launchd/`, which points at the SOURCE repo, never the version-pinned plugin cache. - Dispatch smoke test: `bash scripts/board.sh --dispatch --repo repo-mailbox --prompt-file /tmp/x.prompt --target-pane yes --path known --verification strong --reversibility cheap --scope local --rationale smoke` (read-only; needs a non-empty `/tmp/x.prompt`. Use `--target-pane yes` in a smoke test: it produces no plan file, so nothing can be handed to `morning` by accident) - Route smoke test: `bash scripts/route.sh --path known --verification strong --reversibility cheap --scope local --rationale x` (writes nothing, instant) - Sweep smoke test: `bash scripts/coord-sweep.sh` (dry-run is the default, so this writes nothing; never add `--write` to a smoke test against the real mailbox) ## Release Version must agree across: `.claude-plugin/plugin.json`, `package.json`, README version badge, `skills/coord-send/SKILL.md`, `skills/board/SKILL.md`, `skills/route/SKILL.md` and `skills/dispatch/SKILL.md` frontmatter, git tag `vX.Y.Z`, and the catalog `ref` in `ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json`. Release via the catalog's `scripts/release-plugin.mjs repo-mailbox` (tag + ref bump together); verify with `scripts/check-versions.mjs`. Never hand-edit a ref. Two things that script does that its dry-run label does not suggest: `--create-tag` creates AND pushes the tag even without `--write`, and its closing verification gate runs `check-versions.mjs` over ALL plugins — one unrelated plugin in ERROR aborts it with the catalog edit written but uncommitted. When that happens, commit the catalog's `marketplace.json` + `README.md` by hand and leave every other dirty file in that repo alone. **`--write --commit` does NOT close this window (confirmed by catalog, 2026-08-10).** The gate (`check-versions.mjs`) runs via `execFileSync` before the `--commit` conditional, so it throws on any plugin's ERROR — including one we did not touch — after the catalog files are written and before commit, regardless of whether `--commit` was passed. Catalog is evaluating a pre-flight gate (run the check before writing, abort there) but it is **not implemented yet** — do not assume it exists. Until it ships: before running `--write`, run `node scripts/check-versions.mjs` in the catalog manually and confirm 0 ERROR first, even when the only ERROR belongs to an unrelated plugin. If it still fires mid-release, fall back to the manual-commit recovery above. ## Hardening roadmap Empty — the post-v0.1.0 queue (atomic delivery, `.`/`..` rejection, selftest gaps, uniform `-h`) shipped in v0.2.0; broadcast self-delivery shipped in v0.2.1; broadcast retraction (`coord-send --retract`) shipped in v0.4.0, closing the last monotonically-growing surface. `coord-inbox.sh` still ignores unknown arguments by design (hook context must never fail) but now warns about each one on stderr, which the hook discards. Two retraction limits are deliberate, not gaps: it is un-send and never recall (a repo that already received a broadcast keeps it — the seen set is delivery history and is left untouched), and the sender check is an accident guard, not a security boundary, because `--from` redefines identity here as it does everywhere else in the engine.