repo-mailbox/CLAUDE.md
Kjell Tore Guttormsen 03e712a423 fix(board): plan/brief use actual debt, not raw pending mail
board.sh --plan group 2 and --brief conflated every unhandled inbox
message with an obligation to reply, including ones the sender
declared reply-expected: no (a notice, not a request). Reported by
morning-driver (2026-08-11), independently reproduced against the live
mailbox on 2026-08-13: 27 of 72 pending messages were notices. Both
paths now join against coord-count.sh's owed column instead, so a
done/deferred/blocked repo whose only mail is FYI no longer gets a
plan tab, and --brief no longer counts a notice as an obligation. The
table's raw INN column is unchanged by design.

While extending that join with a second lookup file, found and fixed
a more severe, independent defect: the existing $UNBLOCKS/$RECORDS
join used the NR==FNR awk idiom, which silently empties the entire
plan whenever the first file is empty -- i.e. whenever the repo tree
has zero blocked repos, a common, ordinary state, not an edge case.
Verified against the shipped 0.21.0 script. Fixed by matching on
FILENAME instead of cumulative line counts, for both lookup files.

Also repoints the README's governance link at repo-standard's
canonical GOVERNANCE.md (was pointing at the marketplace's copy),
per org-ops D11.

board-selftest.sh: 142 -> 150 checks. Version 0.21.0 -> 0.22.0.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGWMPskXBsTjMrrQ2GofFx
2026-08-13 20:54:55 +02:00

415 lines
26 KiB
Markdown

# repo-mailbox
Renamed from `coord` in v0.3.0. The plugin/repo is `repo-mailbox`; the CLI
(`coord-send.sh`, `coord-inbox.sh`, `coord-done.sh`), the mailbox root
(`~/.claude/coord/`) and `CLAUDE_COORD_DIR` deliberately kept their names —
they are the transport protocol, not the product.
## Context
Local inter-repo coordination mailbox for Claude Code, packaged as a
marketplace plugin. Three components, one boundary:
- **Engine (`scripts/*.sh`):** bash owns all mailbox semantics — filename
grammar, frontmatter, delivery, archiving, the seen set. `coord-send.sh`
writes, `coord-inbox.sh` reads (formatted for context injection),
`coord-done.sh` archives, `coord-count.sh` counts without delivering,
`coord-sweep.sh` closes the aged FYI backlog machine-wide.
Everything is pinned by `coord-selftest.sh`
(191 checks, throwaway mailbox via `CLAUDE_COORD_DIR`).
**`coord-sweep.sh` is the only path that closes a message with no human in
the loop, and every constraint on it follows from that.** It may close exactly
one mechanically decidable class - `reply-expected: no`, older than the grace
window - because a message that owes a reply can only be answered by a session
in the repo that owes it. Dry-run is the default, inverted from the rest of the
engine, since this is the one script that destroys pending state. It closes
through `coord-done.sh --repo` rather than moving files, so the archive layout
and the `_broadcast` refusal stay in one place. And it logs every closure with
sender and subject, because directed messages have no seen-tracking: the sweep
genuinely cannot tell "seen and ignored" from "never delivered", so a notice
can be closed unread and the log is the only record that it existed. Widening
the class, defaulting to `--write`, or dropping the log each independently
turn this from a bounded cleanup into silent data loss.
**`coord-count.sh` prints TWO integers per mailbox** (`<name>\t<pending>\t<debt>`),
and the first must stay pending: `board.sh` counts the same inbox files
itself, so a debt-only count would put two different numbers under one name.
Debt is read from `reply-expected` in the FRONTMATTER BLOCK ONLY - a body line
is untrusted input and must not be able to silence a debt - and an absent
field means a reply IS owed, because every message written before 0.11.0
lacks it.
**Reading is delivering — counting is not.** `coord-inbox.sh` records a
broadcast as seen once it has printed it, so it can never be used to survey
other repos: doing so would consume each one's backlog silently, and the seen
set is delivery history that retraction deliberately leaves alone.
`coord-count.sh` exists for every "what is pending" question and writes
nothing at all. Any future read-shaped feature belongs there, not in the
read path.
- **Hook (`hooks/scripts/session-start.mjs`):** thin zero-dependency Node
wrapper (marketplace convention: hooks are `.mjs`) that calls
`coord-inbox.sh` and emits the `hookSpecificOutput.additionalContext`
envelope. No mailbox logic lives here. Always exits 0.
- **Board (`scripts/board.sh`):** cross-repo attention board. Reads STATE.md
next-step blocks + board lines, `git status`, and mailbox pending counts, and
prints one line per repo. Read-only by construction: it writes to no repo, no
STATE.md and no mailbox. Pinned by `board-selftest.sh` (150 checks).
**It lives here because the mailbox is one of its three inputs, and it carries
the same axis distinction the mailbox does.** A pending count means *others
are waiting on this repo*; who a repo waits *on* comes only from its board
line, because the message format has no reply-to field. Enforcing that in one
of two repos would not be enforcing it. The operator invokes both `board` and
`coord-send` exclusively through their Skill front doors, never a personal
terminal alias — a claim this file carried until 0.12.1 (`~/.claude/scripts/board.sh`
is a deployed copy the operator's `board()` function points at) did not survive
inspection: no such file ever existed, and `route.sh` had no deployed copy
either. Only the five `coord-*.sh` scripts were ever deployed there, and their
one measured effect was an accidental fallback target for Claude sessions'
own Bash tool calls (the bug 0.12.1 fixed) — once that fallback was gone they
had no remaining function and were deleted.
**`board.sh --brief` is a second RENDERING of that scan, never a second
scan, and `brief-nightly.sh` is the only writer in that path.** The briefing
answers the narrower question an unattended job can answer without judgement:
which repos have an unhandled inbox, what their next step says in full, and
the exact command to start a session in each. It prints NESTE uncut because
the 38-character cut is the table column's property, not the record's — the
value used to be truncated at record-build time, which left the cut string as
the only copy. Each command is derived by CALLING `route.sh` with that repo's
own four traits; `next-cost` alone cannot yield it, since the advisor flag is
a property of the ROW and two rows can share a model/effort pair while
differing on it. A repo with no route line is told so rather than handed a
guess, because a guessed command reads as authoritative.
**`board.sh --plan` is the THIRD rendering, and the only one that takes a
position.** It answers which repos to open a tab for today, in what order,
with which command. The position it takes is the ORDER and nothing else -
there is no cutoff, so the plan hides nothing, and every term is a lookup over
fields the scan already read. Since 0.20.0 it is FIVE ORDERED GROUPS -
chain-root, debt, planned, in-progress, undeclared - ranked within a group by
that group's own quantity, then a Sonnet next-cost, then oldest plan first.
**0.19.0 shipped a weighted score here and 0.20.0 replaced it, and the
objection the score answered is ACCEPTED, not forgotten.** A group order
genuinely cannot express "this repo owes one message and releases two others"
as one quantity; a score could, and that was its point. What a score could not
do was hold still for the second consumer - re-tuning 40 against 15 silently
reorders a parser living in another repo, and no test in THIS repo can catch
that. The operator weighed both and chose the lookup (2026-08-03). Write that
down every time this paragraph is edited: a later session that reads the
objection as an unfixed defect will "restore" the score, and the round trip is
the loop this file exists to stop.
**`planned` ranks ABOVE `in-progress`, inverted at 0.20.0 by operator
decision.** Turning a decision into motion is the slow step; live work is
already moving. Flipping it back is a policy change, not a sort fix.
**Debt is never excluded and never capped, and that is the rule most likely to
be "fixed" into a defect.** Excluding `blocked` or `done` is a claim about a
repo's OWN next step, which by definition cannot be moved, while owing a reply
is the other axis entirely - answering is often what unblocks it. Measured on
the real tree at 0.16.0, two of 26 planned repos were `done` with an unhandled
inbox. A cap on debt was proposed with the chain credit at 0.19.0 and DECLINED
by the operator. Sitting one group below chain-root credit is NOT that cap:
the debtor keeps its tab, its most-owed-first position among the other
debtors, and its `why=inbox:N`. A change that DROPPED a debtor from the plan
would be the declined cap wearing the group order as a disguise, and selftest
section 12 pins both halves - the root outranking four owed messages, and the
debtor keeping everything it had.
**"Debt" means OWED, never raw pending, since 0.22.0 - and this NARROWS what
counts as debt, it does not reopen the paragraph above.** The paragraph above
settles a different question: once a repo has debt, is it ever excluded or
capped (no). This one settles what counts as debt in the first place. Through
0.21.0, group 2's `keep`/`mag`/`why=inbox:N` and `--brief`'s whole "repo som
skylder et svar" listing were computed from the raw pending-file count - every
unhandled message in the inbox, including ones the sender declared
`reply-expected: no`. That is a notice, not a request, and 0.11.0 gave
`coord-count.sh` a second column (`owed`) for exactly this distinction - but
`board.sh` never read it. Reported by morning-driver (2026-08-11) and
independently reproduced against the live mailbox 2026-08-13: 27 of 72
pending messages (37.5%) were notices. The fix joins `--plan` and `--brief`
against `coord-count.sh`'s `owed` column by repo name (same technique as the
chain-root `$UNBLOCKS` join below), so a `done`/`deferred`/`blocked` repo
whose only mail is FYI no longer gets a tab, and `--brief` no longer counts a
notice as an obligation. This reverses a decision from session 41
(2026-08-10) that declined to build this filter, on the premise that "the
arrival of the request IS the admission signal" - a premise that assumed
group 2 already meant requests. It didn't; the code computed pending, the
comments already said "owed" throughout, and the plan's own printed header
("Utelatt naar repoet verken skylder svar...") already claimed the exclusion
was debt-based. The fix makes the code match what its own comments and
header already promised. Pinned by board-selftest.sh section 8/12 fixtures
`repo-done-fyi` (pending 2, owed 0 - excluded) and `repo-blocked-mixed`
(pending 3, owed 2 - planned on 2, not 3). The TABLE's `INN` column and the
raw scan (`RECORDS` field 6) are UNCHANGED - they answer "what is the state
of every repo," not "who is waiting on you," and stay on raw pending by
design.
**The `$UNBLOCKS`/`$RECORDS` join used `NR==FNR` through 0.21.0, and that
idiom silently drops the entire plan whenever the FIRST file is empty - fixed
to `FILENAME==` comparison in 0.22.0, found while adding the `$OWED` join
above.** Verified against the shipped 0.21.0 script: one in-progress repo
with an unhandled inbox message, zero blocked repos anywhere in the tree
(so `$UNBLOCKS` is empty, which is a common, ordinary tree state, not an
edge case) - `--plan` printed "0 tabber". `NR==FNR` is only true for the
FIRST file's own lines; when that file is empty, `FNR` and `NR` stay equal
for the ENTIRE next file too (not just its first line - verified with a
minimal awk reproduction), so every record in it is misrouted into the
`ub[]` branch and dropped via `next`. This was invisible to
board-selftest.sh because the fixture tree has carried at least one
`blocked` repo since the chain-root feature shipped, and it was invisible
on the real tree because `~/repos` currently always has one too - neither
is a guarantee. `FILENAME==UBF`/`FILENAME==OWF` compares the exact path,
never line counts, so an empty lookup file degrades to "nothing matched,"
never to "everything after it is misrouted."
**Chain-root credit lands on the ROOT and nowhere else.** For every `blocked`
repo the `blocked-on` edge is followed transitively to the first repo that is
not itself blocked. Crediting a blocked repo would open a tab that cannot move;
crediting only the direct blocker leaves a two-hop chain's root uncredited,
which is the shape the real tree actually had. A cycle, a `blocked-on` naming
an unscanned repo, and a blocked repo with no target must all credit NOBODY:
inventing a root there produces a plan that looks correct and sends the
operator to the wrong repo.
Repos with no board line rank last and are LABELLED rather than dropped,
because the table already prints a MERK line about them and a plan that
omitted them silently would repeat that defect.
It renders `key=value` blocks, not prose, because it has two consumers: the
operator, and a driver repo consuming the plan. Prose would make the rendered
format an API no test in THIS repo could hold stable for a consumer in
another. `command_missing=` carries both no-command causes (no route line, and
a route line route.sh rejects) because a bare `command=` is the shape of a
runnable command carrying nothing - a driver reading `^command=` would type an
empty line into a live pane. `route_cmd_for()` is the single reader of the
route-line grammar, shared with `--brief`, and distinguishes the two causes by
exit code rather than by an empty string.
**`paste=` and `dir=`/`command=` are the same fact for the two consumers, and
neither is redundant.** A driver moves the pane itself and then types the
command, so it needs them apart; a human needs ONE thing to select. Handing
the operator two fields to join by hand is not a saved output line, it is the
step where a session starts in the wrong repo - and it was measured the moment
the feature met its first user, who could not act on the block at all. `paste=`
is emitted only alongside `command=`: `paste=cd X && ` with nothing after it
would run the cd and then a bare newline, which fails SILENTLY by leaving the
operator in the right directory with no session started.
**Driving a terminal from the plan does NOT belong here, and the measurement
in `docs/ghostty-orchestration-measurement.md` is the argument, not taste.**
It is a version-pinned undocumented composition over a preview API whose
documented path is already broken upstream and whose regression was closed as
not planned, with a blast radius reaching into other repos' live sessions.
None of that is mailbox transport, and none of it may be able to break
`coord-inbox` or `board`. The dependency runs one way: the driver consumes the
plan, the plan never knows a terminal exists.
It also cross-checks itself against `coord-count.sh`, and that is not
belt-and-braces. The repo scan and the mailbox are two different populations:
a mailbox can carry a name no scan will ever produce — a declared non-git
surface (`CLAUDE_COORD_REPO`, e.g. `~/repos` itself) or a checkout outside the
roots. Measured on the real mailbox at 0.15.0: the briefing found 11 repos /
21 messages where `coord-count` saw 12 mailboxes / 22 pending, the missing one
being the declared surface `repos`. A briefing that only walks the scan
answers "who is waiting on you" with a number it quietly knows is short.
**Zero model calls, and that is the load-bearing property, not an
implementation detail.** The operator authenticates by subscription, so a
headless `claude -p` job draws from the same quota pool as interactive work.
Measured against 2.1.220: `--max-budget-usd` DOES bite under subscription auth
(`terminal_reason: budget_exhausted`, exit 1), but it aborts AFTER turn one,
never before it — floor ~0.25 USD-equivalent per turn on `claude-opus-5[1m]`.
It is a runaway brake, not a pre-flight gate. Making the briefing deterministic
removes the question entirely.
`board.sh` stays read-only, which is why the file write lives in the wrapper
instead of behind a `--brief --out FILE` flag. The wrapper renders to a temp
file in the target directory and renames it into place, and treats an EMPTY
render as a FAILED one: board prints nothing at all when its scan roots do not
exist, which is what a mistyped path or a moved home directory looks like, and
a plain `> file` redirect would destroy yesterday's briefing on a bad launchd
environment. A tree where nobody owes anything is a different case — that is a
valid, non-empty briefing saying so, and is written normally.
- **Route (`scripts/route.sh`):** pure calculator for the next session's model
and effort. Takes four scored traits of the next task plus a required
rationale, and prints one block of `key=value` lines: the rubric row, the rule
that fired, the `next-cost` value, a pasteable startup command, the one-row
cheaper fallback, and the STATE.md comment lines. Pinned by
`route-selftest.sh` (69 checks).
**It is here because it is the WRITER for the field `board.sh` already reads.**
`next-cost` had a reader and no writer, so it was hand-typed every session and
drifted into several competing spellings — cleaning the data
could not fix that, because the cause was the missing write path. The row
table is a closed set of six values, so a seventh cannot enter circulation,
and section 6 of the selftest runs the round trip (route emits → board parses)
*inside* one repo rather than across two. `board.sh` itself is untouched: a
calculator that prints to stdout writes nothing, and the session writes
STATE.md.
**The row table is the operator's global rubric, moved here as the single
copy.** It is not a second spec — `board.sh --help` documents the board line's
*grammar* and points here for the *values*. Scoring the traits is judgement
and belongs to the skill; turning scores into a row is a lookup and takes zero
model calls.
**`--advisor opus` is emitted per ROW, on a need, never unconditionally.**
Rows 1-2 always carry it (Sonnet main model, so opus is a capability lift —
and since every fallback is one row cheaper and the cheap rows are Sonnet,
this is what makes the quota fallback safe to take); rows 3-4 only at
`reversibility=costly|one-way` (Opus main model, so it buys peer review where
a mistake is not cheap to undo); rows 5-6 never, because the CLI rejects every
advisor for a Fable main model. The alternative — the global `advisorModel`
setting written by `/advisor` — is what this replaces: it applies to every
session in every repo, which is how it burned quota before. The two triggers
are almost disjoint by construction, since `costly` forces row 3 and `one-way`
forces row 4, so a Sonnet row always has `reversibility=cheap`.
`verification=none` is deliberately NOT a third trigger: beyond the stakes
rule it would only add mistakes that are cheap to reverse, docs sessions
(`known/none/cheap/local`) among them. Section 14 pins the rule and gates the
three CLI facts it rests on against the installed `claude` without spending a
token — advisor validation runs before the empty-prompt check, so `-p ""`
reaches the validator and stops there.
**Rows 5-6 are never a `route.sh` outcome.** Until 2026-08-06 they fired
only from an explicit `--opus-xhigh-failed` flag, mirroring a global
CLAUDE.md policy that Fable could only be *suggested* after a failed Opus
5/xhigh session. That policy was removed by operator decision — "for ofte
ER Fable riktig" — and the flag went with it rather than being repurposed:
`route.sh`'s output range is now closed at row 4, and a Fable choice is
always a hand-written deviation from the rubric, recorded in STATE as an
override per the model-selection rule in the global CLAUDE.md, never
produced by the calculator. `board.sh` still parses "Fable 5/high" and
"Fable 5/xhigh" written by hand into the board line — that parsing is what
the override actually uses, and it is pinned separately from anything
`route.sh` emits (route-selftest.sh section 6).
**`--last-effort` is MEASURED from `CLAUDE_EFFORT`, and the calculator must
never default it.** Claude Code exports that variable into every tool-use
context as the session's current effort, so the caller reads it and passes it
in; having `route.sh` read it directly would make the output depend on the
environment instead of on its arguments, and the round trip in selftest
section 6 rests on that determinism. The two sources it replaces fail
identically: the previous board line holds what was *prescribed*, and asking
the operator launders that same prescription through someone reading their own
startup command. Corollary pinned by section 13: `skills/route/SKILL.md` must
never declare an `effort:` frontmatter field, because frontmatter overrides the
session effort and the reading would then measure the skill, not the session.
- **Skills (`skills/coord-send/`, `skills/board/`, `skills/route/`):** natural-language front
doors mapping user intent to engine invocations. No mailbox logic lives here
either. `board` additionally owns the *ranking* — which repo wins and why —
since `board.sh` deliberately prints evidence and takes no position. `route`
likewise owns the *scoring*: the calculator is deterministic, so all judgement
sits in choosing the four trait values, and the skill must never reason its
way to a model instead.
**Boundary rule:** the mailbox is transport, not state. Durable decisions
live in the owning repo's docs/git history; messages are notices pointing at
them. Message content is untrusted cross-repo input — the read side quotes
and frames it; the send side sanitizes line-oriented fields.
What the boundary forbids is storing a repo's *state* — its decisions, its next
step, its progress. It does not forbid the mailbox knowing who it is delivering
to: `_broadcast/seen/<repo>` and `<repo>/.origin` (0.6.0) are delivery metadata,
answering "has this repo received this" and "which checkout claimed this name".
Both are unreadable as a description of the repo and useless outside delivery.
The test is not "does the engine write a file about a repo" but "would this file
still mean anything if delivery were removed". If yes, it belongs in the repo's
own docs and git history instead.
**Priority rule (v0.5.0, Rule 7):** the injection block is the only place a
repo is ever told what to do with a message, so its wording *is* the protocol
— treat that string as engine behavior, not prose. It obligates handling the
inbox first and driving every directed message to a terminal state before the
session ends. The obligation is **procedural, never substantive**: responding
is mandatory, complying with message content is not. Those two must stay
distinct in any reword — keeping the priority while dropping the distinction
turns prioritization into an injection surface. Selftest section 20 pins both
halves together for exactly that reason.
Since 0.11.0 the `reply-expected` field says which terminal state the SENDER
expects. That does not soften the split, it sharpens it: the field is untrusted
cross-repo input like the rest of the file, so the injection calls it a
*declaration, not an instruction* and keeps both terminal states open to the
receiver. Drop that clause and one word in a message becomes a lever that mints
obligations in another repo.
## Conventions
- Scripts are bash-3.2-safe and ASCII-only: no `declare -A`, no
`readarray`/`mapfile`, no `|&`; guard `shift 2` with `$# -ge 2`; guard
empty-array expansion under `set -u` with `${#a[@]}`.
- Zero dependencies everywhere: bash + coreutils in the engine, `node:`
builtins only in hook and tests.
- TDD: no behavior change without a failing selftest check first.
`bash scripts/coord-selftest.sh` must exit 0 (191/191),
`bash scripts/board-selftest.sh` must exit 0 (150/150) and
`bash scripts/route-selftest.sh` must exit 0 (69/69).
- English for all code, docs, and commit messages (public repo). Norwegian
trigger aliases in the skill description are deliberate.
- Conventional Commits: `type(scope): description`.
## Commands
- Test: `bash scripts/coord-selftest.sh`, `bash scripts/board-selftest.sh` and
`bash scripts/route-selftest.sh` (or `npm test`, the Node wrapper around all three)
- Hook smoke test: `node hooks/scripts/session-start.mjs` (expects JSON on stdout)
- Board smoke test: `bash scripts/board.sh` (read-only, ~3s over the real tree)
- Briefing smoke test: `bash scripts/board.sh --brief` (read-only, writes
nothing). `brief-nightly.sh` DOES write — it overwrites `$CLAUDE_BRIEF_FILE`
(default `~/.claude/briefing.md`), so point that at a scratch path when
testing. Installed as a launchd agent from `launchd/`, which points at the
SOURCE repo, never the version-pinned plugin cache.
- Route smoke test: `bash scripts/route.sh --path known --verification strong
--reversibility cheap --scope local --rationale x` (writes nothing, instant)
- Sweep smoke test: `bash scripts/coord-sweep.sh` (dry-run is the default, so
this writes nothing; never add `--write` to a smoke test against the real
mailbox)
## Release
Version must agree across: `.claude-plugin/plugin.json`, `package.json`,
README version badge, `skills/coord-send/SKILL.md`, `skills/board/SKILL.md` and
`skills/route/SKILL.md` frontmatter, git tag `vX.Y.Z`, and the catalog `ref` in
`ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json`. Release via
the catalog's `scripts/release-plugin.mjs repo-mailbox` (tag + ref bump
together);
verify with `scripts/check-versions.mjs`. Never hand-edit a ref.
Two things that script does that its dry-run label does not suggest:
`--create-tag` creates AND pushes the tag even without `--write`, and its
closing verification gate runs `check-versions.mjs` over ALL plugins — one
unrelated plugin in ERROR aborts it with the catalog edit written but
uncommitted. When that happens, commit the catalog's `marketplace.json` +
`README.md` by hand and leave every other dirty file in that repo alone.
**`--write --commit` does NOT close this window (confirmed by catalog,
2026-08-10).** The gate (`check-versions.mjs`) runs via `execFileSync` before
the `--commit` conditional, so it throws on any plugin's ERROR — including one
we did not touch — after the catalog files are written and before commit,
regardless of whether `--commit` was passed. Catalog is evaluating a
pre-flight gate (run the check before writing, abort there) but it is **not
implemented yet** — do not assume it exists. Until it ships: before running
`--write`, run `node scripts/check-versions.mjs` in the catalog manually and
confirm 0 ERROR first, even when the only ERROR belongs to an unrelated
plugin. If it still fires mid-release, fall back to the manual-commit
recovery above.
## Hardening roadmap
Empty — the post-v0.1.0 queue (atomic delivery, `.`/`..` rejection,
selftest gaps, uniform `-h`) shipped in v0.2.0; broadcast self-delivery
shipped in v0.2.1; broadcast retraction (`coord-send --retract`) shipped in
v0.4.0, closing the last monotonically-growing surface. `coord-inbox.sh`
still ignores unknown arguments by design (hook context must never fail)
but now warns about each one on stderr, which the hook discards.
Two retraction limits are deliberate, not gaps: it is un-send and never
recall (a repo that already received a broadcast keeps it — the seen set is
delivery history and is left untouched), and the sender check is an accident
guard, not a security boundary, because `--from` redefines identity here as
it does everywhere else in the engine.