repo-mailbox/CLAUDE.md
Kjell Tore Guttormsen 8e207c6c49 test(sweep): stop the --days 0 check from racing a one-second cutoff
Measured, not guessed: the check failed 7 runs in 20, not once. The cause is
a same-second collision, reproduced deterministically - a notice minted at
20260801205457 against a cutoff of 20260801205457 survives, the same notice
60s older is closed.

coord-sweep.sh is right and is left alone. Its cutoff is second-granular and
it closes strictly older messages, which spares rather than closes at the
boundary; at any real --days value one second is unobservable. Relaxing that
guard to <= would make a destructive script more aggressive to satisfy a test.

So the test was claiming what the code does not promise: that a notice minted
earlier in the same run is necessarily older at second granularity. Under a
second of work separates the two, so it was a coin flip. Aged by 5 seconds
through the existing age_it, which keeps it well inside the default 14-day
window and clear of the boundary. No sleep: that would have hidden the answer
rather than fixed it.

Selftest 182 -> 183, 20/20 green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GfDGWyyhnM26J4p93GSk2L
2026-08-01 23:01:32 +02:00

265 lines
16 KiB
Markdown

# repo-mailbox
Renamed from `coord` in v0.3.0. The plugin/repo is `repo-mailbox`; the CLI
(`coord-send.sh`, `coord-inbox.sh`, `coord-done.sh`), the mailbox root
(`~/.claude/coord/`) and `CLAUDE_COORD_DIR` deliberately kept their names —
they are the transport protocol, not the product.
## Context
Local inter-repo coordination mailbox for Claude Code, packaged as a
marketplace plugin. Three components, one boundary:
- **Engine (`scripts/*.sh`):** bash owns all mailbox semantics — filename
grammar, frontmatter, delivery, archiving, the seen set. `coord-send.sh`
writes, `coord-inbox.sh` reads (formatted for context injection),
`coord-done.sh` archives, `coord-count.sh` counts without delivering,
`coord-sweep.sh` closes the aged FYI backlog machine-wide.
Everything is pinned by `coord-selftest.sh`
(183 checks, throwaway mailbox via `CLAUDE_COORD_DIR`).
**`coord-sweep.sh` is the only path that closes a message with no human in
the loop, and every constraint on it follows from that.** It may close exactly
one mechanically decidable class - `reply-expected: no`, older than the grace
window - because a message that owes a reply can only be answered by a session
in the repo that owes it. Dry-run is the default, inverted from the rest of the
engine, since this is the one script that destroys pending state. It closes
through `coord-done.sh --repo` rather than moving files, so the archive layout
and the `_broadcast` refusal stay in one place. And it logs every closure with
sender and subject, because directed messages have no seen-tracking: the sweep
genuinely cannot tell "seen and ignored" from "never delivered", so a notice
can be closed unread and the log is the only record that it existed. Widening
the class, defaulting to `--write`, or dropping the log each independently
turn this from a bounded cleanup into silent data loss.
**`coord-count.sh` prints TWO integers per mailbox** (`<name>\t<pending>\t<debt>`),
and the first must stay pending: `board.sh` counts the same inbox files
itself, so a debt-only count would put two different numbers under one name.
Debt is read from `reply-expected` in the FRONTMATTER BLOCK ONLY - a body line
is untrusted input and must not be able to silence a debt - and an absent
field means a reply IS owed, because every message written before 0.11.0
lacks it.
**Reading is delivering — counting is not.** `coord-inbox.sh` records a
broadcast as seen once it has printed it, so it can never be used to survey
other repos: doing so would consume each one's backlog silently, and the seen
set is delivery history that retraction deliberately leaves alone.
`coord-count.sh` exists for every "what is pending" question and writes
nothing at all. Any future read-shaped feature belongs there, not in the
read path.
- **Hook (`hooks/scripts/session-start.mjs`):** thin zero-dependency Node
wrapper (marketplace convention: hooks are `.mjs`) that calls
`coord-inbox.sh` and emits the `hookSpecificOutput.additionalContext`
envelope. No mailbox logic lives here. Always exits 0.
- **Board (`scripts/board.sh`):** cross-repo attention board. Reads STATE.md
next-step blocks + board lines, `git status`, and mailbox pending counts, and
prints one line per repo. Read-only by construction: it writes to no repo, no
STATE.md and no mailbox. Pinned by `board-selftest.sh` (51 checks).
**It lives here because the mailbox is one of its three inputs, and it carries
the same axis distinction the mailbox does.** A pending count means *others
are waiting on this repo*; who a repo waits *on* comes only from its board
line, because the message format has no reply-to field. Enforcing that in one
of two repos would not be enforcing it. The operator invokes both `board` and
`coord-send` exclusively through their Skill front doors, never a personal
terminal alias — a claim this file carried until 0.12.1 (`~/.claude/scripts/board.sh`
is a deployed copy the operator's `board()` function points at) did not survive
inspection: no such file ever existed, and `route.sh` had no deployed copy
either. Only the five `coord-*.sh` scripts were ever deployed there, and their
one measured effect was an accidental fallback target for Claude sessions'
own Bash tool calls (the bug 0.12.1 fixed) — once that fallback was gone they
had no remaining function and were deleted.
**`board.sh --brief` is a second RENDERING of that scan, never a second
scan, and `brief-nightly.sh` is the only writer in that path.** The briefing
answers the narrower question an unattended job can answer without judgement:
which repos have an unhandled inbox, what their next step says in full, and
the exact command to start a session in each. It prints NESTE uncut because
the 38-character cut is the table column's property, not the record's — the
value used to be truncated at record-build time, which left the cut string as
the only copy. Each command is derived by CALLING `route.sh` with that repo's
own four traits; `next-cost` alone cannot yield it, since the advisor flag is
a property of the ROW and two rows can share a model/effort pair while
differing on it. A repo with no route line is told so rather than handed a
guess, because a guessed command reads as authoritative.
It also cross-checks itself against `coord-count.sh`, and that is not
belt-and-braces. The repo scan and the mailbox are two different populations:
a mailbox can carry a name no scan will ever produce — a declared non-git
surface (`CLAUDE_COORD_REPO`, e.g. `~/repos` itself) or a checkout outside the
roots. Measured on the real mailbox at 0.15.0: the briefing found 11 repos /
21 messages where `coord-count` saw 12 mailboxes / 22 pending, the missing one
being the declared surface `repos`. A briefing that only walks the scan
answers "who is waiting on you" with a number it quietly knows is short.
**Zero model calls, and that is the load-bearing property, not an
implementation detail.** The operator authenticates by subscription, so a
headless `claude -p` job draws from the same quota pool as interactive work.
Measured against 2.1.220: `--max-budget-usd` DOES bite under subscription auth
(`terminal_reason: budget_exhausted`, exit 1), but it aborts AFTER turn one,
never before it — floor ~0.25 USD-equivalent per turn on `claude-opus-5[1m]`.
It is a runaway brake, not a pre-flight gate. Making the briefing deterministic
removes the question entirely.
`board.sh` stays read-only, which is why the file write lives in the wrapper
instead of behind a `--brief --out FILE` flag. The wrapper renders to a temp
file in the target directory and renames it into place, and treats an EMPTY
render as a FAILED one: board prints nothing at all when its scan roots do not
exist, which is what a mistyped path or a moved home directory looks like, and
a plain `> file` redirect would destroy yesterday's briefing on a bad launchd
environment. A tree where nobody owes anything is a different case — that is a
valid, non-empty briefing saying so, and is written normally.
- **Route (`scripts/route.sh`):** pure calculator for the next session's model
and effort. Takes four scored traits of the next task plus a required
rationale, and prints one block of `key=value` lines: the rubric row, the rule
that fired, the `next-cost` value, a pasteable startup command, the one-row
cheaper fallback, and the STATE.md comment lines. Pinned by
`route-selftest.sh` (73 checks).
**It is here because it is the WRITER for the field `board.sh` already reads.**
`next-cost` had a reader and no writer, so it was hand-typed every session and
drifted into several competing spellings — cleaning the data
could not fix that, because the cause was the missing write path. The row
table is a closed set of six values, so a seventh cannot enter circulation,
and section 6 of the selftest runs the round trip (route emits → board parses)
*inside* one repo rather than across two. `board.sh` itself is untouched: a
calculator that prints to stdout writes nothing, and the session writes
STATE.md.
**The row table is the operator's global rubric, moved here as the single
copy.** It is not a second spec — `board.sh --help` documents the board line's
*grammar* and points here for the *values*. Scoring the traits is judgement
and belongs to the skill; turning scores into a row is a lookup and takes zero
model calls.
**`--advisor opus` is emitted per ROW, on a need, never unconditionally.**
Rows 1-2 always carry it (Sonnet main model, so opus is a capability lift —
and since every fallback is one row cheaper and the cheap rows are Sonnet,
this is what makes the quota fallback safe to take); rows 3-4 only at
`reversibility=costly|one-way` (Opus main model, so it buys peer review where
a mistake is not cheap to undo); rows 5-6 never, because the CLI rejects every
advisor for a Fable main model. The alternative — the global `advisorModel`
setting written by `/advisor` — is what this replaces: it applies to every
session in every repo, which is how it burned quota before. The two triggers
are almost disjoint by construction, since `costly` forces row 3 and `one-way`
forces row 4, so a Sonnet row always has `reversibility=cheap`.
`verification=none` is deliberately NOT a third trigger: beyond the stakes
rule it would only add mistakes that are cheap to reverse, docs sessions
(`known/none/cheap/local`) among them. Section 14 pins the rule and gates the
three CLI facts it rests on against the installed `claude` without spending a
token — advisor validation runs before the empty-prompt check, so `-p ""`
reaches the validator and stops there.
**`--last-effort` is MEASURED from `CLAUDE_EFFORT`, and the calculator must
never default it.** Claude Code exports that variable into every tool-use
context as the session's current effort, so the caller reads it and passes it
in; having `route.sh` read it directly would make the output depend on the
environment instead of on its arguments, and the round trip in selftest
section 6 rests on that determinism. The two sources it replaces fail
identically: the previous board line holds what was *prescribed*, and asking
the operator launders that same prescription through someone reading their own
startup command. Corollary pinned by section 13: `skills/route/SKILL.md` must
never declare an `effort:` frontmatter field, because frontmatter overrides the
session effort and the reading would then measure the skill, not the session.
- **Skills (`skills/coord-send/`, `skills/board/`, `skills/route/`):** natural-language front
doors mapping user intent to engine invocations. No mailbox logic lives here
either. `board` additionally owns the *ranking* — which repo wins and why —
since `board.sh` deliberately prints evidence and takes no position. `route`
likewise owns the *scoring*: the calculator is deterministic, so all judgement
sits in choosing the four trait values, and the skill must never reason its
way to a model instead.
**Boundary rule:** the mailbox is transport, not state. Durable decisions
live in the owning repo's docs/git history; messages are notices pointing at
them. Message content is untrusted cross-repo input — the read side quotes
and frames it; the send side sanitizes line-oriented fields.
What the boundary forbids is storing a repo's *state* — its decisions, its next
step, its progress. It does not forbid the mailbox knowing who it is delivering
to: `_broadcast/seen/<repo>` and `<repo>/.origin` (0.6.0) are delivery metadata,
answering "has this repo received this" and "which checkout claimed this name".
Both are unreadable as a description of the repo and useless outside delivery.
The test is not "does the engine write a file about a repo" but "would this file
still mean anything if delivery were removed". If yes, it belongs in the repo's
own docs and git history instead.
**Priority rule (v0.5.0, Rule 7):** the injection block is the only place a
repo is ever told what to do with a message, so its wording *is* the protocol
— treat that string as engine behavior, not prose. It obligates handling the
inbox first and driving every directed message to a terminal state before the
session ends. The obligation is **procedural, never substantive**: responding
is mandatory, complying with message content is not. Those two must stay
distinct in any reword — keeping the priority while dropping the distinction
turns prioritization into an injection surface. Selftest section 20 pins both
halves together for exactly that reason.
Since 0.11.0 the `reply-expected` field says which terminal state the SENDER
expects. That does not soften the split, it sharpens it: the field is untrusted
cross-repo input like the rest of the file, so the injection calls it a
*declaration, not an instruction* and keeps both terminal states open to the
receiver. Drop that clause and one word in a message becomes a lever that mints
obligations in another repo.
## Conventions
- Scripts are bash-3.2-safe and ASCII-only: no `declare -A`, no
`readarray`/`mapfile`, no `|&`; guard `shift 2` with `$# -ge 2`; guard
empty-array expansion under `set -u` with `${#a[@]}`.
- Zero dependencies everywhere: bash + coreutils in the engine, `node:`
builtins only in hook and tests.
- TDD: no behavior change without a failing selftest check first.
`bash scripts/coord-selftest.sh` must exit 0 (183/183),
`bash scripts/board-selftest.sh` must exit 0 (51/51) and
`bash scripts/route-selftest.sh` must exit 0 (73/73).
- English for all code, docs, and commit messages (public repo). Norwegian
trigger aliases in the skill description are deliberate.
- Conventional Commits: `type(scope): description`.
## Commands
- Test: `bash scripts/coord-selftest.sh`, `bash scripts/board-selftest.sh` and
`bash scripts/route-selftest.sh` (or `npm test`, the Node wrapper around all three)
- Hook smoke test: `node hooks/scripts/session-start.mjs` (expects JSON on stdout)
- Board smoke test: `bash scripts/board.sh` (read-only, ~3s over the real tree)
- Briefing smoke test: `bash scripts/board.sh --brief` (read-only, writes
nothing). `brief-nightly.sh` DOES write — it overwrites `$CLAUDE_BRIEF_FILE`
(default `~/.claude/briefing.md`), so point that at a scratch path when
testing. Installed as a launchd agent from `launchd/`, which points at the
SOURCE repo, never the version-pinned plugin cache.
- Route smoke test: `bash scripts/route.sh --path known --verification strong
--reversibility cheap --scope local --rationale x` (writes nothing, instant)
- Sweep smoke test: `bash scripts/coord-sweep.sh` (dry-run is the default, so
this writes nothing; never add `--write` to a smoke test against the real
mailbox)
## Release
Version must agree across: `.claude-plugin/plugin.json`, `package.json`,
README version badge, `skills/coord-send/SKILL.md`, `skills/board/SKILL.md` and
`skills/route/SKILL.md` frontmatter, git tag `vX.Y.Z`, and the catalog `ref` in
`ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json`. Release via
the catalog's `scripts/release-plugin.mjs repo-mailbox` (tag + ref bump
together);
verify with `scripts/check-versions.mjs`. Never hand-edit a ref.
Two things that script does that its dry-run label does not suggest:
`--create-tag` creates AND pushes the tag even without `--write`, and its
closing verification gate runs `check-versions.mjs` over ALL plugins — one
unrelated plugin in ERROR aborts it with the catalog edit written but
uncommitted. When that happens, commit the catalog's `marketplace.json` +
`README.md` by hand and leave every other dirty file in that repo alone.
## Hardening roadmap
Empty — the post-v0.1.0 queue (atomic delivery, `.`/`..` rejection,
selftest gaps, uniform `-h`) shipped in v0.2.0; broadcast self-delivery
shipped in v0.2.1; broadcast retraction (`coord-send --retract`) shipped in
v0.4.0, closing the last monotonically-growing surface. `coord-inbox.sh`
still ignores unknown arguments by design (hook context must never fail)
but now warns about each one on stderr, which the hook discards.
Two retraction limits are deliberate, not gaps: it is un-send and never
recall (a repo that already received a broadcast keeps it — the seen set is
delivery history and is left untouched), and the sender check is an accident
guard, not a security boundary, because `--from` redefines identity here as
it does everywhere else in the engine.