repo-mailbox/CLAUDE.md
Kjell Tore Guttormsen c1dabf109d fix(hooks): state-line-guard ratchets against current size, not a flat gate
Advisor review caught this before the v0.23.0 tag landed: the guard
compared the projected line count only against the fixed 60-line max,
never against the file's current size, so trimming an already-oversized
STATE.md (e.g. 156 -> 100 lines, still over 60 but smaller) was denied
exactly like growing it would be.

Verified against the real tree: 23 of the machine's STATE.md files are
already over 60 lines today, one at 1405. Shipped as a flat gate, this
hook would have made most of them un-editable except by a single write
landing at <=60 in one shot -- backwards for a guard meant to make
trimming possible.

Fixed with a ratchet: deny only when the projection is over the max AND
larger than the file's current line count (0 for a file that doesn't
exist yet), for both Write and Edit. A compliant file still cannot grow
past the limit and a new file still cannot be created oversized, but an
oversized file can now be edited toward compliance one write at a time.

state-line-guard-selftest.sh: 16 -> 21 checks (new section 8: shrink
allows, same-size allows, grow-while-oversized still denies, new-oversized
still denies). Suite total: 191 + 152 + 69 + 21 = 433.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186kZGKddxfA9N84HqMLbb2
2026-08-14 17:10:04 +02:00

488 lines
31 KiB
Markdown

# repo-mailbox
Renamed from `coord` in v0.3.0. The plugin/repo is `repo-mailbox`; the CLI
(`coord-send.sh`, `coord-inbox.sh`, `coord-done.sh`), the mailbox root
(`~/.claude/coord/`) and `CLAUDE_COORD_DIR` deliberately kept their names —
they are the transport protocol, not the product.
## Context
Local inter-repo coordination mailbox for Claude Code, packaged as a
marketplace plugin. Three components, one boundary:
- **Engine (`scripts/*.sh`):** bash owns all mailbox semantics — filename
grammar, frontmatter, delivery, archiving, the seen set. `coord-send.sh`
writes, `coord-inbox.sh` reads (formatted for context injection),
`coord-done.sh` archives, `coord-count.sh` counts without delivering,
`coord-sweep.sh` closes the aged FYI backlog machine-wide.
Everything is pinned by `coord-selftest.sh`
(191 checks, throwaway mailbox via `CLAUDE_COORD_DIR`).
**`coord-sweep.sh` is the only path that closes a message with no human in
the loop, and every constraint on it follows from that.** It may close exactly
one mechanically decidable class - `reply-expected: no`, older than the grace
window - because a message that owes a reply can only be answered by a session
in the repo that owes it. Dry-run is the default, inverted from the rest of the
engine, since this is the one script that destroys pending state. It closes
through `coord-done.sh --repo` rather than moving files, so the archive layout
and the `_broadcast` refusal stay in one place. And it logs every closure with
sender and subject, because directed messages have no seen-tracking: the sweep
genuinely cannot tell "seen and ignored" from "never delivered", so a notice
can be closed unread and the log is the only record that it existed. Widening
the class, defaulting to `--write`, or dropping the log each independently
turn this from a bounded cleanup into silent data loss.
**`coord-count.sh` prints TWO integers per mailbox** (`<name>\t<pending>\t<debt>`),
and the first must stay pending: `board.sh` counts the same inbox files
itself, so a debt-only count would put two different numbers under one name.
Debt is read from `reply-expected` in the FRONTMATTER BLOCK ONLY - a body line
is untrusted input and must not be able to silence a debt - and an absent
field means a reply IS owed, because every message written before 0.11.0
lacks it.
**Reading is delivering — counting is not.** `coord-inbox.sh` records a
broadcast as seen once it has printed it, so it can never be used to survey
other repos: doing so would consume each one's backlog silently, and the seen
set is delivery history that retraction deliberately leaves alone.
`coord-count.sh` exists for every "what is pending" question and writes
nothing at all. Any future read-shaped feature belongs there, not in the
read path.
- **Hook (`hooks/scripts/session-start.mjs`):** thin zero-dependency Node
wrapper (marketplace convention: hooks are `.mjs`) that calls
`coord-inbox.sh` and emits the `hookSpecificOutput.additionalContext`
envelope. No mailbox logic lives here. Always exits 0.
- **Hook (`hooks/scripts/pre-state-line-guard.mjs`):** a `PreToolUse` hook on
`Write|Edit` that enforces the STATE.md convention's `maks ~60 linjer`
(global CLAUDE.md) mechanically. It exists because the prose limit alone
failed: a real STATE.md drifted to 155-156 lines before an /insights sweep
of 160 sessions noticed, and one trim pass on it *increased* the line count
instead of shrinking it. org-ops dispatched the work order
(20260814T144553Z) asking for a `PostToolUse` hook — that was the wrong
event, and the fix is not cosmetic: `PostToolUse` fires only after the tool
has already written the file (confirmed against the official hooks docs,
2026-08-14 — "Can block? No", stderr is shown to the model but the write
already landed), so it cannot stop an oversized STATE.md from landing, only
nag about it afterward. `PreToolUse` is the only event that can deny the
call before the file is touched, which is what "enforces" has to mean here.
Denial is stderr + `exit 2`, matching `llm-security`'s
`pre-write-pathguard.mjs` — the only other `PreToolUse` `Write|Edit` guard
in this marketplace — rather than the `hookSpecificOutput.permissionDecision`
JSON form; both block, and matching the sibling convention keeps one idiom
for "block a write" instead of two. For `Write` the projected content is the
call's own `content`; for `Edit` it is the CURRENT on-disk file (read fresh,
since `PreToolUse` fires before the edit is applied) with `old_string`
replaced by `new_string` — every occurrence when `replace_all` is set,
otherwise only the first, mirroring what the real Edit tool does. Getting
`replace_all` wrong in either direction is not a hypothetical: a hook that
only ever replaced the first occurrence would silently pass a bulk edit that
balloons the file, so `state-line-guard-selftest.sh` (21 checks) pins a
fixture where only counting every `replace_all` occurrence produces the
correct denial. Anything the hook cannot project with confidence — a
missing file, an `old_string` that is not present, fields of the wrong
type — is left to the real tool, which reports a clearer error than a guess
here would; the guard only ever touches files named exactly `STATE.md`, at
any depth, matching the same basename rule the global session-start hook's
nearest-STATE-wins search already uses.
**It is a RATCHET against the file's current size, not a flat gate at 60 —
found by advisor review before the tag landed, not by the selftest, which
had no fixture for it.** The first cut compared the projected line count
only against `MAX_LINES`, never against what the file already was, so
trimming an oversized STATE.md from, say, 156 to 100 lines — still over 60,
but strictly smaller — was denied exactly like growing it would have been.
Verified empirically against the real tree (2026-08-14):
`wc -l ~/repos/*/STATE.md ~/repos/*/*/STATE.md | awk '$1 > 60'` found 23
files already over 60 lines, one at 1405. Shipped as a flat gate, this hook
would have made most of the machine's STATE.md files un-editable except by
a single write landing at `<=60` in one shot — backwards for a guard whose
whole point is making the trim the /insights finding asked for actually
possible. The fix reads the file's current line count for BOTH tool types
(previously only `Edit` read the file at all) and denies only when the
projection is over `MAX_LINES` **and** larger than that current count: a
compliant file still cannot grow past the limit, a brand-new file still
cannot be created oversized (current defaults to 0), but an already-oversized
file can always be edited toward compliance, one write at a time, without
ever making it worse. Section 8 of the selftest pins all four cases:
shrink-while-still-over-limit allows, same-size-rewrite allows, grow-an-
already-oversized-file still denies, and create-new-oversized-file still
denies.
- **Board (`scripts/board.sh`):** cross-repo attention board. Reads STATE.md
next-step blocks + board lines, `git status`, and mailbox pending counts, and
prints one line per repo. Read-only by construction: it writes to no repo, no
STATE.md and no mailbox. Pinned by `board-selftest.sh` (152 checks).
**It lives here because the mailbox is one of its three inputs, and it carries
the same axis distinction the mailbox does.** A pending count means *others
are waiting on this repo*; who a repo waits *on* comes only from its board
line, because the message format has no reply-to field. Enforcing that in one
of two repos would not be enforcing it. The operator invokes both `board` and
`coord-send` exclusively through their Skill front doors, never a personal
terminal alias — a claim this file carried until 0.12.1 (`~/.claude/scripts/board.sh`
is a deployed copy the operator's `board()` function points at) did not survive
inspection: no such file ever existed, and `route.sh` had no deployed copy
either. Only the five `coord-*.sh` scripts were ever deployed there, and their
one measured effect was an accidental fallback target for Claude sessions'
own Bash tool calls (the bug 0.12.1 fixed) — once that fallback was gone they
had no remaining function and were deleted.
**`board.sh --brief` is a second RENDERING of that scan, never a second
scan, and `brief-nightly.sh` is the only writer in that path.** The briefing
answers the narrower question an unattended job can answer without judgement:
which repos have an unhandled inbox, what their next step says in full, and
the exact command to start a session in each. It prints NESTE uncut because
the 38-character cut is the table column's property, not the record's — the
value used to be truncated at record-build time, which left the cut string as
the only copy. Each command is derived by CALLING `route.sh` with that repo's
own four traits; `next-cost` alone cannot yield it, since the advisor flag is
a property of the ROW and two rows can share a model/effort pair while
differing on it. A repo with no route line is told so rather than handed a
guess, because a guessed command reads as authoritative.
**`board.sh --plan` is the THIRD rendering, and the only one that takes a
position.** It answers which repos to open a tab for today, in what order,
with which command. The position it takes is the ORDER and nothing else -
there is no cutoff, so the plan hides nothing, and every term is a lookup over
fields the scan already read. Since 0.20.0 it is FIVE ORDERED GROUPS -
chain-root, debt, planned, in-progress, undeclared - ranked within a group by
that group's own quantity, then a Sonnet next-cost, then oldest plan first.
**0.19.0 shipped a weighted score here and 0.20.0 replaced it, and the
objection the score answered is ACCEPTED, not forgotten.** A group order
genuinely cannot express "this repo owes one message and releases two others"
as one quantity; a score could, and that was its point. What a score could not
do was hold still for the second consumer - re-tuning 40 against 15 silently
reorders a parser living in another repo, and no test in THIS repo can catch
that. The operator weighed both and chose the lookup (2026-08-03). Write that
down every time this paragraph is edited: a later session that reads the
objection as an unfixed defect will "restore" the score, and the round trip is
the loop this file exists to stop.
**`planned` ranks ABOVE `in-progress`, inverted at 0.20.0 by operator
decision.** Turning a decision into motion is the slow step; live work is
already moving. Flipping it back is a policy change, not a sort fix.
**Debt is never excluded and never capped, and that is the rule most likely to
be "fixed" into a defect.** Excluding `blocked` or `done` is a claim about a
repo's OWN next step, which by definition cannot be moved, while owing a reply
is the other axis entirely - answering is often what unblocks it. Measured on
the real tree at 0.16.0, two of 26 planned repos were `done` with an unhandled
inbox. A cap on debt was proposed with the chain credit at 0.19.0 and DECLINED
by the operator. Sitting one group below chain-root credit is NOT that cap:
the debtor keeps its tab, its most-owed-first position among the other
debtors, and its `why=inbox:N`. A change that DROPPED a debtor from the plan
would be the declined cap wearing the group order as a disguise, and selftest
section 12 pins both halves - the root outranking four owed messages, and the
debtor keeping everything it had.
**"Debt" means OWED, never raw pending, since 0.22.0 - and this NARROWS what
counts as debt, it does not reopen the paragraph above.** The paragraph above
settles a different question: once a repo has debt, is it ever excluded or
capped (no). This one settles what counts as debt in the first place. Through
0.21.0, group 2's `keep`/`mag`/`why=inbox:N` and `--brief`'s whole "repo som
skylder et svar" listing were computed from the raw pending-file count - every
unhandled message in the inbox, including ones the sender declared
`reply-expected: no`. That is a notice, not a request, and 0.11.0 gave
`coord-count.sh` a second column (`owed`) for exactly this distinction - but
`board.sh` never read it. Reported by morning-driver (2026-08-11) and
independently reproduced against the live mailbox 2026-08-13: 27 of 72
pending messages (37.5%) were notices. The fix joins `--plan` and `--brief`
against `coord-count.sh`'s `owed` column by repo name (same technique as the
chain-root `$UNBLOCKS` join below), so a `done`/`deferred`/`blocked` repo
whose only mail is FYI no longer gets a tab, and `--brief` no longer counts a
notice as an obligation. This reverses a decision from session 41
(2026-08-10) that declined to build this filter, on the premise that "the
arrival of the request IS the admission signal" - a premise that assumed
group 2 already meant requests. It didn't; the code computed pending, the
comments already said "owed" throughout, and the plan's own printed header
("Utelatt naar repoet verken skylder svar...") already claimed the exclusion
was debt-based. The fix makes the code match what its own comments and
header already promised. Pinned by board-selftest.sh section 8/12 fixtures
`repo-done-fyi` (pending 2, owed 0 - excluded) and `repo-blocked-mixed`
(pending 3, owed 2 - planned on 2, not 3). The TABLE's `INN` column and the
raw scan (`RECORDS` field 6) are UNCHANGED - they answer "what is the state
of every repo," not "who is waiting on you," and stay on raw pending by
design.
**The same 0.22.0 patch that switched `n_owe` to OWED also had to fix what
`n_owe == 0` claims.** `--brief`'s empty-debt branch said "Ingen repo har
uhaandtert innboks. Ingen skylder noen et svar i dag." (no repo has
unhandled inbox; nobody owes a reply) - two claims in one branch, and only
the second is what `n_owe == 0` actually proves once `n_owe` means OWED. A
repo can hold FYI-only mail with zero debt, which makes the first sentence
false while it fires - caught in review before release, not by any fixture
(the shared test tree never reaches `n_owe == 0`, since it always carries a
debtor). Fixed to state only the debt claim, and to name any FYI-only
mailboxes found rather than let their existence become invisible again -
the same "labelled, not silently dropped" principle `--plan` already
applies to unknown-status repos. Pinned by board-selftest.sh section 14
with its own isolated root (debt-free, one FYI-only repo).
**The `$UNBLOCKS`/`$RECORDS` join used `NR==FNR` through 0.21.0, and that
idiom silently drops the entire plan whenever the FIRST file is empty - fixed
to `FILENAME==` comparison in 0.22.0, found while adding the `$OWED` join
above.** Verified against the shipped 0.21.0 script: one in-progress repo
with an unhandled inbox message, zero blocked repos anywhere in the tree
(so `$UNBLOCKS` is empty, which is a common, ordinary tree state, not an
edge case) - `--plan` printed "0 tabber". `NR==FNR` is only true for the
FIRST file's own lines; when that file is empty, `FNR` and `NR` stay equal
for the ENTIRE next file too (not just its first line - verified with a
minimal awk reproduction), so every record in it is misrouted into the
`ub[]` branch and dropped via `next`. This was invisible to
board-selftest.sh because the fixture tree has carried at least one
`blocked` repo since the chain-root feature shipped, and it was invisible
on the real tree because `~/repos` currently always has one too - neither
is a guarantee. `FILENAME==UBF`/`FILENAME==OWF` compares the exact path,
never line counts, so an empty lookup file degrades to "nothing matched,"
never to "everything after it is misrouted."
**Chain-root credit lands on the ROOT and nowhere else.** For every `blocked`
repo the `blocked-on` edge is followed transitively to the first repo that is
not itself blocked. Crediting a blocked repo would open a tab that cannot move;
crediting only the direct blocker leaves a two-hop chain's root uncredited,
which is the shape the real tree actually had. A cycle, a `blocked-on` naming
an unscanned repo, and a blocked repo with no target must all credit NOBODY:
inventing a root there produces a plan that looks correct and sends the
operator to the wrong repo.
Repos with no board line rank last and are LABELLED rather than dropped,
because the table already prints a MERK line about them and a plan that
omitted them silently would repeat that defect.
It renders `key=value` blocks, not prose, because it has two consumers: the
operator, and a driver repo consuming the plan. Prose would make the rendered
format an API no test in THIS repo could hold stable for a consumer in
another. `command_missing=` carries both no-command causes (no route line, and
a route line route.sh rejects) because a bare `command=` is the shape of a
runnable command carrying nothing - a driver reading `^command=` would type an
empty line into a live pane. `route_cmd_for()` is the single reader of the
route-line grammar, shared with `--brief`, and distinguishes the two causes by
exit code rather than by an empty string.
**`paste=` and `dir=`/`command=` are the same fact for the two consumers, and
neither is redundant.** A driver moves the pane itself and then types the
command, so it needs them apart; a human needs ONE thing to select. Handing
the operator two fields to join by hand is not a saved output line, it is the
step where a session starts in the wrong repo - and it was measured the moment
the feature met its first user, who could not act on the block at all. `paste=`
is emitted only alongside `command=`: `paste=cd X && ` with nothing after it
would run the cd and then a bare newline, which fails SILENTLY by leaving the
operator in the right directory with no session started.
**Driving a terminal from the plan does NOT belong here, and the measurement
in `docs/ghostty-orchestration-measurement.md` is the argument, not taste.**
It is a version-pinned undocumented composition over a preview API whose
documented path is already broken upstream and whose regression was closed as
not planned, with a blast radius reaching into other repos' live sessions.
None of that is mailbox transport, and none of it may be able to break
`coord-inbox` or `board`. The dependency runs one way: the driver consumes the
plan, the plan never knows a terminal exists.
It also cross-checks itself against `coord-count.sh`, and that is not
belt-and-braces. The repo scan and the mailbox are two different populations:
a mailbox can carry a name no scan will ever produce — a declared non-git
surface (`CLAUDE_COORD_REPO`, e.g. `~/repos` itself) or a checkout outside the
roots. Measured on the real mailbox at 0.15.0: the briefing found 11 repos /
21 messages where `coord-count` saw 12 mailboxes / 22 pending, the missing one
being the declared surface `repos`. A briefing that only walks the scan
answers "who is waiting on you" with a number it quietly knows is short.
**Zero model calls, and that is the load-bearing property, not an
implementation detail.** The operator authenticates by subscription, so a
headless `claude -p` job draws from the same quota pool as interactive work.
Measured against 2.1.220: `--max-budget-usd` DOES bite under subscription auth
(`terminal_reason: budget_exhausted`, exit 1), but it aborts AFTER turn one,
never before it — floor ~0.25 USD-equivalent per turn on `claude-opus-5[1m]`.
It is a runaway brake, not a pre-flight gate. Making the briefing deterministic
removes the question entirely.
`board.sh` stays read-only, which is why the file write lives in the wrapper
instead of behind a `--brief --out FILE` flag. The wrapper renders to a temp
file in the target directory and renames it into place, and treats an EMPTY
render as a FAILED one: board prints nothing at all when its scan roots do not
exist, which is what a mistyped path or a moved home directory looks like, and
a plain `> file` redirect would destroy yesterday's briefing on a bad launchd
environment. A tree where nobody owes anything is a different case — that is a
valid, non-empty briefing saying so, and is written normally.
- **Route (`scripts/route.sh`):** pure calculator for the next session's model
and effort. Takes four scored traits of the next task plus a required
rationale, and prints one block of `key=value` lines: the rubric row, the rule
that fired, the `next-cost` value, a pasteable startup command, the one-row
cheaper fallback, and the STATE.md comment lines. Pinned by
`route-selftest.sh` (69 checks).
**It is here because it is the WRITER for the field `board.sh` already reads.**
`next-cost` had a reader and no writer, so it was hand-typed every session and
drifted into several competing spellings — cleaning the data
could not fix that, because the cause was the missing write path. The row
table is a closed set of six values, so a seventh cannot enter circulation,
and section 6 of the selftest runs the round trip (route emits → board parses)
*inside* one repo rather than across two. `board.sh` itself is untouched: a
calculator that prints to stdout writes nothing, and the session writes
STATE.md.
**The row table is the operator's global rubric, moved here as the single
copy.** It is not a second spec — `board.sh --help` documents the board line's
*grammar* and points here for the *values*. Scoring the traits is judgement
and belongs to the skill; turning scores into a row is a lookup and takes zero
model calls.
**`--advisor opus` is emitted per ROW, on a need, never unconditionally.**
Rows 1-2 always carry it (Sonnet main model, so opus is a capability lift —
and since every fallback is one row cheaper and the cheap rows are Sonnet,
this is what makes the quota fallback safe to take); rows 3-4 only at
`reversibility=costly|one-way` (Opus main model, so it buys peer review where
a mistake is not cheap to undo); rows 5-6 never, because the CLI rejects every
advisor for a Fable main model. The alternative — the global `advisorModel`
setting written by `/advisor` — is what this replaces: it applies to every
session in every repo, which is how it burned quota before. The two triggers
are almost disjoint by construction, since `costly` forces row 3 and `one-way`
forces row 4, so a Sonnet row always has `reversibility=cheap`.
`verification=none` is deliberately NOT a third trigger: beyond the stakes
rule it would only add mistakes that are cheap to reverse, docs sessions
(`known/none/cheap/local`) among them. Section 14 pins the rule and gates the
three CLI facts it rests on against the installed `claude` without spending a
token — advisor validation runs before the empty-prompt check, so `-p ""`
reaches the validator and stops there.
**Rows 5-6 are never a `route.sh` outcome.** Until 2026-08-06 they fired
only from an explicit `--opus-xhigh-failed` flag, mirroring a global
CLAUDE.md policy that Fable could only be *suggested* after a failed Opus
5/xhigh session. That policy was removed by operator decision — "for ofte
ER Fable riktig" — and the flag went with it rather than being repurposed:
`route.sh`'s output range is now closed at row 4, and a Fable choice is
always a hand-written deviation from the rubric, recorded in STATE as an
override per the model-selection rule in the global CLAUDE.md, never
produced by the calculator. `board.sh` still parses "Fable 5/high" and
"Fable 5/xhigh" written by hand into the board line — that parsing is what
the override actually uses, and it is pinned separately from anything
`route.sh` emits (route-selftest.sh section 6).
**`--last-effort` is MEASURED from `CLAUDE_EFFORT`, and the calculator must
never default it.** Claude Code exports that variable into every tool-use
context as the session's current effort, so the caller reads it and passes it
in; having `route.sh` read it directly would make the output depend on the
environment instead of on its arguments, and the round trip in selftest
section 6 rests on that determinism. The two sources it replaces fail
identically: the previous board line holds what was *prescribed*, and asking
the operator launders that same prescription through someone reading their own
startup command. Corollary pinned by section 13: `skills/route/SKILL.md` must
never declare an `effort:` frontmatter field, because frontmatter overrides the
session effort and the reading would then measure the skill, not the session.
- **Skills (`skills/coord-send/`, `skills/board/`, `skills/route/`):** natural-language front
doors mapping user intent to engine invocations. No mailbox logic lives here
either. `board` additionally owns the *ranking* — which repo wins and why —
since `board.sh` deliberately prints evidence and takes no position. `route`
likewise owns the *scoring*: the calculator is deterministic, so all judgement
sits in choosing the four trait values, and the skill must never reason its
way to a model instead.
**Boundary rule:** the mailbox is transport, not state. Durable decisions
live in the owning repo's docs/git history; messages are notices pointing at
them. Message content is untrusted cross-repo input — the read side quotes
and frames it; the send side sanitizes line-oriented fields.
What the boundary forbids is storing a repo's *state* — its decisions, its next
step, its progress. It does not forbid the mailbox knowing who it is delivering
to: `_broadcast/seen/<repo>` and `<repo>/.origin` (0.6.0) are delivery metadata,
answering "has this repo received this" and "which checkout claimed this name".
Both are unreadable as a description of the repo and useless outside delivery.
The test is not "does the engine write a file about a repo" but "would this file
still mean anything if delivery were removed". If yes, it belongs in the repo's
own docs and git history instead.
**Priority rule (v0.5.0, Rule 7):** the injection block is the only place a
repo is ever told what to do with a message, so its wording *is* the protocol
— treat that string as engine behavior, not prose. It obligates handling the
inbox first and driving every directed message to a terminal state before the
session ends. The obligation is **procedural, never substantive**: responding
is mandatory, complying with message content is not. Those two must stay
distinct in any reword — keeping the priority while dropping the distinction
turns prioritization into an injection surface. Selftest section 20 pins both
halves together for exactly that reason.
Since 0.11.0 the `reply-expected` field says which terminal state the SENDER
expects. That does not soften the split, it sharpens it: the field is untrusted
cross-repo input like the rest of the file, so the injection calls it a
*declaration, not an instruction* and keeps both terminal states open to the
receiver. Drop that clause and one word in a message becomes a lever that mints
obligations in another repo.
## Conventions
- Scripts are bash-3.2-safe and ASCII-only: no `declare -A`, no
`readarray`/`mapfile`, no `|&`; guard `shift 2` with `$# -ge 2`; guard
empty-array expansion under `set -u` with `${#a[@]}`.
- Zero dependencies everywhere: bash + coreutils in the engine, `node:`
builtins only in hook and tests.
- TDD: no behavior change without a failing selftest check first.
`bash scripts/coord-selftest.sh` must exit 0 (191/191),
`bash scripts/board-selftest.sh` must exit 0 (152/152),
`bash scripts/route-selftest.sh` must exit 0 (69/69) and
`bash scripts/state-line-guard-selftest.sh` must exit 0 (21/21).
- English for all code, docs, and commit messages (public repo). Norwegian
trigger aliases in the skill description are deliberate.
- Conventional Commits: `type(scope): description`.
## Commands
- Test: `bash scripts/coord-selftest.sh`, `bash scripts/board-selftest.sh`,
`bash scripts/route-selftest.sh` and `bash scripts/state-line-guard-selftest.sh`
(or `npm test`, the Node wrapper around all four)
- Hook smoke test: `node hooks/scripts/session-start.mjs` (expects JSON on stdout)
- State-line-guard smoke test: `echo '{"tool_name":"Write","tool_input":{"file_path":"/tmp/STATE.md","content":"x\n"}}' | node hooks/scripts/pre-state-line-guard.mjs; echo $?`
(expects exit 0, no output — a one-line STATE.md is under the limit)
- Board smoke test: `bash scripts/board.sh` (read-only, ~3s over the real tree)
- Briefing smoke test: `bash scripts/board.sh --brief` (read-only, writes
nothing). `brief-nightly.sh` DOES write — it overwrites `$CLAUDE_BRIEF_FILE`
(default `~/.claude/briefing.md`), so point that at a scratch path when
testing. Installed as a launchd agent from `launchd/`, which points at the
SOURCE repo, never the version-pinned plugin cache.
- Route smoke test: `bash scripts/route.sh --path known --verification strong
--reversibility cheap --scope local --rationale x` (writes nothing, instant)
- Sweep smoke test: `bash scripts/coord-sweep.sh` (dry-run is the default, so
this writes nothing; never add `--write` to a smoke test against the real
mailbox)
## Release
Version must agree across: `.claude-plugin/plugin.json`, `package.json`,
README version badge, `skills/coord-send/SKILL.md`, `skills/board/SKILL.md` and
`skills/route/SKILL.md` frontmatter, git tag `vX.Y.Z`, and the catalog `ref` in
`ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json`. Release via
the catalog's `scripts/release-plugin.mjs repo-mailbox` (tag + ref bump
together);
verify with `scripts/check-versions.mjs`. Never hand-edit a ref.
Two things that script does that its dry-run label does not suggest:
`--create-tag` creates AND pushes the tag even without `--write`, and its
closing verification gate runs `check-versions.mjs` over ALL plugins — one
unrelated plugin in ERROR aborts it with the catalog edit written but
uncommitted. When that happens, commit the catalog's `marketplace.json` +
`README.md` by hand and leave every other dirty file in that repo alone.
**`--write --commit` does NOT close this window (confirmed by catalog,
2026-08-10).** The gate (`check-versions.mjs`) runs via `execFileSync` before
the `--commit` conditional, so it throws on any plugin's ERROR — including one
we did not touch — after the catalog files are written and before commit,
regardless of whether `--commit` was passed. Catalog is evaluating a
pre-flight gate (run the check before writing, abort there) but it is **not
implemented yet** — do not assume it exists. Until it ships: before running
`--write`, run `node scripts/check-versions.mjs` in the catalog manually and
confirm 0 ERROR first, even when the only ERROR belongs to an unrelated
plugin. If it still fires mid-release, fall back to the manual-commit
recovery above.
## Hardening roadmap
Empty — the post-v0.1.0 queue (atomic delivery, `.`/`..` rejection,
selftest gaps, uniform `-h`) shipped in v0.2.0; broadcast self-delivery
shipped in v0.2.1; broadcast retraction (`coord-send --retract`) shipped in
v0.4.0, closing the last monotonically-growing surface. `coord-inbox.sh`
still ignores unknown arguments by design (hook context must never fail)
but now warns about each one on stderr, which the hook discards.
Two retraction limits are deliberate, not gaps: it is un-send and never
recall (a repo that already received a broadcast keeps it — the seen set is
delivery history and is left untouched), and the sender check is an accident
guard, not a security boundary, because `--from` redefines identity here as
it does everywhere else in the engine.