ORDRE 59. A dispatched order used to live only in a scratch prompt file
passed through argv, so it died with the pane it was typed into. Measured
2026-08-17: one order was dispatched three times over 90 minutes before it
was worked, because the first two tabs ran something else and the order
left no trace in the receiving repo at all.
New channel `~/.claude/coord/<repo>/orders/`, beside `inbox/` and never
merged with it. The axis is authorization: inbox content is untrusted
cross-repo data that may never instruct a session (Rule 6), a dispatch
order is operator-authorized work by construction. One channel carrying
both classes would mean either mail that can instruct or orders that
cannot, so the infrastructure is reused and the channel is not.
Four one-verb engines: coord-order-send.sh (write), coord-order-inbox.sh
(read, writes nothing at all), coord-order-claim.sh (atomic claim),
coord-order-done.sh (executed with a commit pointer / --no-commit with a
reason / --return with a reason).
The claim is a rename with no check-then-act step, so of N racing sessions
exactly one finds the source and the rest get ENOENT. The test that proves
it spawns 20 claimers BARRIERED on a start flag - unbarriered children do
not race at all - and runs the identical harness against a deliberately
racy `[ -e src ] && cp && rm` as a known-negative control, which must
produce many winners. Without that control, "exactly one winner" is
indistinguishable from "the race never happened".
Channel separation is pinned structurally, not only behaviourally: no mail
script may contain the string `orders`, with a known-positive control
proving the grep can find. coord-done cannot archive an order and
coord-order-claim cannot claim a message.
board gains an ORDRE column beside INN, counted with the identical idiom
and never summed with it: INN is "others are waiting on YOU", ORDRE is
"work is waiting on this REPO". Claimed orders are excluded - the column
answers what a session can pick up. board.sh --dispatch --order-id emits a
thin starter carrying only the id and the four steps, so the order text has
exactly one home; the id is validated shell-clean and must be pending in
the target's queue.
SessionStart injects the queue as its own block below the mailbox block.
Two channels, two blocks, mail first: it carries Rule 7, and the queue
order is mail -> orders -> STATE's NESTE.
Also folds in dde392d (board prefix-match fix), which landed after the
0.26.0 bump and before any tag. v0.26.0 was never tagged, so 0.27.0 is the
release that carries all of it.
Suites: coord 220, board 237, route 69, orders 97, guard 40; npm test 11/11.
Antakelse 4 (atomic claim) and antakelse 6 (morning --plan-file --dry-run
reports 1 of 1 for the thin starter) both measured, not assumed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0134iB7ipXGgEpv9imYoVmr2
770 lines
50 KiB
Markdown
770 lines
50 KiB
Markdown
# repo-mailbox
|
|
|
|
Renamed from `coord` in v0.3.0. The plugin/repo is `repo-mailbox`; the CLI
|
|
(`coord-send.sh`, `coord-inbox.sh`, `coord-done.sh`), the mailbox root
|
|
(`~/.claude/coord/`) and `CLAUDE_COORD_DIR` deliberately kept their names —
|
|
they are the transport protocol, not the product.
|
|
|
|
## Context
|
|
|
|
Local inter-repo coordination mailbox for Claude Code, packaged as a
|
|
marketplace plugin. Three components, one boundary:
|
|
|
|
- **Engine (`scripts/*.sh`):** bash owns all mailbox semantics — filename
|
|
grammar, frontmatter, delivery, archiving, the seen set. `coord-send.sh`
|
|
writes, `coord-inbox.sh` reads (formatted for context injection),
|
|
`coord-done.sh` archives, `coord-count.sh` counts without delivering,
|
|
`coord-sweep.sh` closes the aged FYI backlog machine-wide.
|
|
Everything is pinned by `coord-selftest.sh`
|
|
(220 checks, throwaway mailbox via `CLAUDE_COORD_DIR`).
|
|
|
|
**`ktg-plugin-marketplace` is a RETIRED `--to` address (operator decision
|
|
2026-08-15), rejected rather than redirected.** It is a polyrepo directory,
|
|
not a git repo, so `basename(git toplevel)` can never resolve to it and no
|
|
session was ever able to hold that identity naturally - mail for it belongs
|
|
to `catalog` instead. A silent redirect was considered and declined: it
|
|
delivers mail somewhere the sender does not believe it landed, which is the
|
|
same misdelivery defect this closes a second time (2 messages sat
|
|
undelivered 2 days on this exact misaddressing before `catalog`'s H4 count
|
|
caught 6 more). Rejection fails loud at the sender, at the moment the
|
|
mistake is made. Only `--to` is retired, not `--from` - the defect was mail
|
|
*arriving* there, never mail claiming to *originate* there.
|
|
|
|
**`coord-send --reply-to` asserts "marked handled" against GROUND TRUTH, and
|
|
the exit code alone is NOT that ground truth.** The line used to print
|
|
unconditionally with `coord-done`'s output discarded (`>/dev/null 2>&1`),
|
|
so the one line a session relies on to close a reply debt was false at the
|
|
moment it was printed — measured with a stub `coord-done` exiting 1: original
|
|
still in the inbox, no `archive/`, exit 0, "marked handled". Every reply this
|
|
repo sent had to be verified by hand afterwards, which is what a false
|
|
success in the TRANSPORT costs. The check is `exit 0` **and** the original no
|
|
longer being at `$COORD/$FROM/inbox/$REPLYTO`, because `coord-done` exits 0
|
|
when it archives nothing (an unknown name is idempotently fine by its own
|
|
contract), so a nonzero-exit test still certifies a message that never moved.
|
|
The path is recomputed rather than reusing `$REPLY_ORIG`, which resolves to
|
|
the inbox OR the archive — replying to an already-archived original moves
|
|
nothing and must not warn. Failure is exit **1**, a new status: the reply WAS
|
|
delivered and re-sending would duplicate it, so 2 stays the
|
|
nothing-was-written status it has always been. Selftest section 34 pins all
|
|
four cases (fails outright / exits 0 without moving / real happy path /
|
|
archive-path reply).
|
|
|
|
**`coord-sweep.sh` is the only path that closes a message with no human in
|
|
the loop, and every constraint on it follows from that.** It may close exactly
|
|
one mechanically decidable class - `reply-expected: no`, older than the grace
|
|
window - because a message that owes a reply can only be answered by a session
|
|
in the repo that owes it. Dry-run is the default, inverted from the rest of the
|
|
engine, since this is the one script that destroys pending state. It closes
|
|
through `coord-done.sh --repo` rather than moving files, so the archive layout
|
|
and the `_broadcast` refusal stay in one place. And it logs every closure with
|
|
sender and subject, because directed messages have no seen-tracking: the sweep
|
|
genuinely cannot tell "seen and ignored" from "never delivered", so a notice
|
|
can be closed unread and the log is the only record that it existed. Widening
|
|
the class, defaulting to `--write`, or dropping the log each independently
|
|
turn this from a bounded cleanup into silent data loss.
|
|
|
|
**`coord-count.sh` prints TWO integers per mailbox** (`<name>\t<pending>\t<debt>`),
|
|
and the first must stay pending: `board.sh` counts the same inbox files
|
|
itself, so a debt-only count would put two different numbers under one name.
|
|
Debt is read from `reply-expected` in the FRONTMATTER BLOCK ONLY - a body line
|
|
is untrusted input and must not be able to silence a debt - and an absent
|
|
field means a reply IS owed, because every message written before 0.11.0
|
|
lacks it.
|
|
|
|
**Reading is delivering — counting is not.** `coord-inbox.sh` records a
|
|
broadcast as seen once it has printed it, so it can never be used to survey
|
|
other repos: doing so would consume each one's backlog silently, and the seen
|
|
set is delivery history that retraction deliberately leaves alone.
|
|
`coord-count.sh` exists for every "what is pending" question and writes
|
|
nothing at all. Any future read-shaped feature belongs there, not in the
|
|
read path.
|
|
- **Order queue (`scripts/coord-order-*.sh`):** a SECOND channel beside the
|
|
mailbox, `~/.claude/coord/<repo>/orders/`, with four one-verb scripts —
|
|
`coord-order-send.sh` (write), `coord-order-inbox.sh` (read for injection),
|
|
`coord-order-claim.sh` (claim), `coord-order-done.sh` (terminal state).
|
|
Pinned by `orders-selftest.sh` (97 checks).
|
|
|
|
**It is a separate CHANNEL, not more mail, and the axis is authorization.**
|
|
Inbox content is untrusted cross-repo data that may never instruct a session
|
|
(Rule 6); a dispatch order is operator-authorized work by construction —
|
|
dispatch IS the operator's authorization. One channel carrying both classes
|
|
would mean either mail that can instruct or orders that cannot, so the
|
|
infrastructure is reused and the channel is not. The separation is pinned
|
|
STRUCTURALLY, not only behaviourally: no mail script may contain the string
|
|
`orders` at all (with a known-positive control proving the grep can find),
|
|
`coord-done` cannot archive an order, `coord-order-claim` cannot claim a
|
|
message. A behavioural test alone samples one case; the claim is that no
|
|
write path exists.
|
|
|
|
**What the queue does NOT claim is that dispatch is its only writer.**
|
|
`--from` redefines identity here exactly as it does in `coord-send.sh`, so
|
|
any session can write an order into any repo's queue. The authority rests on
|
|
a CONVENTION about who writes, not on enforcement, and the injected text says
|
|
so in those words. Asserting the guarantee instead would be this repo's
|
|
oldest defect class (`--reply-to` claiming "marked handled" against the call
|
|
rather than the world).
|
|
|
|
**The claim is a `mv` with NO check-then-act step, and the test that proves
|
|
it is built to be able to fail.** `rename(2)` is atomic, so of N racing
|
|
processes exactly one finds the source and the rest get ENOENT — the source
|
|
is the contended resource, not any lock. The selftest spawns 20 claimers
|
|
BARRIERED on a start flag, because unbarriered children do not race at all
|
|
(the first finishes before the second starts) and the test would go green
|
|
having proven nothing. It then runs the identical harness against a
|
|
deliberately racy `[ -e src ] && cp && rm`, which must produce MANY winners.
|
|
Deleting that known-negative control turns the whole section back into an
|
|
unmeasured assumption wearing a passing test — this is the assumption the
|
|
design document marked RISIKO, and the control is what closes it.
|
|
|
|
**Claimed orders stay visible in the read path.** A session that claims an
|
|
order and dies is the only remaining way an order can evaporate, so the
|
|
injection keeps showing claimed orders with their in-flight age and the
|
|
command that returns them. Visible-again, NOT a lease timer: nothing here
|
|
expires anything, and building expiry would make the engine decide that a
|
|
session is dead, which it cannot know.
|
|
|
|
**`--return` costs a stated reason, and so does `--no-commit`.** A return
|
|
with no reason is a silent drop with extra steps; a `--no-commit` with no
|
|
reason is "trust me" and would quietly become the cheapest way to close any
|
|
order. The reason is written INTO the order, so the next session that picks
|
|
it up sees why the last one put it down.
|
|
- **Hook (`hooks/scripts/session-start.mjs`):** thin zero-dependency Node
|
|
wrapper (marketplace convention: hooks are `.mjs`) that calls
|
|
`coord-inbox.sh` AND `coord-order-inbox.sh` and emits the
|
|
`hookSpecificOutput.additionalContext` envelope. No mailbox logic lives here.
|
|
Always exits 0.
|
|
|
|
**Two channels, two blocks, never merged, mail first.** Each engine owns the
|
|
words its own block is read under; concatenating them, or letting the wrapper
|
|
write a shared header, would put both authorization classes under one framing
|
|
— the exact thing the channel split exists to prevent. Mail goes above the
|
|
queue because the mail block carries Rule 7 and because the convention's
|
|
queue order is mail -> orders -> STATE's NESTE, so printing the queue first
|
|
would invert on the page the order the two are to be worked in. Each engine
|
|
is run in its own try: an order queue that stayed invisible because the
|
|
mailbox threw would be precisely the silent evaporation the queue exists to
|
|
stop.
|
|
- **Hook (`hooks/scripts/pre-state-line-guard.mjs`):** a `PreToolUse` hook on
|
|
`Write|Edit` that enforces the STATE.md convention's `maks ~120 linjer`
|
|
(global CLAUDE.md; raised from `~60` by operator decision 2026-08-14 —
|
|
see the dated paragraph below) mechanically. It exists because the prose limit alone
|
|
failed: a real STATE.md drifted to 155-156 lines before an /insights sweep
|
|
of 160 sessions noticed, and one trim pass on it *increased* the line count
|
|
instead of shrinking it. org-ops dispatched the work order
|
|
(20260814T144553Z) asking for a `PostToolUse` hook — that was the wrong
|
|
event, and the fix is not cosmetic: `PostToolUse` fires only after the tool
|
|
has already written the file (confirmed against the official hooks docs,
|
|
2026-08-14 — "Can block? No", stderr is shown to the model but the write
|
|
already landed), so it cannot stop an oversized STATE.md from landing, only
|
|
nag about it afterward. `PreToolUse` is the only event that can deny the
|
|
call before the file is touched, which is what "enforces" has to mean here.
|
|
Denial is stderr + `exit 2`, matching `llm-security`'s
|
|
`pre-write-pathguard.mjs` — the only other `PreToolUse` `Write|Edit` guard
|
|
in this marketplace — rather than the `hookSpecificOutput.permissionDecision`
|
|
JSON form; both block, and matching the sibling convention keeps one idiom
|
|
for "block a write" instead of two. For `Write` the projected content is the
|
|
call's own `content`; for `Edit` it is the CURRENT on-disk file (read fresh,
|
|
since `PreToolUse` fires before the edit is applied) with `old_string`
|
|
replaced by `new_string` — every occurrence when `replace_all` is set,
|
|
otherwise only the first, mirroring what the real Edit tool does. Getting
|
|
`replace_all` wrong in either direction is not a hypothetical: a hook that
|
|
only ever replaced the first occurrence would silently pass a bulk edit that
|
|
balloons the file, so `state-line-guard-selftest.sh` (21 checks) pins a
|
|
fixture where only counting every `replace_all` occurrence produces the
|
|
correct denial. Anything the hook cannot project with confidence — a
|
|
missing file, an `old_string` that is not present, fields of the wrong
|
|
type — is left to the real tool, which reports a clearer error than a guess
|
|
here would; the guard only ever touches files named exactly `STATE.md`, at
|
|
any depth, matching the same basename rule the global session-start hook's
|
|
nearest-STATE-wins search already uses.
|
|
|
|
**It is a RATCHET against the file's current size, not a flat gate at 60 —
|
|
found by advisor review before the tag landed, not by the selftest, which
|
|
had no fixture for it.** The first cut compared the projected line count
|
|
only against `MAX_LINES`, never against what the file already was, so
|
|
trimming an oversized STATE.md from, say, 156 to 100 lines — still over 60,
|
|
but strictly smaller — was denied exactly like growing it would have been.
|
|
Verified empirically against the real tree (2026-08-14):
|
|
`wc -l ~/repos/*/STATE.md ~/repos/*/*/STATE.md | awk '$1 > 60'` found 23
|
|
files already over 60 lines, one at 1405. Shipped as a flat gate, this hook
|
|
would have made most of the machine's STATE.md files un-editable except by
|
|
a single write landing at `<=60` in one shot — backwards for a guard whose
|
|
whole point is making the trim the /insights finding asked for actually
|
|
possible. The fix reads the file's current line count for BOTH tool types
|
|
(previously only `Edit` read the file at all) and denies only when the
|
|
projection is over `MAX_LINES` **and** larger than that current count: a
|
|
compliant file still cannot grow past the limit, a brand-new file still
|
|
cannot be created oversized (current defaults to 0), but an already-oversized
|
|
file can always be edited toward compliance, one write at a time, without
|
|
ever making it worse. Section 8 of the selftest pins all four cases:
|
|
shrink-while-still-over-limit allows, same-size-rewrite allows, grow-an-
|
|
already-oversized-file still denies, and create-new-oversized-file still
|
|
denies.
|
|
|
|
**`MAX_LINES` raised 60 -> 120, operator decision 2026-08-14 (evening),
|
|
reported via coord by `.claude` after the global CLAUDE.md prose was
|
|
already updated.** The ratchet mechanics above are unchanged - only the
|
|
constant moved, plus every selftest fixture and boundary value that
|
|
encoded 60 as a literal. Re-verified empirically against the real tree at
|
|
the new threshold (2026-08-14): `wc -l ~/repos/*/STATE.md
|
|
~/repos/*/*/STATE.md | awk '$1 > 120'` found 13 files already over 120
|
|
lines, one at 1496 - the historical 23-files-over-60/one-at-1405 figures
|
|
above describe the tree as it was at the moment the ratchet bug was found,
|
|
not the current threshold, and are left as-is rather than rewritten.
|
|
`session-start.mjs`'s 160-line injection window still covers the new
|
|
120-line limit with room to spare, so no change was needed there.
|
|
|
|
**The Edit path's `current.replace(oldStr, newStr)` was a dollar-pattern
|
|
injection bug, found and fixed 2026-08-15.** Passing `newStr` as a STRING
|
|
makes JavaScript interpret `$`-sequences inside it ($&, `` $` ``, `$'`,
|
|
`$$`, `$n`) as special replacement patterns, even though `oldStr` (the
|
|
search side) is a plain string, not a RegExp. A `new_string` documenting
|
|
old backtick-substitution style (`` $`cmd` ``) - exactly the prose a
|
|
STATE.md's shell-conventions section writes routinely - triggers it.
|
|
Measured against the real bug (`.claude/STATE.md`): a 5-line addition on a
|
|
112-line file projected to 219 lines and was wrongly denied. Direction is
|
|
always fail-CLOSED (never fail-open: it can only over-block, never
|
|
under-block a real oversize), but it made exactly the kind of STATE.md
|
|
that documents shell conventions hard to edit. Fix: replace with a
|
|
function, `current.replace(oldStr, () => newStr)` - a function result is
|
|
never pattern-substituted, so this covers every `$`-sequence at once, not
|
|
a `` $` ``-specific escape. The `replace_all` branch (`split`/`join`) was
|
|
never affected - `join` does not interpret its argument as a pattern.
|
|
Pinned by state-line-guard-selftest.sh section 9 (`$\`` as the real repro,
|
|
`$&` as a second sequence proving the fix is general).
|
|
|
|
**Since ORDRE 42 (operator, 2026-08-16) it carries a SECOND invariant: the
|
|
projected content may not claim `status=done` in its board line while the
|
|
repo holds commits the branch's upstream does not have.** Measured that day:
|
|
two sessions had their push refused by the UFW rate limit on port 22, said so
|
|
honestly in the coord inbox, and wrote `status=done` regardless - board line
|
|
green, one commit unpushed, published surface 404. `done` meant "the session
|
|
finished" where every reader takes it to mean "the work landed", and since
|
|
`done` drops a repo from the board plan, `morning --say <repo>` could not
|
|
reach either of them: one defect hid the other.
|
|
|
|
**The order recommended a session-end hook and that direction does not exist
|
|
in the form it assumes - measured against the official hooks docs, not
|
|
reasoned.** `Stop` fires "once per turn", not once when the session ends, with
|
|
no signal marking the last turn; its exit 2 "prevents Claude from stopping,
|
|
continues the conversation", so a repo that genuinely cannot push (the very
|
|
rate limit that caused the incident) would get a session that will not end.
|
|
`SessionEnd` is the once-per-session event and cannot block at all
|
|
("Can block? No" - exit 2 "shows stderr to user only"), which is the
|
|
after-the-fact nagging the order explicitly refused. Warn-on-write plus
|
|
deny-at-session-end inherits the broken half and buys nothing. So the deny
|
|
sits on the write, where the false claim is actually made.
|
|
|
|
**The false-positive trap is real but bounded, and the deny is escapable by
|
|
telling the truth.** STATE.md is written BEFORE the session's final commit, so
|
|
a session that batches its pushes does hold unpushed commits at that moment -
|
|
but the global git rule already requires a push immediately after every
|
|
commit, and the real tree bears that out (2026-08-16: 43 of 44 repos carrying
|
|
a STATE.md had nothing unpushed; the one exception was `status=blocked` and
|
|
honest). `status=blocked` and `status=in-progress` stay writable in the same
|
|
single edit, so a session that cannot push is never wedged - only stopped from
|
|
claiming otherwise.
|
|
|
|
**No ratchet here, unlike the line limit above, and the asymmetry is the
|
|
reason.** An oversized file needs many writes to come back under the limit, so
|
|
denying the intermediate steps would make trimming impossible; a false `done`
|
|
is corrected by changing one token in the write already being made. A
|
|
"deny only the transition into done" variant was rejected outright: the
|
|
common shape is a repo that ended `done` last session and writes `done` again
|
|
this session, which such a rule waves straight through. It **fails OPEN** on
|
|
every git uncertainty - no upstream, detached HEAD, missing remote-tracking
|
|
ref, not a repo, git slow or absent - because 8 of those 44 repos have no
|
|
upstream at all (one already `status=done`), and a confident denial resting on
|
|
a measurement that never happened is the worse error. The board line is
|
|
selected with `board.sh`'s own anchor (`^<!-- board:`) and the status token
|
|
compared exactly, so prose saying `status=done` never triggers it (a STATE.md
|
|
documenting this guard writes that string routinely) and `done2` is not `done`
|
|
here - and, since F3 (see `board.sh` below), not on `board.sh` either
|
|
anymore: both now do the same exact match on the same extracted value, so
|
|
`done2` reads as MALFORMED in both places, never as `done` in one and
|
|
something else in the other.
|
|
Selftest section 10, 17 checks, including the mandatory known-positive: a
|
|
`status=done` with everything pushed must still go through, or the guard is a
|
|
gate that denies everything and proves nothing.
|
|
- **Board (`scripts/board.sh`):** cross-repo attention board. Reads STATE.md
|
|
next-step blocks + board lines, `git status`, and mailbox pending counts, and
|
|
prints one line per repo. Read-only by construction: it writes to no repo, no
|
|
STATE.md and no mailbox. Pinned by `board-selftest.sh` (217 checks).
|
|
|
|
**It lives here because the mailbox is one of its three inputs, and it carries
|
|
the same axis distinction the mailbox does.** A pending count means *others
|
|
are waiting on this repo*; who a repo waits *on* comes only from its board
|
|
line, because the message format has no reply-to field. Enforcing that in one
|
|
of two repos would not be enforcing it. The operator invokes both `board` and
|
|
`coord-send` exclusively through their Skill front doors, never a personal
|
|
terminal alias — a claim this file carried until 0.12.1 (`~/.claude/scripts/board.sh`
|
|
is a deployed copy the operator's `board()` function points at) did not survive
|
|
inspection: no such file ever existed, and `route.sh` had no deployed copy
|
|
either. Only the five `coord-*.sh` scripts were ever deployed there, and their
|
|
one measured effect was an accidental fallback target for Claude sessions'
|
|
own Bash tool calls (the bug 0.12.1 fixed) — once that fallback was gone they
|
|
had no remaining function and were deleted.
|
|
**`board.sh --brief` is a second RENDERING of that scan, never a second
|
|
scan, and `brief-nightly.sh` is the only writer in that path.** The briefing
|
|
answers the narrower question an unattended job can answer without judgement:
|
|
which repos have an unhandled inbox, what their next step says in full, and
|
|
the exact command to start a session in each. It prints NESTE uncut because
|
|
the 38-character cut is the table column's property, not the record's — the
|
|
value used to be truncated at record-build time, which left the cut string as
|
|
the only copy. Each command is derived by CALLING `route.sh` with that repo's
|
|
own four traits; `next-cost` alone cannot yield it, since the advisor flag is
|
|
a property of the ROW and two rows can share a model/effort pair while
|
|
differing on it. A repo with no route line is told so rather than handed a
|
|
guess, because a guessed command reads as authoritative.
|
|
|
|
**`board.sh --plan` is the THIRD rendering, and the only one that takes a
|
|
position.** It answers which repos to open a tab for today, in what order,
|
|
with which command. The position it takes is the ORDER and nothing else -
|
|
there is no cutoff, so the plan hides nothing, and every term is a lookup over
|
|
fields the scan already read. Since 0.20.0 it is FIVE ORDERED GROUPS -
|
|
chain-root, debt, planned, in-progress, undeclared - ranked within a group by
|
|
that group's own quantity, then a Sonnet next-cost, then oldest plan first.
|
|
|
|
**0.19.0 shipped a weighted score here and 0.20.0 replaced it, and the
|
|
objection the score answered is ACCEPTED, not forgotten.** A group order
|
|
genuinely cannot express "this repo owes one message and releases two others"
|
|
as one quantity; a score could, and that was its point. What a score could not
|
|
do was hold still for the second consumer - re-tuning 40 against 15 silently
|
|
reorders a parser living in another repo, and no test in THIS repo can catch
|
|
that. The operator weighed both and chose the lookup (2026-08-03). Write that
|
|
down every time this paragraph is edited: a later session that reads the
|
|
objection as an unfixed defect will "restore" the score, and the round trip is
|
|
the loop this file exists to stop.
|
|
|
|
**`planned` ranks ABOVE `in-progress`, inverted at 0.20.0 by operator
|
|
decision.** Turning a decision into motion is the slow step; live work is
|
|
already moving. Flipping it back is a policy change, not a sort fix.
|
|
|
|
**Debt is never excluded and never capped, and that is the rule most likely to
|
|
be "fixed" into a defect.** Excluding `blocked` or `done` is a claim about a
|
|
repo's OWN next step, which by definition cannot be moved, while owing a reply
|
|
is the other axis entirely - answering is often what unblocks it. Measured on
|
|
the real tree at 0.16.0, two of 26 planned repos were `done` with an unhandled
|
|
inbox. A cap on debt was proposed with the chain credit at 0.19.0 and DECLINED
|
|
by the operator. Sitting one group below chain-root credit is NOT that cap:
|
|
the debtor keeps its tab, its most-owed-first position among the other
|
|
debtors, and its `why=inbox:N`. A change that DROPPED a debtor from the plan
|
|
would be the declined cap wearing the group order as a disguise, and selftest
|
|
section 12 pins both halves - the root outranking four owed messages, and the
|
|
debtor keeping everything it had.
|
|
|
|
**"Debt" means OWED, never raw pending, since 0.22.0 - and this NARROWS what
|
|
counts as debt, it does not reopen the paragraph above.** The paragraph above
|
|
settles a different question: once a repo has debt, is it ever excluded or
|
|
capped (no). This one settles what counts as debt in the first place. Through
|
|
0.21.0, group 2's `keep`/`mag`/`why=inbox:N` and `--brief`'s whole "repo som
|
|
skylder et svar" listing were computed from the raw pending-file count - every
|
|
unhandled message in the inbox, including ones the sender declared
|
|
`reply-expected: no`. That is a notice, not a request, and 0.11.0 gave
|
|
`coord-count.sh` a second column (`owed`) for exactly this distinction - but
|
|
`board.sh` never read it. Reported by morning-driver (2026-08-11) and
|
|
independently reproduced against the live mailbox 2026-08-13: 27 of 72
|
|
pending messages (37.5%) were notices. The fix joins `--plan` and `--brief`
|
|
against `coord-count.sh`'s `owed` column by repo name (same technique as the
|
|
chain-root `$UNBLOCKS` join below), so a `done`/`deferred`/`blocked` repo
|
|
whose only mail is FYI no longer gets a tab, and `--brief` no longer counts a
|
|
notice as an obligation. This reverses a decision from session 41
|
|
(2026-08-10) that declined to build this filter, on the premise that "the
|
|
arrival of the request IS the admission signal" - a premise that assumed
|
|
group 2 already meant requests. It didn't; the code computed pending, the
|
|
comments already said "owed" throughout, and the plan's own printed header
|
|
("Utelatt naar repoet verken skylder svar...") already claimed the exclusion
|
|
was debt-based. The fix makes the code match what its own comments and
|
|
header already promised. Pinned by board-selftest.sh section 8/12 fixtures
|
|
`repo-done-fyi` (pending 2, owed 0 - excluded) and `repo-blocked-mixed`
|
|
(pending 3, owed 2 - planned on 2, not 3). The TABLE's `INN` column and the
|
|
raw scan (`RECORDS` field 6) are UNCHANGED - they answer "what is the state
|
|
of every repo," not "who is waiting on you," and stay on raw pending by
|
|
design.
|
|
|
|
**The same 0.22.0 patch that switched `n_owe` to OWED also had to fix what
|
|
`n_owe == 0` claims.** `--brief`'s empty-debt branch said "Ingen repo har
|
|
uhaandtert innboks. Ingen skylder noen et svar i dag." (no repo has
|
|
unhandled inbox; nobody owes a reply) - two claims in one branch, and only
|
|
the second is what `n_owe == 0` actually proves once `n_owe` means OWED. A
|
|
repo can hold FYI-only mail with zero debt, which makes the first sentence
|
|
false while it fires - caught in review before release, not by any fixture
|
|
(the shared test tree never reaches `n_owe == 0`, since it always carries a
|
|
debtor). Fixed to state only the debt claim, and to name any FYI-only
|
|
mailboxes found rather than let their existence become invisible again -
|
|
the same "labelled, not silently dropped" principle `--plan` already
|
|
applies to unknown-status repos. Pinned by board-selftest.sh section 14
|
|
with its own isolated root (debt-free, one FYI-only repo).
|
|
|
|
**The `$UNBLOCKS`/`$RECORDS` join used `NR==FNR` through 0.21.0, and that
|
|
idiom silently drops the entire plan whenever the FIRST file is empty - fixed
|
|
to `FILENAME==` comparison in 0.22.0, found while adding the `$OWED` join
|
|
above.** Verified against the shipped 0.21.0 script: one in-progress repo
|
|
with an unhandled inbox message, zero blocked repos anywhere in the tree
|
|
(so `$UNBLOCKS` is empty, which is a common, ordinary tree state, not an
|
|
edge case) - `--plan` printed "0 tabber". `NR==FNR` is only true for the
|
|
FIRST file's own lines; when that file is empty, `FNR` and `NR` stay equal
|
|
for the ENTIRE next file too (not just its first line - verified with a
|
|
minimal awk reproduction), so every record in it is misrouted into the
|
|
`ub[]` branch and dropped via `next`. This was invisible to
|
|
board-selftest.sh because the fixture tree has carried at least one
|
|
`blocked` repo since the chain-root feature shipped, and it was invisible
|
|
on the real tree because `~/repos` currently always has one too - neither
|
|
is a guarantee. `FILENAME==UBF`/`FILENAME==OWF` compares the exact path,
|
|
never line counts, so an empty lookup file degrades to "nothing matched,"
|
|
never to "everything after it is misrouted."
|
|
|
|
**Chain-root credit lands on the ROOT and nowhere else.** For every `blocked`
|
|
repo the `blocked-on` edge is followed transitively to the first repo that is
|
|
not itself blocked. Crediting a blocked repo would open a tab that cannot move;
|
|
crediting only the direct blocker leaves a two-hop chain's root uncredited,
|
|
which is the shape the real tree actually had. A cycle, a `blocked-on` naming
|
|
an unscanned repo, and a blocked repo with no target must all credit NOBODY:
|
|
inventing a root there produces a plan that looks correct and sends the
|
|
operator to the wrong repo.
|
|
|
|
Repos with no board line rank last and are LABELLED rather than dropped,
|
|
because the table already prints a MERK line about them and a plan that
|
|
omitted them silently would repeat that defect.
|
|
|
|
It renders `key=value` blocks, not prose, because it has two consumers: the
|
|
operator, and a driver repo consuming the plan. Prose would make the rendered
|
|
format an API no test in THIS repo could hold stable for a consumer in
|
|
another. `command_missing=` carries both no-command causes (no route line, and
|
|
a route line route.sh rejects) because a bare `command=` is the shape of a
|
|
runnable command carrying nothing - a driver reading `^command=` would type an
|
|
empty line into a live pane. `route_cmd_for()` is the single reader of the
|
|
route-line grammar, shared with `--brief`, and distinguishes the two causes by
|
|
exit code rather than by an empty string.
|
|
|
|
**`paste=` and `dir=`/`command=` are the same fact for the two consumers, and
|
|
neither is redundant.** A driver moves the pane itself and then types the
|
|
command, so it needs them apart; a human needs ONE thing to select. Handing
|
|
the operator two fields to join by hand is not a saved output line, it is the
|
|
step where a session starts in the wrong repo - and it was measured the moment
|
|
the feature met its first user, who could not act on the block at all. `paste=`
|
|
is emitted only alongside `command=`: `paste=cd X && ` with nothing after it
|
|
would run the cd and then a bare newline, which fails SILENTLY by leaving the
|
|
operator in the right directory with no session started.
|
|
|
|
**`board.sh --dispatch` is the FOURTH rendering, and the generator-ownership
|
|
question it settles was open for two sessions.** It carries a task INTO
|
|
another repo — "start a session in repo X, on order Y, at cost Z" — and emits
|
|
either a plan block or a single paste line. It lives in `board.sh` rather
|
|
than in a script of its own for one reason, and it is the reason the operator
|
|
and `.claude` both named first: the block format has exactly ONE generator,
|
|
and this file already is it. A second emitter of
|
|
`tab=`/`repo=`/`dir=`/`command=`/`paste=` would be two copies of one file
|
|
format, drifting apart, with a second place to get `paste=` wrong. Read-only
|
|
survives untouched: every check is a read, and the two writes a dispatch needs
|
|
(prompt file, plan file) stay with the caller — the same split
|
|
`brief-nightly.sh` already carries for the briefing.
|
|
|
|
**The cost comes from `route.sh`'s row table, and `--dispatch` deliberately
|
|
refuses a `--model`/`--effort` pair.** `--advisor opus` is a property of the
|
|
ROW; two rows share a model/effort pair while differing on it, and the CLI
|
|
accepts a wrong advisor silently. A dispatch taking the model directly would
|
|
have no honest source for that flag, and both available guesses produce the
|
|
same failure — a session that looks peer-reviewed without being. A Fable
|
|
dispatch is therefore not a `--dispatch` outcome at all, exactly as it is not
|
|
a `route.sh` outcome; it is a hand-written override.
|
|
|
|
**`--target-pane yes|no` is REQUIRED, with no default, and that is the same
|
|
rule `--last-effort` carries.** It is a measurement of the world — does the
|
|
target repo already have a Ghostty pane — and this repo must never learn to
|
|
look for a terminal itself; the caller measures with `morning --probe-panes`
|
|
and passes the fact in. Defaulting would be worst at `no`: that is the
|
|
plan-file form, and `morning`'s `plan_drop_open` (morning:1788) silently drops
|
|
a plan block for a repo that already has a pane, reporting "0 of 1" — which
|
|
reads as a broken plan file. Measured four times on 2026-08-16 by two
|
|
different repos. The `yes` form therefore emits **no `tab=` key at all**:
|
|
`plan_parse` discards a block without one, so the wrong use is impossible
|
|
rather than merely discouraged.
|
|
|
|
**The dry-run is a parse check, not the pane gate, and the difference was
|
|
measured (2026-08-16).** `morning --plan-file <f> --dry-run` proves the block
|
|
parses and yields a command. Run from a Claude session there is no tty, so
|
|
`morning` prints "window: unknown ... assuming an empty window" and
|
|
`plan_drop_open` never fires — a gate built on it would pass the self-dispatch
|
|
case every single time, which is the one case it would exist to catch.
|
|
`--probe-panes`, by contrast, DOES work without a tty: it cannot identify the
|
|
anchor pane, but the `DIR` column is there.
|
|
|
|
**The prompt goes in argv, and only the PATH has to be shell-clean.** Verified
|
|
directly: `"$(cat f)"` hands the file's bytes to the session as one argv
|
|
element with no re-evaluation, so `$(...)`, backticks, quotes and UTF-8 in the
|
|
BODY are inert — which is precisely why the prompt is passed this way instead
|
|
of inlined. The path sits inside those quotes and IS evaluated, so it must be
|
|
absolute (a relative one resolves against the pane's directory, not the
|
|
emitter's) and drawn from a safe character class. An empty prompt file is
|
|
refused with `test -s`: it would start a session and tell it nothing, which
|
|
from the far end is indistinguishable from one waiting for a Go.
|
|
|
|
**Driving a terminal from the plan does NOT belong here, and the measurement
|
|
in `docs/ghostty-orchestration-measurement.md` is the argument, not taste.**
|
|
It is a version-pinned undocumented composition over a preview API whose
|
|
documented path is already broken upstream and whose regression was closed as
|
|
not planned, with a blast radius reaching into other repos' live sessions.
|
|
None of that is mailbox transport, and none of it may be able to break
|
|
`coord-inbox` or `board`. The dependency runs one way: the driver consumes the
|
|
plan, the plan never knows a terminal exists.
|
|
|
|
It also cross-checks itself against `coord-count.sh`, and that is not
|
|
belt-and-braces. The repo scan and the mailbox are two different populations:
|
|
a mailbox can carry a name no scan will ever produce — a declared non-git
|
|
surface (`CLAUDE_COORD_REPO`, e.g. `~/repos` itself) or a checkout outside the
|
|
roots. Measured on the real mailbox at 0.15.0: the briefing found 11 repos /
|
|
21 messages where `coord-count` saw 12 mailboxes / 22 pending, the missing one
|
|
being the declared surface `repos`. A briefing that only walks the scan
|
|
answers "who is waiting on you" with a number it quietly knows is short.
|
|
|
|
**Zero model calls, and that is the load-bearing property, not an
|
|
implementation detail.** The operator authenticates by subscription, so a
|
|
headless `claude -p` job draws from the same quota pool as interactive work.
|
|
Measured against 2.1.220: `--max-budget-usd` DOES bite under subscription auth
|
|
(`terminal_reason: budget_exhausted`, exit 1), but it aborts AFTER turn one,
|
|
never before it — floor ~0.25 USD-equivalent per turn on `claude-opus-5[1m]`.
|
|
It is a runaway brake, not a pre-flight gate. Making the briefing deterministic
|
|
removes the question entirely.
|
|
|
|
`board.sh` stays read-only, which is why the file write lives in the wrapper
|
|
instead of behind a `--brief --out FILE` flag. The wrapper renders to a temp
|
|
file in the target directory and renames it into place, and treats an EMPTY
|
|
render as a FAILED one: board prints nothing at all when its scan roots do not
|
|
exist, which is what a mistyped path or a moved home directory looks like, and
|
|
a plain `> file` redirect would destroy yesterday's briefing on a bad launchd
|
|
environment. A tree where nobody owes anything is a different case — that is a
|
|
valid, non-empty briefing saying so, and is written normally.
|
|
|
|
**F3+F4 (2026-08-16): the status classifier and `route_cmd_for()`'s four-trait
|
|
extraction both matched a PREFIX of the closed vocabulary, not the exact
|
|
token.** Both used a `[a-z-]*` sed capture, which stops at the first byte
|
|
outside that class instead of running to the field's real boundary. Two
|
|
distinct failure shapes came out of the same defect: `status=done2` (a valid
|
|
token plus one byte) captured as `done` and was silently classified as a real
|
|
`done` - excluded from `--plan` the same way a genuine done repo is, and shown
|
|
green in the table, never flagged; `status=Planned` (a case variant) captured
|
|
as empty and was silently classified as `?` - "no board line at all" - which
|
|
fed the MERK footer a false count for a repo that has one. `route_cmd_for()`
|
|
carried the identical class on all four route traits, and there the failure is
|
|
worse: `path=known2` truncated to `known`, which `route.sh`'s own exact-match
|
|
validation then ACCEPTS, producing a safely-worded but WRONG startup command
|
|
(row 1) instead of the refusal a typo like `knwon` (a whole different word,
|
|
already handled correctly) already got. Fixed by capturing to the next `;` or
|
|
the closing `-->` instead - the same `[^;>]*` + trim shape `next-cost` already
|
|
used for its own reason (spec-conformant values contain spaces and capitals) -
|
|
so the classifier and `route.sh`'s case statement see the value un-truncated
|
|
and the exact-match either accepts it or correctly calls it MALFORMED /
|
|
`command_missing=`. `blocked-on`'s `[A-Za-z0-9._-]*` capture is a different,
|
|
wider class feeding a different mechanism (matched against scanned repo
|
|
names for chain-root credit, not a closed vocabulary) and was checked, not
|
|
touched - grepping `scripts/board.sh` for `[a-z-]*` after the fix returns
|
|
nothing. Selftest fixtures `repo-status-prefix`, `repo-status-case`,
|
|
`repo-trait-prefix` and `repo-revers-prefix` pin all four cases, plus the
|
|
known-positive controls that a genuinely absent board line still reads `?`
|
|
and a genuinely valid route line still derives a real command.
|
|
|
|
- **Route (`scripts/route.sh`):** pure calculator for the next session's model
|
|
and effort. Takes four scored traits of the next task plus a required
|
|
rationale, and prints one block of `key=value` lines: the rubric row, the rule
|
|
that fired, the `next-cost` value, a pasteable startup command, the one-row
|
|
cheaper fallback, and the STATE.md comment lines. Pinned by
|
|
`route-selftest.sh` (69 checks).
|
|
|
|
**It is here because it is the WRITER for the field `board.sh` already reads.**
|
|
`next-cost` had a reader and no writer, so it was hand-typed every session and
|
|
drifted into several competing spellings — cleaning the data
|
|
could not fix that, because the cause was the missing write path. The row
|
|
table is a closed set of six values, so a seventh cannot enter circulation,
|
|
and section 6 of the selftest runs the round trip (route emits → board parses)
|
|
*inside* one repo rather than across two. `board.sh` itself is untouched: a
|
|
calculator that prints to stdout writes nothing, and the session writes
|
|
STATE.md.
|
|
|
|
**The row table is the operator's global rubric, moved here as the single
|
|
copy.** It is not a second spec — `board.sh --help` documents the board line's
|
|
*grammar* and points here for the *values*. Scoring the traits is judgement
|
|
and belongs to the skill; turning scores into a row is a lookup and takes zero
|
|
model calls.
|
|
|
|
**`--advisor opus` is emitted per ROW, on a need, never unconditionally.**
|
|
Rows 1-2 always carry it (Sonnet main model, so opus is a capability lift —
|
|
and since every fallback is one row cheaper and the cheap rows are Sonnet,
|
|
this is what makes the quota fallback safe to take); rows 3-4 only at
|
|
`reversibility=costly|one-way` (Opus main model, so it buys peer review where
|
|
a mistake is not cheap to undo); rows 5-6 never, because the CLI rejects every
|
|
advisor for a Fable main model. The alternative — the global `advisorModel`
|
|
setting written by `/advisor` — is what this replaces: it applies to every
|
|
session in every repo, which is how it burned quota before. The two triggers
|
|
are almost disjoint by construction, since `costly` forces row 3 and `one-way`
|
|
forces row 4, so a Sonnet row always has `reversibility=cheap`.
|
|
`verification=none` is deliberately NOT a third trigger: beyond the stakes
|
|
rule it would only add mistakes that are cheap to reverse, docs sessions
|
|
(`known/none/cheap/local`) among them. Section 14 pins the rule and gates the
|
|
three CLI facts it rests on against the installed `claude` without spending a
|
|
token — advisor validation runs before the empty-prompt check, so `-p ""`
|
|
reaches the validator and stops there.
|
|
|
|
**Rows 5-6 are never a `route.sh` outcome.** Until 2026-08-06 they fired
|
|
only from an explicit `--opus-xhigh-failed` flag, mirroring a global
|
|
CLAUDE.md policy that Fable could only be *suggested* after a failed Opus
|
|
5/xhigh session. That policy was removed by operator decision — "for ofte
|
|
ER Fable riktig" — and the flag went with it rather than being repurposed:
|
|
`route.sh`'s output range is now closed at row 4, and a Fable choice is
|
|
always a hand-written deviation from the rubric, recorded in STATE as an
|
|
override per the model-selection rule in the global CLAUDE.md, never
|
|
produced by the calculator. `board.sh` still parses "Fable 5/high" and
|
|
"Fable 5/xhigh" written by hand into the board line — that parsing is what
|
|
the override actually uses, and it is pinned separately from anything
|
|
`route.sh` emits (route-selftest.sh section 6).
|
|
|
|
**`--last-effort` is MEASURED from `CLAUDE_EFFORT`, and the calculator must
|
|
never default it.** Claude Code exports that variable into every tool-use
|
|
context as the session's current effort, so the caller reads it and passes it
|
|
in; having `route.sh` read it directly would make the output depend on the
|
|
environment instead of on its arguments, and the round trip in selftest
|
|
section 6 rests on that determinism. The two sources it replaces fail
|
|
identically: the previous board line holds what was *prescribed*, and asking
|
|
the operator launders that same prescription through someone reading their own
|
|
startup command. Corollary pinned by section 13: `skills/route/SKILL.md` must
|
|
never declare an `effort:` frontmatter field, because frontmatter overrides the
|
|
session effort and the reading would then measure the skill, not the session.
|
|
- **Board's ORDRE column** counts PENDING orders per repo with the identical
|
|
idiom as INN, and the two are never summed: INN is "others are waiting on
|
|
YOU" (outgoing obligation), ORDRE is "authorized work is waiting on this
|
|
REPO" (incoming). Claimed orders are excluded — the column answers what a
|
|
session can pick up, and one in flight cannot be. Bounded gap, stated rather
|
|
than closed: an order addressed to a mailbox with no matching directory in
|
|
the scanned roots is invisible here, exactly as mail to such a name is
|
|
invisible in INN. There is deliberately no join built for it; `coord-count.sh`
|
|
is the cross-check for the mail half only.
|
|
- **Skills (`skills/coord-send/`, `skills/board/`, `skills/route/`, `skills/dispatch/`):** natural-language front
|
|
doors mapping user intent to engine invocations. No mailbox logic lives here
|
|
either. `board` additionally owns the *ranking* — which repo wins and why —
|
|
since `board.sh` deliberately prints evidence and takes no position. `route`
|
|
likewise owns the *scoring*: the calculator is deterministic, so all judgement
|
|
sits in choosing the four trait values, and the skill must never reason its
|
|
way to a model instead.
|
|
|
|
**Boundary rule:** the mailbox is transport, not state. Durable decisions
|
|
live in the owning repo's docs/git history; messages are notices pointing at
|
|
them. Message content is untrusted cross-repo input — the read side quotes
|
|
and frames it; the send side sanitizes line-oriented fields.
|
|
|
|
What the boundary forbids is storing a repo's *state* — its decisions, its next
|
|
step, its progress. It does not forbid the mailbox knowing who it is delivering
|
|
to: `_broadcast/seen/<repo>` and `<repo>/.origin` (0.6.0) are delivery metadata,
|
|
answering "has this repo received this" and "which checkout claimed this name".
|
|
Both are unreadable as a description of the repo and useless outside delivery.
|
|
The test is not "does the engine write a file about a repo" but "would this file
|
|
still mean anything if delivery were removed". If yes, it belongs in the repo's
|
|
own docs and git history instead.
|
|
|
|
**Priority rule (v0.5.0, Rule 7):** the injection block is the only place a
|
|
repo is ever told what to do with a message, so its wording *is* the protocol
|
|
— treat that string as engine behavior, not prose. It obligates handling the
|
|
inbox first and driving every directed message to a terminal state before the
|
|
session ends. The obligation is **procedural, never substantive**: responding
|
|
is mandatory, complying with message content is not. Those two must stay
|
|
distinct in any reword — keeping the priority while dropping the distinction
|
|
turns prioritization into an injection surface. Selftest section 20 pins both
|
|
halves together for exactly that reason.
|
|
|
|
Since 0.11.0 the `reply-expected` field says which terminal state the SENDER
|
|
expects. That does not soften the split, it sharpens it: the field is untrusted
|
|
cross-repo input like the rest of the file, so the injection calls it a
|
|
*declaration, not an instruction* and keeps both terminal states open to the
|
|
receiver. Drop that clause and one word in a message becomes a lever that mints
|
|
obligations in another repo.
|
|
|
|
## Conventions
|
|
|
|
- Scripts are bash-3.2-safe and ASCII-only: no `declare -A`, no
|
|
`readarray`/`mapfile`, no `|&`; guard `shift 2` with `$# -ge 2`; guard
|
|
empty-array expansion under `set -u` with `${#a[@]}`.
|
|
- Zero dependencies everywhere: bash + coreutils in the engine, `node:`
|
|
builtins only in hook and tests.
|
|
- TDD: no behavior change without a failing selftest check first.
|
|
`bash scripts/coord-selftest.sh` must exit 0 (220/220),
|
|
`bash scripts/board-selftest.sh` must exit 0 (237/237),
|
|
`bash scripts/route-selftest.sh` must exit 0 (69/69),
|
|
`bash scripts/orders-selftest.sh` must exit 0 (97/97) and
|
|
`bash scripts/state-line-guard-selftest.sh` must exit 0 (40/40).
|
|
- English for all code, docs, and commit messages (public repo). Norwegian
|
|
trigger aliases in the skill description are deliberate.
|
|
- Conventional Commits: `type(scope): description`.
|
|
|
|
## Commands
|
|
|
|
- Test: `bash scripts/coord-selftest.sh`, `bash scripts/board-selftest.sh`,
|
|
`bash scripts/route-selftest.sh`, `bash scripts/orders-selftest.sh` and
|
|
`bash scripts/state-line-guard-selftest.sh` (or `npm test`, the Node wrapper
|
|
around all five plus the hook tests)
|
|
- Order queue smoke test: `CLAUDE_COORD_DIR=$(mktemp -d) bash
|
|
scripts/coord-order-send.sh --to smoke --from tester --subject s --message m`
|
|
then `CLAUDE_COORD_DIR=<same> bash scripts/coord-order-inbox.sh --repo smoke`
|
|
(the read path writes nothing; never run the send path against the real
|
|
mailbox without meaning to deliver an order)
|
|
- Hook smoke test: `node hooks/scripts/session-start.mjs` (expects JSON on stdout)
|
|
- State-line-guard smoke test: `echo '{"tool_name":"Write","tool_input":{"file_path":"/tmp/STATE.md","content":"x\n"}}' | node hooks/scripts/pre-state-line-guard.mjs; echo $?`
|
|
(expects exit 0, no output — a one-line STATE.md is under the limit)
|
|
- Board smoke test: `bash scripts/board.sh` (read-only, ~3s over the real tree)
|
|
- Briefing smoke test: `bash scripts/board.sh --brief` (read-only, writes
|
|
nothing). `brief-nightly.sh` DOES write — it overwrites `$CLAUDE_BRIEF_FILE`
|
|
(default `~/.claude/briefing.md`), so point that at a scratch path when
|
|
testing. Installed as a launchd agent from `launchd/`, which points at the
|
|
SOURCE repo, never the version-pinned plugin cache.
|
|
- Dispatch smoke test: `bash scripts/board.sh --dispatch --repo repo-mailbox
|
|
--prompt-file /tmp/x.prompt --target-pane yes --path known --verification
|
|
strong --reversibility cheap --scope local --rationale smoke` (read-only;
|
|
needs a non-empty `/tmp/x.prompt`. Use `--target-pane yes` in a smoke test:
|
|
it produces no plan file, so nothing can be handed to `morning` by accident)
|
|
- Route smoke test: `bash scripts/route.sh --path known --verification strong
|
|
--reversibility cheap --scope local --rationale x` (writes nothing, instant)
|
|
- Sweep smoke test: `bash scripts/coord-sweep.sh` (dry-run is the default, so
|
|
this writes nothing; never add `--write` to a smoke test against the real
|
|
mailbox)
|
|
|
|
## Release
|
|
|
|
Version must agree across: `.claude-plugin/plugin.json`, `package.json`,
|
|
README version badge, `skills/coord-send/SKILL.md`, `skills/board/SKILL.md`,
|
|
`skills/route/SKILL.md` and `skills/dispatch/SKILL.md` frontmatter, git tag `vX.Y.Z`, and the catalog `ref` in
|
|
`ktg-plugin-marketplace/catalog/.claude-plugin/marketplace.json`. Release via
|
|
the catalog's `scripts/release-plugin.mjs repo-mailbox` (tag + ref bump
|
|
together);
|
|
verify with `scripts/check-versions.mjs`. Never hand-edit a ref.
|
|
|
|
Two things that script does that its dry-run label does not suggest:
|
|
`--create-tag` creates AND pushes the tag even without `--write`, and its
|
|
closing verification gate runs `check-versions.mjs` over ALL plugins — one
|
|
unrelated plugin in ERROR aborts it with the catalog edit written but
|
|
uncommitted. When that happens, commit the catalog's `marketplace.json` +
|
|
`README.md` by hand and leave every other dirty file in that repo alone.
|
|
|
|
**`--write --commit` does NOT close this window (confirmed by catalog,
|
|
2026-08-10).** The gate (`check-versions.mjs`) runs via `execFileSync` before
|
|
the `--commit` conditional, so it throws on any plugin's ERROR — including one
|
|
we did not touch — after the catalog files are written and before commit,
|
|
regardless of whether `--commit` was passed. Catalog is evaluating a
|
|
pre-flight gate (run the check before writing, abort there) but it is **not
|
|
implemented yet** — do not assume it exists. Until it ships: before running
|
|
`--write`, run `node scripts/check-versions.mjs` in the catalog manually and
|
|
confirm 0 ERROR first, even when the only ERROR belongs to an unrelated
|
|
plugin. If it still fires mid-release, fall back to the manual-commit
|
|
recovery above.
|
|
|
|
## Hardening roadmap
|
|
|
|
Empty — the post-v0.1.0 queue (atomic delivery, `.`/`..` rejection,
|
|
selftest gaps, uniform `-h`) shipped in v0.2.0; broadcast self-delivery
|
|
shipped in v0.2.1; broadcast retraction (`coord-send --retract`) shipped in
|
|
v0.4.0, closing the last monotonically-growing surface. `coord-inbox.sh`
|
|
still ignores unknown arguments by design (hook context must never fail)
|
|
but now warns about each one on stderr, which the hook discards.
|
|
|
|
Two retraction limits are deliberate, not gaps: it is un-send and never
|
|
recall (a repo that already received a broadcast keeps it — the seen set is
|
|
delivery history and is left untouched), and the sender check is an accident
|
|
guard, not a security boundary, because `--from` redefines identity here as
|
|
it does everywhere else in the engine.
|