feat(route,board): strike the advisor rule, add board.sh --row <repo>
Order 20260912T202210Z-7588027378-from-.claude, operator decision 2026-09-12 (helhetlig vurdering av arbeidssystemet, cut row 3 and the board.sh --row improvement row). One order, two parts, one version bump. THE ADVISOR RULE IS STRUCK. route.sh and board.sh --dispatch emit no --advisor at all. The rule fired per ROW on a need - always on the Sonnet rows (a capability lift, which is what made the quota fallback safe to take), and on the Opus rows at reversibility=costly|one-way - and it read well. It was killed by a MEASUREMENT, not by taste: of 54 dispatches the PM issued 08.-12.09, ZERO carried the flag, because sessions are started by hand from the model and effort rather than from the whole emitted line. A rule nothing honours is not a policy, and an emitted value nobody acts on is decoration in a field whose only job is to be evidence. The advisor is now what it already was in practice: an operator decision per session, said in one sentence in route.sh --help. The comments that rested on the rule were REWRITTEN, not left standing. board.sh --dispatch still refuses a --model/--effort pair, but the reason is no longer "the advisor is a property of the ROW": it is that the rubric has exactly one copy, and a dispatch taking the model directly would be a second, unscored way to reach the same decision - recording no traits, no rationale and no next-cost, so nothing afterwards could say whether the routing or the scoring was wrong. A comment defending a removed mechanism is how the next session restores it. Both skills carry the correction. Pinned as an ABSENCE over the whole trait space - 81 combinations, every line of output, with a known-positive control proving the sweep's grep can find a planted advisor - rather than on four sampled rows, because the claim is that no path emits it. board.sh --dispatch at reversibility=costly is pinned separately: that is the exact input a reintroduced rule would fire on. The literal string is absent from route.sh entirely, including the paragraph recording what was struck (it says "an opus advisor flag" in words), because a blunt grep cannot tell a description from a specification. Backward compatibility is pinned rather than assumed: a route line carrying a legacy advisor= field still parses and still yields a command - measured, 0 of 48 route lines in ~/repos carry one, but a reader that broke on an unknown field would turn last month's STATE.md into "that repo has no route line". The three CLI gates section 14 carried went with the rule; the suite no longer depends on the installed claude at all. board.sh --row <repo> IS THE SEVENTH RENDERING of the same scan, never a second scan, read-only like every other one. (The order calls it the sixth; by this file's own numbering --inbox-plan is the fourth and --dispatch the fifth. Corrected rather than carried wrong.) It exists because the columns WERE misread: on 11.09 the PM read FLY off the table by eye and got it wrong, while every other rendering a program consumes is already key=value. inn, ordre and fly are three separate fields because they are three separate facts; status is the bare token, never the table's blocked>target display, with blocked-on beside it; neste is last and uncut. An unknown repo exits 2 and writes NOTHING to stdout - an empty block would read as a repo whose every column is blank, which is a real and different state. upushet is the ONE field that is not a rendering of the scan, and it is named rather than blended in: nothing in the scan measures it, so it is read once, for the named repo only, and never enters the table, the plan or the briefing. It reads the remote-TRACKING ref, not the remote, so upushet=N honestly means "the local ref says N"; a repo with no upstream reports ?, never 0. The row fixture's three counts are three DIFFERENT integers (3/2/1), and that is the finding worth recording. Built first with 2/1/1, it was mutation-tested by making fly read the ORDRE field - the exact 11.09 misreading - and the check stayed GREEN, because the two fields held the same digit. A fixture that cannot tell two columns apart is the defect wearing a passing test, inside the section written to prevent it. Suites under /bin/bash 3.2, before -> after: coord 257 -> 257, board 393 -> 427, route 73 -> 73 (13 advisor checks and 3 CLI gates out, 15 absence/legacy checks in, and it no longer varies with claude being on PATH), orders 116 -> 116, state-line-guard 54 -> 54. Sum 893 -> 927, README badge updated to the measured sum. npm test 12/12, fail 0. Verified live against the real tree, not only fixtures: --row repo-mailbox reports fly=1 beside ordre=0 (the distinction that was misread), --row on the nested key from-ai-to-chitta/content-sadhguru resolves, and an unknown repo exits 2. No tag, no push, no catalog change - that is the operator's release-plugin.mjs run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
671e275a97
commit
5e5bc4a66e
13 changed files with 689 additions and 271 deletions
|
|
@ -29,10 +29,7 @@
|
|||
# departure from it is recorded in STATE as an OVERRIDE, never produced here.
|
||||
# Write "Fable 5/high" or "Fable 5/xhigh" into the board line by hand when
|
||||
# that is the right call; board.sh still parses both (route-selftest.sh
|
||||
# section 6 pins that half). The one fact worth carrying: a Fable session
|
||||
# runs without an advisor (row_advisor() below, and gated against the
|
||||
# installed claude by route-selftest.sh section 14) - informational, never a
|
||||
# gate on reaching the row, since there is no longer a gate to reach.
|
||||
# section 6 pins that half).
|
||||
#
|
||||
# Cheapest first, so the rubric's "always name one row cheaper as the quota
|
||||
# fallback" is row minus one, floored at row 1, correct by construction.
|
||||
|
|
@ -66,32 +63,14 @@
|
|||
# overkill costs quota every session - but a wrong architecture decision in a
|
||||
# published plugin costs more than either.
|
||||
#
|
||||
# THE ADVISOR is emitted into the command as '--advisor opus' - a second,
|
||||
# stronger model consulted at key moments during the session. It is added on a
|
||||
# NEED, never unconditionally: an always-on advisor is the global advisorModel
|
||||
# setting, which burns quota on every session in every repo and is the thing
|
||||
# this rule exists to replace. Two independent needs qualify:
|
||||
#
|
||||
# rows 1-2 ALWAYS. The main model is Sonnet, so opus is a capability LIFT
|
||||
# rather than a peer - opus judgement at sonnet cost. This is what
|
||||
# makes the FALLBACK safe to take: every fallback is one row
|
||||
# cheaper, and the cheapest rows are the Sonnet ones.
|
||||
# rows 3-4 only at reversibility=costly|one-way. The main model is already
|
||||
# Opus, so the advisor buys peer review, worth paying for when a
|
||||
# mistake is not cheap to undo.
|
||||
#
|
||||
# Rows 5-6 never reach this logic at all - the calculator cannot select them
|
||||
# (see above). Informational only: were the operator to hand-write a Fable
|
||||
# command, it would carry no advisor either way, since the CLI rejects every
|
||||
# advisor for a Fable main model.
|
||||
#
|
||||
# The two triggers barely overlap: costly forces row 3 and one-way forces row
|
||||
# 4, so a Sonnet row always has reversibility=cheap. verification=none is
|
||||
# deliberately not a third trigger - beyond the stakes rule it would only add
|
||||
# mistakes that are cheap to reverse, docs sessions among them.
|
||||
#
|
||||
# Applied per ROW, so 'fallback-command' carries its own correct answer rather
|
||||
# than the winning row's.
|
||||
# THE ADVISOR IS NOT EMITTED, and that is a decision rather than an omission:
|
||||
# the advisor is an operator decision per session, never the rubric's. Until
|
||||
# 2026-09-12 this calculator appended an opus advisor flag on a NEED - always on
|
||||
# the Sonnet rows, and on the Opus rows at costly|one-way stakes. It was struck
|
||||
# on a measurement: of 54 dispatches issued 08.-12.09 not one carried the flag,
|
||||
# because sessions are started by hand from the model and effort, not from the
|
||||
# whole line. A rule nothing honours is not policy, and an emitted value that
|
||||
# nobody acts on is decoration in a field whose only job is to be evidence.
|
||||
#
|
||||
# WHERE IT DISAGREES WITH THE RUBRIC'S EXAMPLES. The rows are task-type labels;
|
||||
# the traits are a different classification over the same six outcomes. They
|
||||
|
|
@ -288,41 +267,14 @@ row_base_cmd() {
|
|||
esac
|
||||
}
|
||||
|
||||
# THE ADVISOR is a second, stronger model consulted mid-task. It costs real
|
||||
# tokens per session, so it fires on a NEED and nowhere else - an unconditional
|
||||
# advisor is just the global advisorModel setting, which is the thing this
|
||||
# replaces. Two independent needs qualify, and they are almost disjoint:
|
||||
#
|
||||
# rows 1-2 (Sonnet) ALWAYS. opus is a capability LIFT here, not a peer:
|
||||
# opus judgement at sonnet cost. This half is what makes
|
||||
# the fallback-command safe, since every fallback is one
|
||||
# row cheaper and the cheapest rows are Sonnet.
|
||||
# rows 3-4 (Opus) only at costly|one-way stakes, where the advisor is a
|
||||
# peer review and being wrong is not cheap to undo.
|
||||
#
|
||||
# Rows 5-6 (Fable) never reach this function - $ROW can only be 1-4 (see
|
||||
# SELECTION above). Informational only: the only advisor this script ever
|
||||
# emits is opus (pinned by selftest 14's "the only advisor value ever emitted
|
||||
# is opus"), and opus is refused as under-capable for a fable main model -
|
||||
# measured against the installed claude, still true at CC 2.1.226 - so a
|
||||
# hand-written Fable command carries no advisor either way.
|
||||
#
|
||||
# costly forces row 3 and one-way forces row 4, so a Sonnet row always has
|
||||
# reversibility=cheap - the stakes rule can never reach rows 1-2, and the model
|
||||
# rule never reaches rows 3-4. verification=none is deliberately NOT a trigger:
|
||||
# beyond the stakes rule it would only add cheap-to-reverse mistakes, and it
|
||||
# would put an advisor on every docs session (known/none/cheap/local).
|
||||
#
|
||||
# Applied per ROW rather than once, because the fallback is a real command the
|
||||
# operator pastes under quota pressure and must carry its own correct answer.
|
||||
row_advisor() {
|
||||
case "$1" in
|
||||
1|2) echo " --advisor opus" ;;
|
||||
3|4) case "$2" in costly|one-way) echo " --advisor opus" ;; *) echo "" ;; esac ;;
|
||||
*) echo "" ;;
|
||||
esac
|
||||
}
|
||||
row_cmd() { printf '%s%s\n' "$(row_base_cmd "$1")" "$(row_advisor "$1" "$REVERS")"; }
|
||||
# No advisor is appended here or anywhere else - the advisor is an operator
|
||||
# decision per session, not a property this rubric computes (struck
|
||||
# 2026-09-12, see the header). row_cmd() is therefore the row's base command
|
||||
# and nothing more; it stays a function rather than collapsing into
|
||||
# row_base_cmd() because the emitted command and the row table are two
|
||||
# separate things that happened to converge, and a later flag would attach
|
||||
# here, to one place, for both the winning row and its fallback.
|
||||
row_cmd() { row_base_cmd "$1"; }
|
||||
|
||||
FB=$((ROW - 1)); [ "$FB" -ge 1 ] || FB=1
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue