fix(route): measure the effort a session ran with, instead of asking for it

The last-session record exists to make the routing policy falsifiable, and
it only is if --last-effort is measured. route.sh documented the opposite as
settled fact: that the effort a session ran with is not observable from
inside that session. That was true when written and is not now. Claude Code
exports CLAUDE_EFFORT into every tool-use context as the session's current
effort level, so a Bash call reads it directly.

The premise had a cost. With effort unobservable, the record could only be
completed by asking the operator at session end, which made it block on their
presence -- all four fields or none. That is also the weaker measurement, and
in the same way the previous board line is: the operator reads the effort off
the startup command they typed, so both sources report what was PRESCRIBED
rather than what was RUN. They come apart exactly when the record would be
most interesting, which is what a session that silently ran xhigh under a
board line saying high already showed.

Reading it makes all four fields knowable from inside the ending session, so
the record no longer waits on anyone. The skill does the reading; route.sh
deliberately does NOT default from the variable, because a calculator that
consults its environment is no longer deterministic from its arguments and
the route->board round trip in selftest section 6 rests on that.

Section 13 also pins the trap this opens: skill frontmatter overrides the
session effort while that skill is active, so an effort: field in route's own
SKILL.md would make the reading report the skill instead of the session --
a measurement quietly measuring itself, with nothing in the output to show
it happened.

Also corrects the neighbouring claim that the model is readable from the
environment. There is no CLAUDE_MODEL; the session takes it from what it
knows itself to be running as.

route-selftest 50 -> 56.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gfa1nvwGXdST2MHvbs6htD
This commit is contained in:
Kjell Tore Guttormsen 2026-07-31 17:21:10 +02:00
commit ac62a38b40
5 changed files with 96 additions and 17 deletions

View file

@ -109,15 +109,39 @@ in the session that is ending — never from what STATE.md prescribed:
- `--last-corrections <n>` — how many rounds of rework it took. This is the
cheap proxy for whether the routing was right.
**Effort is not observable from inside a running session.** The model is
readable from the environment; the effort level the operator launched with is
not. So do not infer it, and in particular do not read it back from the previous
board line — that measures what was *prescribed*, not what was *run*, and the
two come apart exactly when the record would be most interesting. Ask the
operator for the effort level when closing the session. If they are not there to
ask, **omit the record entirely** — it is all four fields or none. A guessed
effort is worse than a missing one, because it reads back later as a
measurement.
**Measure the effort, never infer it.** Read what this session actually
resolved:
```bash
echo "$CLAUDE_EFFORT"
```
`CLAUDE_EFFORT` is Claude Code's own *current* effort level, exported into every
tool-use context — which is why a Bash call can read it. Pass it verbatim as
`--last-effort`. The model is not in the environment (there is no
`CLAUDE_MODEL`); take it from what this session knows itself to be running as.
Two sources are **wrong on purpose**, and both fail the same way. The previous
board line holds what was *prescribed*, not what was *run* — the two come apart
exactly when the record would be most interesting. Asking the operator launders
that same prescription through a human, who is reading it off the startup
command they typed rather than off the running process. Confirming a measured
value with them is fine; sourcing it from them is not.
So the record no longer waits on anyone: all four fields are knowable from
inside the session that is ending. Still omit it entirely — all four or none —
if any one of them is genuinely unknown. A guessed value is worse than a missing
one, because it reads back later as a measurement.
**This file must never declare an `effort:` frontmatter field.** Skill
frontmatter overrides the session effort while the skill is active, so the
reading above would report *this skill's* effort instead of the session's — a
measurement measuring itself, with nothing in the output to show it happened.
Pinned by `route-selftest.sh` section 13.
One honest limit: `CLAUDE_EFFORT` is the *current* level, so if the operator
changed it mid-session with `/effort`, "the effort this session ran with" is not
a single value. Record the level the work was actually done at and say so.
Read the previous `route-last` line out of STATE.md before overwriting it.
Pass `--opus-xhigh-failed` **only** when it says an `Opus 5`/`xhigh` session ran