feat(scanners): model/effort routing becomes a lever, not a 25th dimension (C4)
New GAP finding CA-GAP-028: authored subagents exist and not one of them names `model:` or `effort:`, so every delegated task runs on the main conversation's model (`model` defaults to `inherit`). Cites BP-MODEL-001/002, landed in C1. `whats-active` and `manifest` now carry `model`/`effort` per agent. Shipped as a conditional LEVER rather than a 25th dimension, and the choice was made by measurement: as a t3 dimension the agent-less marketplace-medium fixture would count it vacuously-present, moving the denominators 41->42 and utilization 44->45 — which flips `segment` "Developing"->"Competent" in the frozen v5.0.0 posture baseline, a field strip-retired-gap.mjs does not mask. A lever never enters those denominators. The general rule is now an invariant in CLAUDE.md. One check across both axes, not one per axis: it fires only when neither is used anywhere, so a deliberate everything-on-one-model policy stays silent. Cost is recall, chosen for precision. Found by dogfooding, fixed red-first: `model: inherit` is the documented default spelled out, so it must not count as routing — otherwise a config opts out of the opportunity without changing anything real. Two pre-existing defects surfaced and closed on the way: - The humanizer guard asserted TRANSLATIONS.GAP.static EQUALS the dimension titles, which forbade humanizing any lever — all three existing levers fell through to the generic "feature opportunity" default, wrong for a budget lever. Guard now requires coverage of every emittable title, seen red against those three before the entries were written. - Two hand-written copies of the lever list (finding-codes guard, humanizer guard) merged into one exported LEVERS registry carrying code AND title. - suppression-validation pinned CA-GAP-028 as an unoccupied number; C4 claimed it. Fixed structurally with a derived first-free id, not by picking a new literal — same class as #60's "bump this again". Suite 1596/0. Frozen v5.0.0 snapshots untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Pq3nye21RVYk4pZLeT8pGz
This commit is contained in:
parent
e861e63a7b
commit
9ae4be26d2
16 changed files with 512 additions and 31 deletions
|
|
@ -1,9 +1,10 @@
|
|||
/**
|
||||
* GAP Scanner — Feature Gap Scanner
|
||||
* Compares actual configuration against complete Claude Code feature register.
|
||||
* 25 gap dimensions across 4 tiers, plus a conditional disableBundledSkills
|
||||
* budget-lever check (remediation companion to SKL CA-SKL-002, fires only under
|
||||
* measured skill-listing pressure). Always runs with includeGlobal: true.
|
||||
* 24 gap dimensions across 4 tiers, plus four conditional levers (bundled-skills
|
||||
* budget, CLI-over-MCP, hook-output filtering, agent model/effort routing) which
|
||||
* fire only under a measured condition and are therefore NOT dimensions: they
|
||||
* stay out of the scoring denominators. Always runs with includeGlobal: true.
|
||||
* Finding IDs: CA-GAP-NNN
|
||||
*/
|
||||
|
||||
|
|
@ -115,6 +116,36 @@ const TIER_SEVERITY = {
|
|||
t4: SEVERITY.info,
|
||||
};
|
||||
|
||||
/**
|
||||
* Titles of the conditional levers — findings this scanner emits that are NOT
|
||||
* dimensions in GAP_CHECKS. They fire only under a measured condition, so they
|
||||
* carry no tier and never enter the scoring denominators (TIER_COUNTS /
|
||||
* TOTAL_DIMENSIONS) or the scoring TITLE_TO_ID map.
|
||||
*
|
||||
* Exported as the single source of both the code and the title: the
|
||||
* finding-code registry guard needs the codes, the humanizer coverage guard
|
||||
* needs the titles, and a hand-maintained copy of either list in a test is the
|
||||
* two-copies-drift class. One object so the two cannot disagree.
|
||||
*/
|
||||
export const LEVERS = {
|
||||
bundledSkills: {
|
||||
code: 'bundled-skills-lever',
|
||||
title: 'Bundled skills add to an over-budget skill listing',
|
||||
},
|
||||
cliOverMcp: {
|
||||
code: 'cli-over-mcp-lever',
|
||||
title: 'Prefer CLI over MCP for common operations',
|
||||
},
|
||||
filterHookOutput: {
|
||||
code: 'filter-hook-output-lever',
|
||||
title: 'Filter hook output before it enters context',
|
||||
},
|
||||
agentModelRouting: {
|
||||
code: 'agent-model-routing-lever',
|
||||
title: 'Subagents pin neither model nor effort',
|
||||
},
|
||||
};
|
||||
|
||||
/**
|
||||
* Lazily read and cache file content.
|
||||
* @param {CheckContext} ctx
|
||||
|
|
@ -177,8 +208,8 @@ export function bundledSkillsLeverFinding({ leverPulled, aggregate }) {
|
|||
return finding({
|
||||
scanner: SCANNER,
|
||||
severity: SEVERITY.low,
|
||||
code: 'bundled-skills-lever',
|
||||
title: 'Bundled skills add to an over-budget skill listing',
|
||||
code: LEVERS.bundledSkills.code,
|
||||
title: LEVERS.bundledSkills.title,
|
||||
description:
|
||||
`Your ${aggregate.scanned} active skills already carry ~${aggregate.aggregateTokens} tokens of ` +
|
||||
`description text, over the ${aggregate.budgetTokens}-token listing budget Claude Code allots the ` +
|
||||
|
|
@ -222,8 +253,8 @@ export function cliOverMcpLeverFinding({ assessment } = {}) {
|
|||
return finding({
|
||||
scanner: SCANNER,
|
||||
severity: SEVERITY.low,
|
||||
code: 'cli-over-mcp-lever',
|
||||
title: 'Prefer CLI over MCP for common operations',
|
||||
code: LEVERS.cliOverMcp.code,
|
||||
title: LEVERS.cliOverMcp.title,
|
||||
description:
|
||||
`Your active project MCP tool schemas (~${assessment.aggregateTokens} tokens) are forced into the ` +
|
||||
'always-loaded prefix every turn rather than deferred (see CA-TOK-006). CLI tools (gh, aws, gcloud, …) ' +
|
||||
|
|
@ -260,8 +291,8 @@ export function filterHookLeverFinding({ flaggedHooks } = {}) {
|
|||
return finding({
|
||||
scanner: SCANNER,
|
||||
severity: SEVERITY.info,
|
||||
code: 'filter-hook-output-lever',
|
||||
title: 'Filter hook output before it enters context',
|
||||
code: LEVERS.filterHookOutput.code,
|
||||
title: LEVERS.filterHookOutput.title,
|
||||
description:
|
||||
`${hooks.length} active hook${hooks.length === 1 ? '' : 's'} build hookSpecificOutput.additionalContext ` +
|
||||
"from un-grepped command output (see HKV advisory). That field enters Claude's context on every fire, " +
|
||||
|
|
@ -276,6 +307,101 @@ export function filterHookLeverFinding({ flaggedHooks } = {}) {
|
|||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Agent model/effort routing lever (C4) — cites BP-MODEL-001/002.
|
||||
*
|
||||
* A LEVER rather than a GAP_CHECKS dimension, and deliberately so. The question
|
||||
* "do your subagents route model/effort?" has no meaningful reading on a config
|
||||
* with no subagents — the `No custom subagents` dimension (t2_6) owns that case,
|
||||
* and firing here too would just double-report it. A dimension can only express
|
||||
* "not applicable" as "present", which would also inflate the utilization
|
||||
* denominator for every agent-less config.
|
||||
*
|
||||
* ONE check across BOTH axes, not two: it fires only when NEITHER `model:` nor
|
||||
* `effort:` appears on ANY authored agent. A deliberate all-on-one-model setup
|
||||
* therefore stays silent, which is the precision the opportunity framing needs.
|
||||
* The cost is recall — a config that pins `model:` everywhere but never uses
|
||||
* `effort:` gets no nudge. That trade is the v1 boundary, not an oversight.
|
||||
*
|
||||
* Pure and exported for unit testing.
|
||||
*
|
||||
* @param {{ agentCount: number, modelPinned: number, effortPinned: number }} counts
|
||||
* @returns {object|null} a GAP finding, or null when there is no opportunity
|
||||
*/
|
||||
export function agentModelRoutingLeverFinding({ agentCount, modelPinned, effortPinned }) {
|
||||
if (!agentCount) return null;
|
||||
if (modelPinned > 0 || effortPinned > 0) return null;
|
||||
|
||||
return finding({
|
||||
scanner: SCANNER,
|
||||
severity: SEVERITY.info,
|
||||
code: LEVERS.agentModelRouting.code,
|
||||
title: LEVERS.agentModelRouting.title,
|
||||
description:
|
||||
`All ${agentCount} of your subagents name neither a \`model:\` nor an \`effort:\` in their frontmatter. ` +
|
||||
'The `model` field defaults to `inherit`, so each one runs on the main conversation\'s model — omitting ' +
|
||||
'it is not a neutral default but a choice to pay the session\'s rate for every delegated task ' +
|
||||
'(BP-MODEL-001, https://code.claude.com/docs/en/sub-agents). Reasoning effort is a separate axis with ' +
|
||||
'its own frontmatter field and its own default, so a subagent can be routed on either or both ' +
|
||||
'(BP-MODEL-002, https://code.claude.com/docs/en/model-config).',
|
||||
evidence:
|
||||
`authored_agents=${agentCount}; model_pinned=${modelPinned}; effort_pinned=${effortPinned}; ` +
|
||||
'lever=agent frontmatter `model:` / `effort:` (plugin-bundled and fixture agents excluded)',
|
||||
recommendation:
|
||||
'Pin a cheaper `model:` on the subagents whose work is mechanical or read-only (search, extraction, ' +
|
||||
'summarisation) and leave the orchestrating session on the stronger model; pin a lower `effort:` on the ' +
|
||||
'same ones and reserve the high levels for work whose product is judgement. If running everything on one ' +
|
||||
'model is a deliberate policy, suppress this with `CA-GAP-028` in `.config-audit-ignore`.',
|
||||
category: 'model-fit',
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Count authored agents and how many pin each routing axis.
|
||||
* Frontmatter-only read; an unparseable or frontmatter-less file counts as an
|
||||
* agent that pins nothing, matching what Claude Code would load.
|
||||
* @param {CheckContext} ctx
|
||||
* @returns {Promise<{ agentCount: number, modelPinned: number, effortPinned: number }>}
|
||||
*/
|
||||
async function countAgentRouting(ctx) {
|
||||
let agentCount = 0;
|
||||
let modelPinned = 0;
|
||||
let effortPinned = 0;
|
||||
for (const file of ctx.files.filter(f => f.type === 'agent-md')) {
|
||||
agentCount++;
|
||||
const content = await getContent(ctx, file.absPath);
|
||||
if (!content) continue;
|
||||
const { frontmatter } = parseFrontmatter(content);
|
||||
if (!frontmatter) continue;
|
||||
if (isRoutingValue(frontmatter.model) && !isDefaultModel(frontmatter.model)) modelPinned++;
|
||||
if (isRoutingValue(frontmatter.effort)) effortPinned++;
|
||||
}
|
||||
return { agentCount, modelPinned, effortPinned };
|
||||
}
|
||||
|
||||
/**
|
||||
* True for a frontmatter value that actually names something. An empty or
|
||||
* whitespace-only `model:` is a no-op in Claude Code, so it must not read as a pin.
|
||||
* @param {*} v
|
||||
* @returns {boolean}
|
||||
*/
|
||||
function isRoutingValue(v) {
|
||||
return typeof v === 'string' ? v.trim().length > 0 : v != null && v !== false;
|
||||
}
|
||||
|
||||
/**
|
||||
* `inherit` IS the documented default for a subagent's `model` (BP-MODEL-001),
|
||||
* so writing it explicitly routes nothing — the agent still runs on the main
|
||||
* conversation's model. Spelling out a default must not buy silence, or a config
|
||||
* can opt out of the opportunity without changing a single thing about cost.
|
||||
* Effort has no documented sentinel of this kind, so it has no counterpart here.
|
||||
* @param {*} v
|
||||
* @returns {boolean}
|
||||
*/
|
||||
function isDefaultModel(v) {
|
||||
return typeof v === 'string' && v.trim().toLowerCase() === 'inherit';
|
||||
}
|
||||
|
||||
/** @type {GapCheck[]} */
|
||||
export const GAP_CHECKS = [
|
||||
// --- Tier 1: Foundation ---
|
||||
|
|
@ -608,6 +734,13 @@ export async function scan(targetPath, sharedDiscovery) {
|
|||
const hookLever = filterHookLeverFinding({ flaggedHooks });
|
||||
if (hookLever) findings.push(hookLever);
|
||||
|
||||
// Agent model/effort routing lever (C4) — fires only when authored agents
|
||||
// exist and not one of them uses either routing axis. Reads the SAME authored
|
||||
// set as the presence checks, so plugin-bundled and fixture agents cannot
|
||||
// make a machine look routed (M-BUG-13).
|
||||
const routingLever = agentModelRoutingLeverFinding(await countAgentRouting(ctx));
|
||||
if (routingLever) findings.push(routingLever);
|
||||
|
||||
const filesScanned = discovery.files.length;
|
||||
return scannerResult(SCANNER, 'ok', findings, filesScanned, Date.now() - start);
|
||||
}
|
||||
|
|
|
|||
|
|
@ -817,7 +817,7 @@ export async function enumerateRules(repoPath, pluginList = []) {
|
|||
*
|
||||
* @param {string} repoPath
|
||||
* @param {Array<{name:string, path:string}>} [pluginList]
|
||||
* @returns {Promise<Array<{name:string, source:string, pluginName:string|null, path:string, bytes:number, estimatedTokens:number, loadPattern:string, survivesCompaction:string, derivationConfidence:string}>>}
|
||||
* @returns {Promise<Array<{name:string, source:string, pluginName:string|null, path:string, bytes:number, estimatedTokens:number, model:string|null, effort:string|null, loadPattern:string, survivesCompaction:string, derivationConfidence:string}>>}
|
||||
*/
|
||||
export async function enumerateAgents(repoPath, pluginList = []) {
|
||||
const out = [];
|
||||
|
|
@ -842,6 +842,11 @@ export async function enumerateAgents(repoPath, pluginList = []) {
|
|||
path: f.path,
|
||||
bytes: f.size,
|
||||
estimatedTokens: estimateTokens(f.size, 'frontmatter'),
|
||||
// Routing axes (C4). Explicit null rather than an absent key: `model`
|
||||
// defaults to `inherit` and `effort` to the session level, so a consumer
|
||||
// must be able to read "not pinned" without guessing (BP-MODEL-001/002).
|
||||
model: hasText(frontmatter && frontmatter.model) ? frontmatter.model.trim() : null,
|
||||
effort: hasText(frontmatter && frontmatter.effort) ? frontmatter.effort.trim() : null,
|
||||
...lp,
|
||||
});
|
||||
}
|
||||
|
|
|
|||
|
|
@ -209,8 +209,11 @@ export const FINDING_CODES = {
|
|||
},
|
||||
|
||||
// ── GAP: feature-gap-scanner ────────────────────────────────────────────
|
||||
// Keys are GAP_CHECKS[].id (already stable). Dimensions 1–24 in table order,
|
||||
// then the three conditional levers, which the scanner emits after the loop.
|
||||
// Keys are GAP_CHECKS[].id for dimensions, and the lever code for the
|
||||
// conditional levers the scanner emits after the loop. Numbers 1–24 happen to
|
||||
// follow the current table order because that is how the dimensions were first
|
||||
// published — NOT because position determines the number. A new check takes the
|
||||
// next free number wherever it sits in the file (M-BUG-28).
|
||||
GAP: {
|
||||
t1_1: 1,
|
||||
t1_2: 2,
|
||||
|
|
@ -239,6 +242,7 @@ export const FINDING_CODES = {
|
|||
'bundled-skills-lever': 25,
|
||||
'cli-over-mcp-lever': 26,
|
||||
'filter-hook-output-lever': 27,
|
||||
'agent-model-routing-lever': 28,
|
||||
},
|
||||
};
|
||||
|
||||
|
|
|
|||
|
|
@ -524,6 +524,30 @@ export const TRANSLATIONS = {
|
|||
description: 'Language-server connections let Claude see types, error messages, and definitions the same way your editor does.',
|
||||
recommendation: 'Set up LSP integration if you work in a typed language.',
|
||||
},
|
||||
// Conditional levers. These are not "a feature you haven't set up" — they
|
||||
// fire only under a measured condition, so the generic _default would
|
||||
// misdescribe them. Every title the scanner can emit needs an entry here
|
||||
// (guarded in tests/scanners/feature-gap-scanner.test.mjs).
|
||||
'Bundled skills add to an over-budget skill listing': {
|
||||
title: 'Built-in skills are crowding an already-full skill list',
|
||||
description: 'Claude Code loads its own built-in skills into the same limited list as yours. Your list is already over budget, so entries risk being cut off and Claude may miss the right skill.',
|
||||
recommendation: 'Turn off the built-in skills to free up room — unless you use them, in which case shorten your own skill descriptions instead.',
|
||||
},
|
||||
'Prefer CLI over MCP for common operations': {
|
||||
title: 'Some connected services load their full tool list every turn',
|
||||
description: 'Most connected services only cost tokens when used, but yours are set to load everything upfront. That weight is there whether you use them or not.',
|
||||
recommendation: 'For services with a command-line equivalent (like `gh` or `aws`), the command line costs nothing until you run it.',
|
||||
},
|
||||
'Filter hook output before it enters context': {
|
||||
title: 'An automation is pasting its full output into the conversation',
|
||||
description: 'An automation that injects its output adds it to every turn that follows. Unfiltered command output can be much larger than the part that actually matters.',
|
||||
recommendation: 'Trim the output inside the script itself, so only the useful lines reach the conversation.',
|
||||
},
|
||||
'Subagents pin neither model nor effort': {
|
||||
title: 'Your helper agents all run at the same cost as your main session',
|
||||
description: 'A subagent that names no model inherits the one you are using, so routine delegated work costs the same as your hardest work. Reasoning effort is a separate dial with the same default.',
|
||||
recommendation: 'Give mechanical agents (search, extraction, summarizing) a smaller model or a lower effort level, and keep the strong settings for the work that needs judgement.',
|
||||
},
|
||||
},
|
||||
patterns: [],
|
||||
_default: {
|
||||
|
|
|
|||
|
|
@ -119,6 +119,10 @@ export function buildManifest(activeConfig) {
|
|||
name: a.name,
|
||||
source: sourceLabel(a, 'project'),
|
||||
estimated_tokens: a.estimatedTokens || 0,
|
||||
// Routing axes (C4) — named explicitly because withLoadPattern copies the
|
||||
// row plus the load-pattern triple, nothing else from the enumeration.
|
||||
model: a.model ?? null,
|
||||
effort: a.effort ?? null,
|
||||
}, a));
|
||||
}
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue