feat(tokens): MCP tool-schema deferral check + CLI-over-MCP lever (v5.10 B4)
By default Claude Code defers MCP tool schemas (names-only, ~120 tok; full
schemas on demand via tool search). CA-TOK-006 detects config-file signals that
force the FULL schemas into the always-loaded prefix every turn:
- settings.json env.ENABLE_TOOL_SEARCH="false" (high)
- "ToolSearch" in permissions.deny (high)
- configured model is a Haiku model (medium)
- per-server .mcp.json alwaysLoad:true (CC v2.1.121+) (high)
auto[:N] is threshold mode (info, not a trigger).
New engine lib/mcp-deferral.mjs: pure assessMcpDeferral({settings,mcpServers})
(unit-tested, no IO) + thin IO wrapper assessMcpDeferralForRepo shared by TOK and
GAP. Severity scales with aggregate forced-upfront tokens (medium-confidence
reasons cap at medium). feature-gap cliOverMcpLeverFinding fires only as a
companion to CA-TOK-006 (prefer gh/aws/gcloud over MCP for common ops).
Honest scoping (Verifiseringsplikt): triggers on config files ONLY — never
process.env shell vars. Vertex / custom ANTHROPIC_BASE_URL / runtime /model
switch are launch state (would flap snapshots machine-dependently), so they are
DISCLOSED in every finding, not triggered. Tool-level anthropic/alwaysLoad and
claude.ai connectors likewise disclosed. Mechanism verified 2026-06-23 against
code.claude.com/docs (context-window.md, mcp.md#configure-tool-search +
#exempt-a-server-from-deferral, costs.md); the prefix-cache invalidation claim
was NOT-CONFIRMED in docs and is not asserted.
alwaysLoad added to CA-MCP VALID_SERVER_FIELDS (no longer flagged as unknown).
active-config-reader surfaces per-server alwaysLoad. Byte-stable: CA-TOK-006
fires only on new conditions; frozen v5.0.0 + SC-5 snapshots untouched.
Tests 1215 -> 1239 (engine 16, integration 5, lever 2, mcp-field guard 1).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
42e48514e4
commit
8f7e196046
13 changed files with 692 additions and 2 deletions
|
|
@ -13,6 +13,7 @@ import { finding, scannerResult } from './lib/output.mjs';
|
|||
import { SEVERITY } from './lib/severity.mjs';
|
||||
import { findImports, parseJson, parseFrontmatter } from './lib/yaml-parser.mjs';
|
||||
import { measureActiveSkillListing, isBundledSkillsDisabled, BUDGET_CALIBRATION_NOTE } from './lib/skill-listing-budget.mjs';
|
||||
import { assessMcpDeferralForRepo } from './lib/mcp-deferral.mjs';
|
||||
|
||||
const SCANNER = 'GAP';
|
||||
|
||||
|
|
@ -134,6 +135,46 @@ export function bundledSkillsLeverFinding({ leverPulled, aggregate }) {
|
|||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* CLI-over-MCP lever — remediation companion to CA-TOK-006 (v5.10 B4).
|
||||
*
|
||||
* Fires ONLY when MCP tool schemas are forced into the always-loaded prefix
|
||||
* (tool search disabled, or a per-server alwaysLoad), i.e. when MCP is actually
|
||||
* costing always-loaded tokens. When schemas are deferred (the default), MCP is
|
||||
* effectively free until used, so there is nothing to recommend and we stay
|
||||
* silent — opportunity, not noise (mirrors the bundledSkills lever's "fire only
|
||||
* under measured pressure" contract). CLI tools (gh / aws / gcloud) add zero
|
||||
* context tokens until invoked, so they are the lever the deferral mechanism
|
||||
* cannot reach for the forced-upfront servers.
|
||||
*
|
||||
* Pure and exported for unit testing.
|
||||
*
|
||||
* @param {{ assessment: (import('./lib/mcp-deferral.mjs').assessMcpDeferral)|null }} args
|
||||
* @returns {object|null} a GAP finding, or null when nothing is forced upfront
|
||||
*/
|
||||
export function cliOverMcpLeverFinding({ assessment } = {}) {
|
||||
if (!assessment || !assessment.forcedUpfront) return null;
|
||||
const names = (assessment.affectedServers || []).map((m) => m.name).join(', ');
|
||||
return finding({
|
||||
scanner: SCANNER,
|
||||
severity: SEVERITY.low,
|
||||
title: 'Prefer CLI over MCP for common operations',
|
||||
description:
|
||||
`Your active project MCP tool schemas (~${assessment.aggregateTokens} tokens) are forced into the ` +
|
||||
'always-loaded prefix every turn rather than deferred (see CA-TOK-006). CLI tools (gh, aws, gcloud, …) ' +
|
||||
'add ZERO context tokens until you actually call them, so moving common operations off MCP and onto a ' +
|
||||
'CLI reclaims always-loaded budget the deferral mechanism cannot.',
|
||||
evidence:
|
||||
`forced_schema_tokens~${assessment.aggregateTokens}; servers=${names}; ` +
|
||||
`reason=${assessment.reason || 'alwaysLoad'} (companion to CA-TOK-006)`,
|
||||
recommendation:
|
||||
'For operations a CLI already covers (GitHub → gh, AWS → aws, GCP → gcloud), prefer the CLI over an ' +
|
||||
'MCP server — CLI output enters context only when invoked. Keep MCP for capabilities with no CLI ' +
|
||||
'equivalent, disable unused servers via /mcp, and re-enable tool-search deferral so the rest stay names-only.',
|
||||
category: 'token-efficiency',
|
||||
});
|
||||
}
|
||||
|
||||
/** @type {GapCheck[]} */
|
||||
const GAP_CHECKS = [
|
||||
// --- Tier 1: Foundation ---
|
||||
|
|
@ -441,6 +482,13 @@ export async function scan(targetPath, sharedDiscovery) {
|
|||
const leverFinding = bundledSkillsLeverFinding({ leverPulled, aggregate });
|
||||
if (leverFinding) findings.push(leverFinding);
|
||||
|
||||
// CLI-over-MCP lever — companion to CA-TOK-006: fires only when project-local
|
||||
// MCP tool schemas are forced into the always-loaded prefix (tool search
|
||||
// disabled or a per-server alwaysLoad). Reuses the same static assessment.
|
||||
const mcpAssessment = await assessMcpDeferralForRepo(ctx.targetPath);
|
||||
const cliLever = cliOverMcpLeverFinding({ assessment: mcpAssessment });
|
||||
if (cliLever) findings.push(cliLever);
|
||||
|
||||
const filesScanned = discovery.files.length;
|
||||
return scannerResult(SCANNER, 'ok', findings, filesScanned, Date.now() - start);
|
||||
}
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue