BREAKING CHANGE: the {NNN} in CA-{SCANNER}-{NNN} identifies the check that
produced the finding. It used to be the finding's position in that scanner's
output for that run, which made it unstable across CONFIGURATIONS, not just
across releases as STATE framed it. Measured on two fixtures: "No custom
subagents" was CA-GAP-007 on minimal-project and CA-GAP-004 on healthy-project.
A user who fixed an unrelated earlier gap silently renumbered every later one,
so a .config-audit-ignore pin retargeted to a neighbouring finding with no
version change at all.
Second measured arm: README already documented the opposite scheme. It and the
scanner headers describe ~20 numbers as check codes (CA-SKL-003 = oversized
body, CA-PLH-015 = folder shadowing, CA-TOK-006 = schema deferral), and the
counter could only produce those in the all-fire case -- source-order positions
are 4, 3 and 8. The documentation described the scheme; the implementation was
what was wrong. Every published number is preserved by construction and pinned
exhaustively in tests/lib/finding-codes.test.mjs.
scanners/lib/finding-codes.mjs is the single authority. Every finding() call
passes a `code`; an undeclared or missing one THROWS. No counter fallback --
that would reproduce D1's findGapId -> 'unknown' silent degradation and let a
half-converted scanner ship IDs that look valid. findingCounter/resetCounter
are deleted outright, not left as no-ops. Retirement is now a mechanism:
RETIRED_CODES tombstones a withdrawn key so its number is never reissued,
seeded with GAP t3_8 -- the D1 removal that opened this chunk.
IDs are consequently NOT unique per finding: one check failing in three files
emits three findings sharing an ID. That inverts which consumer is correct, so
every f.id/findingId site was classified before the change. diff-engine and
most of fix-engine already keyed on scanner+title+file (drift was never lying);
fix-engine's verification did not, and keyed on the ID alone -- fixing one of
two sibling instances marked both fixed, and the untouched one, still present
in the re-scan, was reported as a REGRESSION. Red test first, then keyed on
(findingId, file), which both planFixes and applyFixes already carry.
plugin-health's crossIds Set was measured and is a clean negative: cross
findings are allFindings.slice(crossPluginStart) and codes 18/19 are emitted
only in that tail, so the partition holds by construction.
unknownSuppressions() reports a pin that names no declared check, in the
--output-file payload (ux-rules rule 2 -- a stderr-only warning is invisible to
the commands) and only when one exists, so a clean config is byte-identical.
That is what makes the break safe: a stale pin goes loud instead of dying quiet.
Frozen tests/snapshots/v5.0.0/ untouched on disk. IDs are masked out of that
comparison (mask-finding-ids.mjs) rather than re-derived -- re-deriving
positional IDs would assert the retired scheme against itself, and #58's
isGapEntry off-by-one is the measured example of that misfiring. The dead
re-derivation is removed from strip-retired-gap.mjs. default-output snapshots
re-approved after confirming the diff is IDs and nothing else.
Guards, each seen red against its own defect: a missing code (scanner errors
out mid-sweep), an orphan declaration, a resurrected retired key, and a
documented ID naming no check. The sweep asserts the union across all 16
scanners, never per scanner -- a per-scanner assertion goes green on a partial
conversion.
Fasit written before implementation: docs/mbug28-id-semantics-fasit.local.md,
including one correction made before running (CML has 12 checks over 13 call
sites -- the anchored and calibrated char-budget arms are one check, which a
repeated-title sweep found and my call-site count had missed).
Suite 1535 -> 1573, 0 failing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MyqCQKK2ornJ1jFWwqx17E
118 lines
5.7 KiB
JavaScript
118 lines
5.7 KiB
JavaScript
/**
|
|
* AGT Scanner — Agent-listing always-loaded token budget
|
|
*
|
|
* Claude Code injects a listing of every active agent's name+description into the
|
|
* system prompt so it knows which subagents it can delegate to. That listing is
|
|
* re-sent on EVERY turn, whether or not a delegation happens — so with many
|
|
* installed agents it is a large always-loaded cost (on a heavily-plugged machine
|
|
* the dominant single always-loaded source).
|
|
*
|
|
* Detection:
|
|
* CA-AGT-NNN per-agent description over the soft bloat cap (low, advisory)
|
|
* CA-AGT-NNN aggregate agent-listing estimate exceeds the listing budget (low)
|
|
*
|
|
* Per-agent advisories are emitted FIRST (mirroring SKL 001→002). They are a
|
|
* soft bloat heuristic (the 500-char threshold TOK pattern F uses for SKILL.md
|
|
* descriptions), NOT a truncation finding — agents have no verified per-
|
|
* description cap, so nothing is dropped; the description is simply re-sent in
|
|
* full every turn.
|
|
*
|
|
* INTELLECTUAL-HONESTY CONTRACT (the reason this is `low`, not a hard finding):
|
|
* unlike the skill listing, the agent-listing mechanism is NOT documented (agents
|
|
* are absent from Claude Code's published context breakdown). So the finding is an
|
|
* INFERRED, UPPER-BOUND ESTIMATE, the budget is a config-audit heuristic (no
|
|
* documented agent allotment) anchored on a conservative 200k window, and the
|
|
* evidence discloses all of that. The cap, budget, and enumerate-and-measure step
|
|
* live in `lib/agent-listing-budget.mjs` — AGT only constructs the finding.
|
|
*
|
|
* Zero external dependencies.
|
|
*/
|
|
|
|
import { finding, scannerResult } from './lib/output.mjs';
|
|
import { SEVERITY } from './lib/severity.mjs';
|
|
import {
|
|
AGGREGATE_BUDGET_TOKENS,
|
|
BUDGET_CALIBRATION_NOTE,
|
|
PER_AGENT_DESC_SOFT_CAP,
|
|
measureActiveAgentListing,
|
|
} from './lib/agent-listing-budget.mjs';
|
|
|
|
const SCANNER = 'AGT';
|
|
|
|
/**
|
|
* Main scanner entry point.
|
|
*
|
|
* @param {string} _targetPath unused (agent listing is HOME-scoped)
|
|
* @param {object} _discovery unused (ignores project discovery)
|
|
*/
|
|
export async function scan(_targetPath, _discovery) {
|
|
const start = Date.now();
|
|
const findings = [];
|
|
|
|
const { agents, aggregate } = await measureActiveAgentListing();
|
|
|
|
// Per-agent advisory (emitted FIRST so the common "long agent + aggregate"
|
|
// case reads 001=per-agent, 002=aggregate, mirroring SKL). This is a soft
|
|
// heuristic, NOT a truncation finding — agents have no verified per-description
|
|
// cap, so the framing is "this is large and re-sent every turn", never "Claude
|
|
// Code drops the tail".
|
|
for (const agent of agents) {
|
|
if (agent.descLength <= PER_AGENT_DESC_SOFT_CAP) continue;
|
|
|
|
const sourceLabel = agent.source === 'plugin'
|
|
? `plugin:${agent.pluginName}`
|
|
: 'user';
|
|
|
|
findings.push(finding({
|
|
scanner: SCANNER,
|
|
code: 'description-bloat',
|
|
severity: SEVERITY.low,
|
|
title: 'Agent description is long (re-sent every turn in the always-loaded listing)',
|
|
description:
|
|
`Agent "${agent.name}" (${sourceLabel}) has a description of ${agent.descLength} ` +
|
|
`characters (>${PER_AGENT_DESC_SOFT_CAP}). Claude Code injects every active agent's ` +
|
|
'name+description into the agent listing on every turn so it knows which subagents it ' +
|
|
'can delegate to, so every character of this description re-enters context each turn ' +
|
|
'whether or not you delegate. Unlike the skill listing there is no verified ' +
|
|
'per-description cap, so nothing is dropped — this is a bloat advisory, not a ' +
|
|
'hard-cap finding.',
|
|
file: agent.path,
|
|
evidence:
|
|
`description_chars=${agent.descLength}; soft_cap=${PER_AGENT_DESC_SOFT_CAP} ` +
|
|
`(heuristic, same bloat threshold TOK pattern F uses for SKILL.md descriptions; agents ` +
|
|
`have NO verified per-description cap); agent="${agent.name}"; source=${sourceLabel}`,
|
|
recommendation:
|
|
'Trim the description toward its trigger phrases / "when to use this agent" cues and move ' +
|
|
'long examples into the agent body, or disable the plugin if you never delegate to this ' +
|
|
'agent (the whole agent block then leaves the always-loaded listing).',
|
|
category: 'token-efficiency',
|
|
}));
|
|
}
|
|
|
|
if (aggregate.overBudget) {
|
|
findings.push(finding({
|
|
scanner: SCANNER,
|
|
code: 'aggregate-listing-budget',
|
|
severity: SEVERITY.low,
|
|
title: 'Aggregate agent listing may exceed the always-loaded budget',
|
|
description:
|
|
`The ${aggregate.scanned} active agents carry about ${aggregate.aggregateTokens} tokens of ` +
|
|
'name+description text that Claude Code injects every turn so it knows which subagents it can ' +
|
|
`delegate to — above the ${AGGREGATE_BUDGET_TOKENS}-token budget this scanner anchors on a 200k ` +
|
|
'context window. Every one of those tokens is re-sent on every turn whether or not you delegate. ' +
|
|
'Note: unlike the skill listing, the agent-listing always-loaded mechanism is inferred (not ' +
|
|
'documented), so this is an upper-bound estimate (see evidence).',
|
|
evidence:
|
|
`active_agents_scanned=${aggregate.scanned}; description_chars=${aggregate.aggregateChars}; ` +
|
|
`description_tokens~${aggregate.aggregateTokens}; budget@200k=${AGGREGATE_BUDGET_TOKENS} tok; ` +
|
|
`over_by~${aggregate.overBy} tok - ${BUDGET_CALIBRATION_NOTE}`,
|
|
recommendation:
|
|
'Shrink the always-loaded agent listing: disable plugins whose agents you do not use (the whole ' +
|
|
'agent block leaves the listing), remove dead user agents from ~/.claude/agents/, and trim long ' +
|
|
'agent descriptions toward their trigger phrases.',
|
|
category: 'token-efficiency',
|
|
}));
|
|
}
|
|
|
|
return scannerResult(SCANNER, 'ok', findings, aggregate.scanned, Date.now() - start);
|
|
}
|