Compare commits

...

140 commits

Author SHA1 Message Date
caea8aca23 fix(commands): stop answering questions the caller did not ask
Dogfooding `campaign` + `knowledge-refresh` against a throwaway ledger. Seven
defects, all found by running the commands as written and measuring, not by
reading them.

The headline pair only existed together. `knowledge-refresh` built
`STALE_AFTER="--stale-after 30"` and expanded it unquoted, trusting the shell to
split it in two. bash does; zsh — the macOS default, and what the Bash tool runs
here — does not. The CLI got one argv entry, matched no flag, and because it had
no unknown-flag branch, silently kept the 90-day default and reported "✓ All 14
register entries were re-verified within the last 90 days": a true-sounding
sentence about a threshold the user had just overridden. Fixing either half alone
leaves a silent wrong answer or a loud one; both are fixed, and a guard now
rejects any template that packs a flag and its value into one variable.

`knowledge-refresh` also read one register and wrote another: step 6 named an
unanchored `knowledge/best-practices.json` while the CLI reads
`${CLAUDE_PLUGIN_ROOT}/…`, which for an installed plugin is the cache. The
validation gate then ran the cached test against the cached register — green no
matter what was written. The two copies were byte-identical that day, which is
exactly why it was invisible.

`campaign` vouched for repos it could not read. `add /finnes/ikke` returned
`added` + exit 0; `refresh-tokens` then put the phantom in `swept[]` with a
0-token delta and left `skipped[]` empty, so the machine-wide bill claimed
coverage of three repos on a machine with two. Paths stay tracked — an unmounted
volume is a legitimate absence — but are reported as `addedUnverified`, and the
command names them.

Two class sweeps, both measured rather than assumed. `posture` was the single
scanner (1 of 14) whose fatal catch exited 1, which ux-rules defines as a normal
WARNING grade — a crash indistinguishable from a result. And all 13 payload
writers failed on a `--output-file` whose parent did not exist, which on a fresh
machine turned `campaign`'s first run into "the ledger may be corrupt"; they now
share `scanners/lib/write-output.mjs`.

Predicted breadth was too wide for the first time in five sessions: 6 of 8 CLIs
predicted to lack unknown-flag rejection, 4 measured. `drift` and `fix` already
reject them, via a construct the grep did not recognise — a grep matches an
implementation, the invariant is a behaviour. The sweep was rewritten to run each
CLI with a bogus flag and read the exit code.

Suite 1453 → 1469/0. Frozen snapshots untouched. `optimize-lens-cli` and
`token-hotspots-cli` share the unknown-flag defect and are deferred to the v5.14
argument-handling chunk with their positional-swallow arm; the count is recorded
in the guard rather than rounded down to zero.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012NHWjN8EnoxSqRvMTLK2NE
2026-08-01 21:26:39 +02:00
acd1cf1248 fix(commands): stop writing files no later step can read, and payloads nobody asked for
Dogfooding the four read commands (posture, tokens, manifest, whats-active)
surfaced four defect classes, all in the seam between what a command template
promises and what the scanner behind it actually does.

M-BUG-40, fifth arm: posture wrote four temp files it could never read back.
#49 closed the $$/cross-block class in four commands, but posture survived it —
and so did the guard written to prevent exactly this. The guard compared each
$$ path to the block that created it, so a path written once and then read via
prose had no second occurrence to flag. Measured live: written from PID 21614,
read attempted from PID 23772. The invariant is now blanket (no $$ in any temp
path), which also caught fix.md and feature-gap.md.

M-BUG-43: 6 of 7 scanners write their payload to stdout when --raw/--json is
set even when --output-file was given, and the templates redirected only
stderr. Measured: posture 255 182 B, whats-active 35 922 B, drift 28 316 B,
manifest 23 825 B, tokens 8 768 B. fix and feature-gap never read the file they
wrote, so both recovered one letter grade from a quarter-megabyte dump.

tokens swallowed --json and --with-telemetry-recipe: documented, never
threaded, so --json returned the humanized payload where the docs promise
byte-stable v5.0.0 output.

M-BUG-42: manifest's render contract asked for {load}; the payload carries
loadPattern, so the Load column rendered blank for all 96 rows.

Four new tests (1449 -> 1453), each verified red before the fix. The
render-contract test checks {field} names against a live payload from a
fixture, since a hardcoded key list would drift. Frozen v5.0.0 snapshots
untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VGCk9o27eWo9uXLjkZTXEq
2026-08-01 20:40:07 +02:00
09f817977c fix(commands): stop assuming shell state survives between blocks
Dogfooding `plan` + `implement` against a throwaway config surfaced one root
defect with many arms: the command templates treat consecutive fenced blocks as
one shell. They are not. Every ```bash fence runs as its own Bash call in its own
process, so a variable set in one block is empty in the next, and `$$` is a
different PID (measured: 21710 vs 22109).

The planner agent confirmed the sharpest arm at runtime, reporting that
`Mode: $RAW_FLAG` "arrived literally unsubstituted" — `--raw` was documented in
three command files while being functionally dead. A machine sweep found the same
root in 20 places across 9 files, well past the two the written fasit predicted:

  - `$RAW_FLAG` read from non-shell agent prompts (analyze, plan, implement)
  - `$TMPFILE` read across blocks (tokens, manifest, whats-active,
    plugin-health) — each command could not read the file it had just written
  - `$GLOBAL_FLAG` across blocks (fix)
  - `$TODAY` never assigned in any block (campaign), passing
    `--reference-date ""` to a write CLI in six places
  - three `$$` temp paths handed to the Read tool (fix), which expands neither

All now follow the hardened drift.md pattern: a fixed literal path, or a
re-derivation inside each block that needs it.

Also fixed, all confirmed against ground truth rather than inferred:

  - `implement` printed a rollback ID it never captured (the timestamp lived only
    inside a command substitution) — the one message a user reads after a bad run
  - `plan` reported "No analysis results found" for valid sessions, because Read
    was pointed at a glob it cannot expand; now uses Glob and verifies the
    analysis report exists before spawning the agent
  - five phase commands wrote state.yaml with two of four required fields; since
    the agent writes all four, a follow-up write silently deleted the rest
  - `implement` promised rollback deletes created files; rollback deliberately
    leaves them (M-BUG-26 still open) — the doc, not the engine, was wrong
  - `implement` claimed a score delta with no pre-change measurement
  - `verifier-agent` was told to write a report it has no tool to write
  - dead `Task` tool name in always-loaded rule context; planner-agent template
    demonstrated the inline file content its own line 110 forbids

The sweeps land as tests/commands/command-shell-state-shape.test.mjs, verified
red before the fix and proven able to fail by reintroducing the defect. Two
existing tests asserted the old bash-block mechanism rather than the intent and
were updated. Suite 1449/0; frozen v5.0.0 snapshots and all scanner code
untouched.

Not fixed, deliberately: neither command scope-gates its actions to the audit
target. The generated plan included an edit to a real file under ~/.claude,
outside the throwaway target, because the skill/agent scanners are machine-wide.
That is a design change, not a side fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0195udHgCcFegzm7ecKku2Yc
2026-08-01 20:12:17 +02:00
de8a7b5d51 fix(scanners): stop discarding our own stdout when it is a pipe
process.exit() terminates immediately, but Node writes stdout asynchronously
when stdout is a pipe — everything still buffered is dropped. scan-orchestrator
measured 246 854 bytes to a file against 65 536 to a pipe (131 072 on another
run; the cut point is a flush race), so every machine consumer that pipes the
envelope got truncated, unparseable JSON. The failure reads like a corrupt file,
not like a cut-off, which is what made it survive this long. Reported by
org-ops, whose census pipes our output.

Closes the class rather than the one CLI where it was visible. campaign-cli,
campaign-export-cli, campaign-write-cli, knowledge-refresh-cli, drift-cli and
fix-cli all exited the same way on their success paths and were green only
because their payloads fit the pipe buffer today; size is not correctness. All
38 sites across 14 files now set process.exitCode and return, which is the
pattern self-audit.mjs already used.

Two contracts needed care rather than substitution: fail() is a never-returns
guard at ~25 call sites, so it throws a CliUsageError the top-level catch
renders with the identical "Error: " prefix and exit code 3; the path guards
needed an explicit return so main() stops instead of running on. Exit codes and
stderr text are unchanged, and the frozen v5.0.0 snapshots are untouched.

The class sweep is landed as a test, not as fourteen edits — it caught one site
this commit had missed. Suite 1443/0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B8sS1DuDV6bUJcyumLwbvj
2026-07-31 21:40:15 +02:00
b85919f2ec fix(commands): close the promises the command templates could not keep
DEL B chunk `interview` (+ discover/status/cleanup/help). Fasit written before
the run predicted 8 defects and refuted 4 candidates; all 8 confirmed, all 4
refutations held, and three predictions turned out too narrow.

- M-BUG-36: `drift --list` reached the command as 0 bytes. drift-cli accepted
  --output-file but list mode ignored it, and the listing goes to stderr, which
  the command discards per ux-rules rule 2. Fixing the caller alone would not
  have helped.
- M-BUG-37: feature-gap's "Create backup" step ran fix-cli without --apply.
  Dry-run is the default, so no backup existed (backupId: null) while the
  command went on to edit config believing it could roll back.
- M-BUG-38: fix-cli told users to recover with scanners/rollback-cli.mjs, which
  does not exist. Dead reference in the one message read after a bad fix.
- M-BUG-21 fourth arm: five templates carried literal [--global]/[--full-machine]
  inside executable bash blocks. A bracketed placeholder does not start with a
  dash, so every scanner's arg loop takes it as the scan target.
- interview and analyze never said which session they act on; interview could
  rewind a finished session; cleanup interpolated an unvalidated id into rm -rf
  (an empty id deletes every session); status advertised a `resume` command that
  does not exist and documented an `all` argument it never parsed.

TDD: 9 red tests first, including a machine sweep for dead /config-audit
references and for bracketed flags in bash blocks. Suite 1432 -> 1441/0.
Frozen v5.0.0 snapshots untouched; --raw/--json contracts unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UGvA1uUQn2hPBPMaCKK6x3
2026-07-31 21:27:07 +02:00
001090261e fix(plugin-health): make the command able to read what the scanner found
Dogfooding `/config-audit plugin-health` against a fasit registered before the
run: 11 of 12 predictions confirmed, 1 refuted with evidence, 0 deviations.
The command's default path could not produce the report it documents.

M-BUG-21 (third arm): the argument loop ended in
`else if (!args[i].startsWith('-')) targetPath = args[i]` with no unknown-flag
branch, so `--output-file /tmp/x.json` was dropped and its value became the scan
target. Worse than in drift-cli: a non-existent path discovers no plugins, so the
scanner answered "No plugins found" (info) with exit 0 — a reassuring answer, not
an error. Unknown options and a value-less `--output-file` now exit 3.

M-BUG-33: the scanner had no `--output-file` and its default-mode report goes to
stderr, which `commands/plugin-health.md` discards with `2>/dev/null` before
telling the agent to read stdout. Zero bytes captured.

M-BUG-34: per-plugin rows and the grade formula never left `scan()` — the only
grade code, `formatPluginHealthReport`, had no caller — and cross-plugin findings
were flattened behind a `category` they share with per-plugin findings. The
mandated table and Cross-Plugin section were unbuildable, so the command had to
fabricate them. `scanDetailed()` now returns them; `scan()`'s frozen v5.0.0
envelope is unchanged by construction.

M-BUG-35: `.claude-plugin/marketplace.json` was flagged as an unknown file. It is
the documented catalog location, and `"source": "./"` makes the repo root its own
plugin, so one `.claude-plugin/` legitimately holds both.

Also: `commands/posture.md` ran both optional scanners in default mode under
`2>/dev/null` and read stdout — the same class as feature-gap.md:133 in the fix
chunk. A CLI-side flag fix does not close its callers.

Tests 1420 -> 1432, red first. Frozen v5.0.0 snapshots untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XhhZ8zg1amR7YCAPqiZWdt
2026-07-31 21:08:32 +02:00
05f1e954d0 fix(fix): validate the arguments, back up renames, and verify the scope it fixed
Dogfooding `/config-audit fix` against a throwaway repo copy. All eight
predictions registered in the fasit before the run were confirmed, and three
further defects surfaced that were not predicted.

- M-BUG-21, third arm: the argument loop ended in `!arg.startsWith('-') =>
  targetPath`, so an unknown flag was dropped and its value became the target.
  In `fix` that is the WRITE target under `--apply`. Unknown options and a
  value-less `--output-file` now exit 3.
- `--dry-run` was documented in the command's argument-hint and never
  implemented; `--output-file` did not exist, so `commands/fix.md` told the
  agent to Read a file nothing produced. Both now exist.
- M-BUG-31: `file-rename` was excluded from the backup set, so a renamed rule
  file had no backup entry while the command promised one and returned a
  backupId that could not restore it.
- M-BUG-32: `verifyFixes` hardcoded `includeGlobal: false`, so after a
  `--global` run every untouched user-scope finding was reported as verified.
  Reproduced against an unmodified ~/.claude/CLAUDE.md.
- M-BUG-29: a rename was applied before other fixes on the same file, which
  then failed with ENOENT while the run still exited 0. Renames sort last.
- M-BUG-30: `severityOrder[s] || 4` maps critical (0) to 4, so critical fixes
  sorted last. The old test used the same falsy fallback and agreed with the
  bug. Now `?? 4`.
- A failed fix exits 2 instead of 0, matching the other scanners' convention.

Frozen tests/snapshots/v5.0.0/ untouched; --json/--raw stdout byte-identical.
Suite 1420/0 (+10).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YJ3MCDCnyw7wZSPnUXVhYS
2026-07-31 18:41:27 +02:00
1182f85767 fix(drift): validate the arguments and the baseline anchor drift never checked
Dogfooding `/config-audit drift` against the machine. Fasit written before the
run; 6/6 predictions plus both F7 arms confirmed, 0 deviations.

M-BUG-21 (both arms):
The arg loop ended in `else if (!arg.startsWith('-')) targetPath = arg` with no
unknown-flag branch, so an unrecognised flag was dropped silently and its VALUE
became the scan target. `--output-file /tmp/x.json` scanned /tmp/x.json — a path
that does not exist — and reported the near-empty scan as drift, forever. The
same silence was destructive for `--save --name` with the value omitted: the
name stayed `default` and an existing baseline was overwritten. And the flag
ux-rules rule 2 requires did not exist at all: commands/drift.md ran the CLI
under `2>/dev/null` while telling the agent to read stdout, but the default
report, the --save confirmation and --list all write to stderr. All three modes
captured nothing.

M-BUG-27 (found during the run, not predicted):
diff-engine never compared the baseline's stored target_path against the current
target. The machine's `default` baseline is anchored to a test fixture, so
`/config-audit drift` diffed two unrelated trees, marked all 20 baseline
findings resolved and all 15 current ones new, and reported trend "improving".
A reassuring, entirely false signal — and the default path.

Root cause is one thing, not three: the CLI validated neither its flags nor its
anchor. Same class as the rollback chunk's "nothing agreed where a backup lives".

Fix: unknown options and value-less --name/--baseline/--output-file exit 3;
--output-file follows the posture.mjs pattern; `_baselineAnchor` rides in the
default-mode payload (stderr alone is invisible under `2>/dev/null`) while
--json/--raw stdout stays v5.0.0-shaped.

Suite 1410/0 (+12). Frozen snapshots untouched; raw/json/default backcompat green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015RGrL3noVdFSUhohaKTMSN
2026-07-31 18:20:04 +02:00
c60f849d2e chore(release): v5.13.0 — pipeline hardening batch (M-BUG-11..25 + --subtract)
Batch release of everything since v5.12.5: one new lens mode and 14 real bugs,
all of them dogfooding finds — either from running the plugin against the
maintainer's real machine, or from walking analyze -> plan -> implement ->
rollback end-to-end on a throwaway repo copy.

Minor, not patch. STATE recorded this batch as "fix: only"; the log says
otherwise — e9921d3 ships `optimize --subtract`, a user-facing opt-in flag, so
semver requires a minor. Verified by reading `git log v5.12.5..HEAD` rather than
trusting the note: 11 fix, 1 feat, 8 docs.

Consequence: the planned v5.13 work (model routing, effort awareness, dead
references) now targets v5.14. docs/v5.13-model-routing-effort-deadref-plan.md
keeps its filename so existing references resolve, and says so at the top —
leaving a doc named for a version that shipped something else is exactly the
dead-reference class that plan is about.

Gates, all re-run against ground truth before writing anything:
- node --test 'tests/**/*.test.mjs' -> 1398 pass / 0 fail
- scanners/self-audit.mjs --check-readme -> passed (tests badge 1344 -> 1398;
  the path in STATE said scripts/, which does not exist)
- catalog check-versions.mjs -> 0 ERROR

Counts unchanged: scanners 16, agents 7, commands 21, hooks 4.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AA5LT1UaDzctNkMi414qzA
2026-07-31 17:33:48 +02:00
8f149891c9 fix(rollback): restore the backup path contract the engine and the commands disagreed on
Pipeline step 4 dogfood. `/config-audit rollback` could not see a single one of
the four real backups on this machine, and reported "Backup not found" for one
that was sitting right there.

Four defects, one root: nothing agreed on where a backup lives or what its
manifest looks like.

- M-BUG-22 `lib/backup.mjs` resolved `~/.config-audit/backups` (pre-v2.2.0)
  while every command, agent and doc uses `~/.claude/config-audit/backups`.
  The auto-backup hook and fix-cli wrote to the first, implement to the second,
  rollback read only the first. Canonical root now, with the legacy root kept
  readable so older backups stay listable and restorable (`legacy: true`).
- M-BUG-25 `parseManifest` understood only the engine's quoted `original_path:`
  spelling, but implement hand-builds its manifest with `- backup:`/`original:`/
  `sha256:`. Every implement-made backup parsed to zero files and restoreBackup
  returned `{restored: [], failed: []}` — a success-shaped no-op. Both formats
  parse now, and a manifest with unparseable entries throws instead of
  pretending to succeed.
- M-BUG-23 both session hooks watched `~/.config-audit/sessions`, which does not
  exist; sessions live under `~/.claude/`. "Check for active sessions" had never
  fired once. It fires now.
- M-BUG-24 the suite called createBackup() against the developer's real home —
  it had left nine stray backups there, and cleanupOldBackups() deletes past ten.
  Root is overridable via CONFIG_AUDIT_BACKUP_ROOT; both test files use it.

Rollback still cannot delete files implement CREATED — no backup can hold a file
that never existed. It no longer does so silently: manifests carry a `created:`
list, restoreBackup returns `createdNotRemoved`, and rollback.md requires the
report. Automatic deletion is a destructive action and needs its own design.

Verified against backup 20260717_032636 on a throwaway copy: all three files
restore byte-exact (sha256 match), zero writes outside the copy, backup dir
unmodified. Suite 1382 -> 1398/0; frozen v5.0.0 snapshots untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SejM9RQAa1Hfuq7Ek2WfFr
2026-07-31 17:23:30 +02:00
b8cbbc0f5b docs(subtract): correct the ground-truth counts and keep the gate re-runnable
Two record fixes, no behaviour change.

The README quoted "35 blocks, 13 genuinely ambiguous" for the hand-built ground
truth. Counting its own rows gives 48 classified blocks and 19 ambiguous — the
headline in the fasit disagreed with the table beneath it, apparently by
collapsing letter-suffixed sub-blocks for the summary while listing them
separately. Labels are untouched in both files; only the counts are restated,
and now as row counts, which are reproducible with a grep rather than by
recounting a classification.

The dogfood gate script that proves the blocking §8 criterion lived only in the
session scratchpad, which would have made "zero load-bearing blocks proposed,
11/18 groups, ~756 tok" unverifiable claims the moment the session ended —
precisely the premise-not-fact class this repo's own rules warn about. It now
lives at scripts/dogfood-subtraction-gate.local.mjs, gitignored via a new
*.local.mjs pattern because it indexes the operator's private config by line
number and must never reach the public mirror.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RW2haJXbxZpKivKHseSXNh
2026-07-31 16:28:38 +02:00
e9921d3c9d feat(optimize): add --subtract, the subtraction axis, behind a deterministic floor
Every command so far asked an addition question — what to add, what to move,
what it costs. Nothing asked what is no longer earning its always-loaded rent.
This adds that axis as a fourth lensCheck on the existing hybrid motor rather
than a new scanner or a 22nd command: the measured payoff (~18% of one file)
justifies a mode, not machinery.

It is the only lens that proposes REMOVING config, so it carries a guarantee
the others don't need: a load-bearing block is never a candidate. Precision is
asymmetric — a missed dead line costs a few tokens per turn, a deleted one
costs a wrong remote or a broken script — so the floor is decided in code
(lib/floor-exclusion.mjs) before the opus judge sees anything, never in prose.

Granularity is the leaf block, with two structural exceptions: a paragraph
ending in ':' merges with the list it introduces, and an ordered list is a
contract whose steps inherit floor from any sibling. Unordered lists
deliberately do not inherit — a load-bearing bullet and a disposable one
routinely share a list, and container-reasoning is the error the hand-built
ground truth exists to catch.

Verified against that ground truth (built before any classifier existed), with
the comparison machine-checked rather than read by eye: zero load-bearing
blocks proposed, 11/18 deletable groups surfaced, ~756 tok ~ 18% of a ~4300
token file — inside the pre-registered band. The first run found five floor
violations the synthesized fixture missed; each got a structural rule and a
fixture shape so it cannot regress.

Three real bugs the dogfood run exposed, all now covered:
- JS \b is ASCII-only, so /\bunngå\b/ never matches — every Norwegian keyword
  ending in æ/ø/å was silently dead.
- A bare word/word is not a path; "pros/cons" vetoed the largest deletable
  block until PATH_RE was tightened to rooted paths and globs.
- "Mid-sentence" must key on a preceding lowercase letter; the loose version
  read **bold labels:** and quoted openers as entities, costing 4 of 11 groups.

BP-SUB-001 is grounded entirely in the Anthropic steering blog already cited by
BP-MECH-001..004 and asserts nothing from the talk that motivated the feature —
no "80%", no ablation figure.

Suite 1365 -> 1382/0. Frozen v5.0.0 snapshots untouched; plain optimize output
byte-identical on identical input (--subtract adds keys only when passed).
knowledge-refresh-cli's reference date moved to 2026-08-01: its premise that
every seed entry was verified 2026-06-20 expired when BP-SUB-001 got a genuine
verification date, and backdating the entry to fit the test would have been a
lie about when its source was checked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RW2haJXbxZpKivKHseSXNh
2026-07-31 16:24:51 +02:00
c0625c568f docs(plan): age signal measured dead; close §7 q1-q3 on optimize --subtract
Ran the brief's own §8 pre-build checks before writing any code, and two of
them changed the design.

Age signal (§5A, §8): per-line git blame over four real instruction files gives
single-date shares of 88% / 55% / 100% / 58%. Instruction blocks trace back to
bulk commits, and blame reports last-touch rather than vintage — a reformatting
commit (this repo's own 96e32df) makes old instructions look young, so the
signal is biased, not merely sparse. That also kills the mtime fallback. §8
pre-registered this exact outcome and its consequence, so shape A does not ship
as a CA-VIN-* vintage scanner.

Premise correction (§6.1): ~/.claude IS git-tracked as of 2026-07-26 (7 commits,
remote on an external backup volume, 47 files). The rule it justified — mv to
_archive/, never rm — stands on different grounds and is unchanged.

§7 q1: both scopes, user-level mandatory in v1 (the floor test is defined there
and the always-loaded cost sits there).
§7 q2: not deterministically classifiable — the deciding blocks require reading
content against container. Deterministic pre-filter -> precision-gated judge,
with floor-exclusion running BEFORE the judge so a load-bearing block is never a
candidate.
§7 q3: ships as /config-audit optimize --subtract, a fourth lensCheck class
emitting CA-OPT-* with a BP-SUB-001 register rule. No new scanner, no new
command, no badge bump, no snapshot risk. Proportionality decided it: the
hand-built fasit puts the honest payoff at ~850-1400 always-loaded tokens on a
~4300-token file (~20%, not 80%), with 26 of 34 blocks classified floor.

The fasit itself (34 blocks, 13 marked ambiguous per §7.2) is local-only — it
quotes the operator's global CLAUDE.md verbatim and this repo's only remote is
the public open/ mirror.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011x2i9NDEtMR9FU7ufaCnBX
2026-07-31 15:43:03 +02:00
31073c2178 docs(plan): fasit must cover the ambiguous middle, not the poles (§7.2)
§8's floor gate rests on §6.0's classification test, which is a judgment call
rather than a mechanical one. The four named must-survive items ("only Forgejo",
"bash is 3.2", the test command, "~/.claude is not git-tracked") are clear-cut —
any mechanism gets them right, so a fasit built from them proves nothing.

The gate is actually decided by blocks like "Conventional Commits:
type(scope): beskrivelse" (local convention or a nag the model follows anyway?),
"commit ofte med beskrivende meldinger", or the model-routing rubric — local
policy that reads like generic advice. §7.2 now requires 3-5 such blocks in the
fasit deliberately.

Cross-references verified: every section pointer in the brief (§1-§8, §6.0, §5A,
§3.1, §7.2) resolves to an existing heading. Shipping a dead prose reference in
the brief that proposes detecting them would have been an odd artifact — that is
CA-CML's finding class, v5.13 chunk 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNrtHo9hKSLKNyMS6b4Zuy
2026-07-29 09:57:07 +02:00
36c55fb167 docs(plan): add the floor constraint — compensatory vs load-bearing (§6.0)
Operator corrected two things about the brief committed in 3086e8b/3252b51.

Provenance: the operator watched the recording and identifies Boris Cherny on
stage, so the attribution is confirmed by direct observation, not a channel's
claim. The verbatim figures (80 %, "more intelligent without the prompts") still
reach us through the summary's editing and stay at that confidence level. §1 now
carries both levels separately, and the register source string reflects the split
instead of flattening to "unverified".

Design: "start with what it must have" is the constraint the whole feature turns
on, so it is a hard constraint (§6.0), not a candidate-shape detail. Model
capability erodes compensatory instructions ("read the whole file first") and
does nothing to load-bearing local facts ("only Forgejo", "bash is 3.2", the test
command) — the model isn't failing at intelligence there, it cannot know. A tool
that treats them alike deletes the Forgejo constraint because Opus 5 "is smart
enough now". Rebuild is therefore three tiers: floor restored immediately, earned
returns on repeated stumbling, dead never comes back. Policy prohibitions stay in
the floor by decision rather than classification — asymmetric cost, cheap to keep.

Consequences threaded through: §5 disqualifies any shape that cannot express the
distinction, §7.2 becomes the core open question (age and class are independent
signals, so age alone can never carry the call), and §8 gains a blocking floor
test with a hand-built fasit and named must-survive items.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNrtHo9hKSLKNyMS6b4Zuy
2026-07-29 09:55:22 +02:00
3252b514ed docs(plan): close ordering question, pin the untested age premise
Two corrections to the brief committed in 3086e8b.

The ordering question (§7.4) still pointed at STATE.md for a decision the
operator had already made, and STATE.md is gitignored — so the tracked artifact
carried a stale queue and deferred to a file a future session cannot read. It
now records the decision inline: delete-and-rebuild goes ahead of pipeline step
4, prior order stands underneath.

§8 gains the premise the brief was quietly resting on: §5A claims project-level
CLAUDE.md/rules yield usable per-block git ages. That is untested. If instruction
blocks trace to one bulk commit, the age signal carries no information and shape
A collapses to the phrasing heuristic — which would answer §7.2 for us. Checked
before any scanner code, not after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNrtHo9hKSLKNyMS6b4Zuy
2026-07-29 09:50:33 +02:00
3086e8bb11 docs(plan): brief for delete-and-rebuild (config subtraction axis)
Operator relayed the "delete your CLAUDE.md every six months" idea from a
third-party summary of a Boris Cherny talk. Assessed rather than adopted: the
provenance is secondhand and deliberately kept non-load-bearing (this repo has
one scar from treating a plausible quote as fact), while the feature is argued
from the repo's own logic.

The gap is real and verified, not assumed: feature-gap has no inverse, and grep
over scanners/ confirms nothing measures instruction AGE — 'stale' appears only
for knowledge-register entries and plugin-cache versions. drift's saveBaseline/
diffEnvelopes plus backup.mjs/rollback-engine.mjs are already the undo
machinery that makes deletion a measurement rather than a gamble.

CLI ground truth checked against claude --help: --bare (sets
CLAUDE_CODE_SIMPLE=1), --system-prompt, --setting-sources, --add-dir. So the
env var the video calls undocumented is a documented flag here, which is what
would make an ablation harness buildable.

Brief only — no code, no chunk breakdown, no version committed. Three candidate
shapes with a likely landing (deterministic CA-VIN vintage scanner first),
hard constraints, open questions, and testable verification criteria.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JNrtHo9hKSLKNyMS6b4Zuy
2026-07-29 09:49:25 +02:00
b4d819b72d fix(acr): pin Bash >> append discipline on shared implementation log (M-BUG-20)
implement.md spawns implementer agents in parallel batches, all appending to
the same implementation-log.md. Dogfooding showed agents satisfying 'Append
result to:' with a full-file Write — the last writer clobbered 4 of 6 entries.
Pin the mechanism in both contracts: append with Bash >> heredoc, never the
Write/Edit tool on the shared log. Shape tests pin the instruction in both
files (empirically verified: agents given the >> instruction appended safely).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 03:53:06 +02:00
0cd87e0597 fix(rul): globToRegex corrupts mid-pattern /**/ globs (M-BUG-19)
The ? -> [^/] replacement ran AFTER the {{GLOBSTAR_SLASH}} placeholder was
restored to '(?:/.+/|/)', corrupting the group opener '(?:' into '([^/]:' —
every rule pattern containing a mid-pattern '/**/' silently matched only the
zero-dir branch and live rules were flagged 'matches no files' (CA-RUL).
Found by dogfooding /config-audit implement on a throwaway repo copy: the
implementer agent's correct 'posts/**/post.md' rule was flagged dead.

Fix: run the ? replacement before placeholder restoration. Fixture outcomes
byte-identical; frozen v5.0.0 baselines untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 03:53:06 +02:00
4b7b2d9c48 fix(acr): analyze persists agent-returned report (M-BUG-18)
The Claude Code subagent harness instructs spawned agents NOT to write
report/summary/findings/analysis .md files — the parent reads the final
text message. Verified live: analyzer-agent skipped Write entirely and
returned the report inline, so analysis-report.md never landed on disk
and the plan/interview/status phases would find nothing to read.

New contract (orchestrator-writes pattern): analyzer-agent returns the
complete report as its final message; the analyze command saves it
verbatim to the session directory before presenting the summary.

Same class exists in plan/feature-gap/optimize/scanner agent pairs —
deliberately left for their own dogfood chunks (plan is judged as-is
first per the pipeline sequence).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CTontYwY5JGS4nL2AuiASy
2026-07-16 20:24:12 +02:00
69a4654dd7 docs(plan): v5.13 plan — model routing, effort awareness, dead references
Video-derived audit ('The Model Isn't the Moat') cross-checked against
primary sources. Verified: orchestrator+cheap-worker pattern and 5-level
per-agent effort tuning (official docs); rejected: the 'Fable low ≈ Opus
high' chart claim (contradicted by Anthropic's own pages). Five chunks:
register entries BP-MODEL-001/002, fix-engine xhigh hygiene, CA-CML dead
prose references, feature-gap model/effort opportunity, planner-agent
adversarial gate. Sequenced AFTER DEL B pipeline dogfood + batch release.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CTontYwY5JGS4nL2AuiASy
2026-07-14 10:45:57 +02:00
97867dbf37 fix(acr): AGT findings humanize to "Wasted tokens" not "Other" (M-BUG-17)
The agent-listing scanner (AGT) emits an always-loaded per-turn token cost
("Agent description is long, re-sent every turn in the always-loaded listing";
scanner category 'token-efficiency', "the dominant single always-loaded
source"). But SCANNER_TO_CATEGORY in humanizer.mjs had no AGT entry, so its
findings fell through to the 'Other' fallback (humanizer.mjs:140) — a bucket
that isn't even in the analyzer-agent's category list. Neither the scanner
prefix nor the per-finding category ('token-efficiency' is not in
CATEGORY_TO_IMPACT) resolved AGT to its true impact.

Same class as M-BUG-16/15: a finding type without its matching humanizer
mapping landing on a default that mismatches its own evidence. The analogous
SKL body finding correctly buckets "Wasted tokens"; AGT (the same always-loaded
token-waste mechanism) silently landed under the meaningless "Other".

Found during analyze-prep premise-verification of the linkedin-posts scan: the
3 AGT findings bucketed "Other" while the analogous SKL findings bucketed
"Wasted tokens".

Fix: add AGT: 'Wasted tokens' to SCANNER_TO_CATEGORY, alongside TOK/CPS/SKL.
RED-first (extended the Wasted-tokens category test to include AGT; the 'Other'
fallback test still uses a synthetic 'XXX' scanner, unaffected). Frozen v5.0.0
untouched (AGT post-dates it; humanizer bypassed for --raw/--json); no
default-output snapshot contains AGT -> 0 regen. Suite 1359/0.

Verified end-to-end on linkedin-posts: 3 AGT findings now "Wasted tokens", 0
"Other" remaining, 28 findings unchanged (category-only). All 16 orchestrator
scanner prefixes now covered by the category map (class closed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01683eAqVecv9VZfQzL8CQ9h
2026-06-30 13:51:55 +02:00
239e88cecb fix(acr): on-demand copy for oversized skill-body finding (M-BUG-16)
The skill-listing check emits a third finding for an oversized skill BODY
(v5.11 B7, RAW title "Skill body is large (loads on demand when the skill
runs)"). The body is an ON-DEMAND cost — it loads only when the skill is
invoked, not the always-loaded listing Claude reads every turn. The scanner is
careful to distinguish the two (RAW title + comment + evidence note).

But the humanizer-data SKL.static map had no entry for this title, so it fell
through to SKL._default ("A skill is using more of the listing budget than it
should"). The humanized title therefore claimed a listing-budget cost and
directly contradicted the finding's own humanized evidence ("loads ON DEMAND
only ... NOT every turn like the always-loaded listing") — the same internal
contradiction class as M-BUG-15/M-BUG-14, and the same "new finding type added
without a matching humanizer entry" gap the scanner checklist warns about.

Found by finding-granularity premise-verification of the linkedin-posts scan
before feeding it to the analyze pipeline (the prior session's pass focused on
the GAP findings and did not catch the SKL fall-through).

- humanizer-data.mjs: add SKL.static entry for "Skill body is large (loads on
  demand when the skill runs)" with on-demand-correct title ("A skill's body is
  large (it loads only when that skill runs)"), description, and recommendation.
  No listing-budget language; tier1/tier3 forbidden-word checks pass.
- RED-first tests at both layers: humanizer.test.mjs (humanizeFinding path:
  title is not the listing-budget _default, conveys on-demand body) and
  humanizer-data.test.mjs (static entry exists, on-demand-correct).

RAW envelope unaffected (humanizer bypassed for --raw/--json), frozen v5.0.0
snapshots untouched, default-output fixtures contain no oversized-body skill so
no snapshot regen. Suite 1357->1359/0 (+2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01683eAqVecv9VZfQzL8CQ9h
2026-06-30 13:30:22 +02:00
2975b0563f fix(acr): honest absence-state copy for two GAP enhancement findings (M-BUG-15)
The t2_3 ("No path-scoped rules") and t3_6 ("No subagent isolation")
feature-gap checks iterate a collection (rule files / agent files) and return
false for an EMPTY one — so they fire even when the user has zero rules / zero
subagents, the same state their presence-gap siblings flag. But the humanized
titles presupposed the feature already exists:
  - "Your rules all load on every conversation" (with zero rules)
  - "Your subagents share Claude's main work folder" (with zero subagents)
The second directly contradicts GAP-005 "You haven't set up any specialized
helper agents yet" in the same report — a user cannot simultaneously have no
subagents and have subagents that lack isolation. Found by dogfooding the
analyze pipeline against linkedin-posts (premise-verifying each finding before
trusting the analyzer-agent's report).

Decided fix = align the two titles with the house "You haven't set up X yet"
absence framing (state-neutral: honest for both the zero-state and the
has-but-unconfigured state), NOT a base-feature presence gate in the scanner —
that would change the frozen v5.0.0 marketplace-medium baseline (zero
rules/agents, freezes these gaps firing) and break the RAW byte contract.
Mirrors M-BUG-14's humanizer-layer, copy-only approach.

- humanizer-data.mjs: title "Your rules all load on every conversation" ->
  "You haven't set up path-scoped rules yet"; title "Your subagents share
  Claude's main work folder" -> "You haven't set up subagent isolation yet".
  description + recommendation unchanged. t3_5/t3_7 ("Your skills don't ...")
  left as-is: their possessive is correct in the common case where skills
  exist; the zero-skills edge is latent, not manifest here.
- RED-first unit test pins both titles existence-neutral (forbids "your rules
  all load" / "your subagents"; feature still named).

Frozen v5.0.0 snapshots untouched (RAW bypasses humanizer); default-output
snapshot regenerated (2 titles only, no collateral). Verified on linkedin-posts:
GAP-004/005/011 now all consistently absence-framed. Suite 1356->1357/0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01683eAqVecv9VZfQzL8CQ9h
2026-06-30 12:57:40 +02:00
eb0b3fd29d fix(acr): size-neutral copy for "CLAUDE.md not modular" GAP (M-BUG-14)
The t2_2 feature-gap check is a pure presence check (rules file OR @import
present) with no length gate, consistent with its t2_3/t2_4/t2_5 siblings.
But the humanized copy claimed the file is "one big block" and that splitting
makes it "easier on the loading time" — an unconditional size/load-cost
overclaim. For a ~625-token non-modular CLAUDE.md the load saving is trivial,
so the description lied. Found by dogfooding feature-gap against linkedin-posts.

Decided fix = copy-softening, NOT a length gate in t2_2 (that would make it the
only size-gated check among the presence-check siblings + risks byte-stability
if a short non-modular CLAUDE.md fixture exists in snapshots).

- humanizer-data.mjs: title "one big block" -> "all live in one file";
  description drops "long"/"loading time", keeps the honest structural
  "split into linked files with @import or .claude/rules/" framing.
  recommendation unchanged.
- RED-first unit test pins the size-neutral copy (no "big"/"long"/"loading
  time" in title/desc, structural split framing kept).

Frozen v5.0.0 snapshots untouched (RAW envelope bypasses humanizer); SC-5
default-output byte-stable (no non-modular GAP finding in those fixtures).
Suite 1355->1356/0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01683eAqVecv9VZfQzL8CQ9h
2026-06-30 11:02:45 +02:00
f4bf3ae2cb fix(acr): feature-gap scopes presence checks to authored config + reads settings cascade (M-BUG-13)
The GAP scanner's 25 presence checks ran over the full includeGlobal discovery, so
this plugin's own examples/optimal-setup (vendored across plugin-cache versions)
satisfied every tier-3 check — masking real feature gaps to GAP=0 on ANY target.
And the real ~/.claude/settings.json is invisible to the settings-key checks
(includeGlobal gotcha + maxFiles cap), which would flip statusLine/autoMode to
false positives once the maskers were removed.

- isAuthoredConfig: exclude plugin-bundled (~/.claude/plugins/) + nested examples/
  and tests/fixtures/ (relPath-relative, so a fixture scanned AS the target keeps
  its own files) from ctx.files + parsedSettings.
- readSettingsCascade: read the user->project->local settings cascade directly and
  merge into parsedSettings — immune to the discovery cap/gotcha.

Empty target: ~0 (masked) -> 18 humanized opportunities; no statusLine/autoMode
false positives. Frozen v5.0.0 snapshots + SC-5/6/7 byte-stable (marketplace-medium
has no nested demo trees; hermetic-HOME cascade adds nothing). Suite 1350->1355/0.
Found by dogfooding feature-gap against the machine.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01683eAqVecv9VZfQzL8CQ9h
2026-06-30 10:09:10 +02:00
b58393099a fix(acr): posture --output-file humanizes findings in default mode (M-BUG-12)
feature-gap.md (Step 3-4) and posture.md (Step 3-4) read findings from
`posture.mjs --output-file` and group on the humanizer fields
(userActionLanguage / userImpactCategory / relevanceContext). But posture.mjs
only humanized the stderr scorecard — its --output-file JSON wrote the raw
v5.0.0-shape `result`, so every finding's humanizer fields were `undefined`.
Both commands silently degraded to the raw tier-fallback: v5.1.0 plain-language
output was dead for feature-gap and for posture's finding-level grouping.

Re-derived on tests/fixtures/marketplace-medium: 17 GAP findings, all three
humanizer fields undefined in the default --output-file JSON.

Fix (posture-CLI-local, surgical): humanize the output-file payload in default
mode, mirroring scan-orchestrator.mjs:277 — but posture nests the scanner
envelope under `result.scannerEnvelope` (its `result` has no top-level
`scanners` array), so humanizeEnvelope is applied to `result.scannerEnvelope`,
not `result` (the latter would no-op). --json / --raw stay raw, so the
explicit-v5.0.0-shape contract and snapshot byte-compat are preserved.

TDD: red-first test in posture-humanizer.test.mjs default-mode block asserts
GAP findings in the output file carry userActionLanguage/userImpactCategory;
a --raw --output-file guard asserts the raw shape is unchanged.

Suite 1350/0 (+2). Frozen v5.0.0 + SC-5/6/7 + default-output snapshots
byte-stable: --json/--raw bypass the humanizer (their snapshot tests use those
flags), and the default --output-file JSON is not snapshot-pinned. Committed,
not released — batches with M-BUG-11 in a later hardening release.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01683eAqVecv9VZfQzL8CQ9h
2026-06-30 09:10:07 +02:00
1d63492617 fix(acr): optimize lens scopes out plugin-bundled CLAUDE.md + unique candidate paths (M-BUG-11)
The optimize lens CLI fed its precision-gate agent every CLAUDE.md that discovery
returned, including the 256 files under ~/.claude/plugins/ — vendored plugin
CLAUDE.md across every cached version (7 config-audit, 6 ms-ai-architect, 5 okr,
...) plus their bundled tests/fixtures and examples. Running `optimize --global`
on this machine produced 454 candidates across 92 "files", ~250 of them sourced
from plugin-internal files a user cannot act on (the plugin overwrites them on
update). Same class as M-BUG-2: plugin-bundled config is not the user's cascade.

Second defect: candidates were keyed by `relPath || absPath`, but relPath
collides across scopes — a repo-root `CLAUDE.md` and the user-global
`~/.claude/CLAUDE.md` both relPath to `CLAUDE.md`. The two files that actually
matter were merged into one indistinguishable bucket (21 candidates), the agent's
Read(file) would resolve the wrong one, and cache-file relPaths were not readable
relative to cwd at all.

Fix (lens-CLI-local, surgical):
- Filter isPluginBundled (absPath under `.claude/plugins/`) from discovery for
  BOTH halves of the motor (candidate loop + the OPT scanner, which reads
  discovery.files directly). Drops vendored files regardless of active/stale
  version, so excludeCache is unnecessary here.
- Key each candidate by absPath: unique + readable.
No change to file-discovery.mjs or the OPT scanner, so their byte-stable
snapshots are untouched.

Suite 1348/0 (+4: candidate scoping, real-config survives, deterministic scoping,
absolute-path identity). Frozen v5.0.0 + SC-5 snapshots untouched (the lens CLI
has no snapshot; the command is agent-driven, not byte-stable). Dogfood ~/.claude
`optimize --global`: candidates 454->45, deterministic 2->0 (both were stale
plugin-cache copies), distinct files 92->11, repo vs user-global now distinct
(12 + 9 = 21). Residual 45 includes config-audit's own tests/fixtures CLAUDE.md
(repo-specific dogfooding artifact, not a general bug — left alone).
2026-06-30 06:44:36 +02:00
96e32df87b docs(claude-md): trim project CLAUDE.md to invariants (−662 always-tok)
The plugin's own CLAUDE.md is loaded every turn while working in this repo
(measured 2,178 always-loaded tokens via `manifest`, the largest slice of the
2,745-tok project delta). Much of it was reference-grade prose that duplicates
README.md, `/config-audit help`, and docs/ — verbose per-command feature lists,
the full plain-language-output spec, the session-dir ASCII tree — none of which
is invariant for working on the plugin.

Trimmed to what is invariant: command names + one-line purpose, the agent /
hook tables (model/color/tools, script/event), finding-ID format, enforced
.claude/rules conventions, coding style, test command, gotchas. Detail now
points to README / docs/. Also dropped two stale badge-duplicate figures that
had already rotted (the README badge owns them): "18 commands" (now 21) and
"1279 tests / 72 files" (now 1344) — removed rather than re-pinned so they
can't go stale again.

Measured: CLAUDE.md 134->102 lines, 8844->6170 B, 2,178->1,516 tok (-30%);
config-audit project always-delta 2,745->2,083 tok (-24%). Full suite 1344/0;
frozen v5.0.0 + SC-5 + default-output snapshots byte-stable (no scanner/snapshot
reads this repo's CLAUDE.md). Docs-only — no version bump, no catalog ref change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01683eAqVecv9VZfQzL8CQ9h
2026-06-29 08:32:44 +02:00
1bdaefc268 release: v5.12.5 — "Dogfood denoise" (M-BUG-2/6/7/8/10 scanner false-positive batch)
Version-sync for the Fase-3 scanner false-positive batch (code already shipped in
bfd577a / dd9db60 / 7e94910 / 3cf5c71 / e8afb14):
- plugin.json 5.12.4 -> 5.12.5
- README version badge -> 5.12.5, tests badge 1307 -> 1344, new version-history row
- CHANGELOG [5.12.5] section (per-bug Fixed entries)

Batch theme: five scanners stop counting non-user / non-live config as the user's authored
cascade (plugin-bundled config, frozen backups, doc examples, forward-compatible settings keys).

checkReadmeBadges: passed:true (tests 1344, scanners 16, commands 21, agents 7, hooks 4 — all
match filesystem). Full suite 1344/0. Frozen v5.0.0 + SC-5 + default-output snapshots byte-stable;
no re-seed across all five fixes (each affected fixture's findings are genuinely unchanged).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EnUvKEqyEa1m9gy6Aqhdqq
2026-06-26 18:04:20 +02:00
e8afb148d3 fix(acr): conflict-detector segregates plugin-bundled configs (M-BUG-2)
CNF compared every discovered settings.json/hooks.json pairwise regardless of
origin, so it treated installed plugins' bundled configs — each plugin's own
settings.json/hooks.json plus its shipped test fixtures and examples under
~/.claude/plugins/ — as if they were the user's authored cascade. A "conflict"
between two plugins' bundled test fixtures is not something a user can resolve,
yet these dominated the count: 339 CNF findings on this machine (315 high-sev
permission allow/deny "conflicts", 18 duplicate-hook, 6 settings-key), almost all
sourced from plugin-internal fixtures. The Conflicts grade was F on pure noise.
Fix: CNF excludes any file whose path is under `.claude/plugins/` from conflict
analysis (new isPluginBundled predicate; absPath marker). Kept CNF-local rather
than a discovery-level skip on purpose: an active plugin's contributed
hooks.json/.mcp.json legitimately lives in plugins/cache and other scanners need
it — only conflict analysis must ignore plugin-bundled files. Same class as
M-BUG-8 (non-live config trees treated as live). Suite 1344/0 (+3: plugin-bundled
exclusion, discovery-side sanity, over-exclusion guard). Frozen v5.0.0 + SC-5
snapshots untouched (marketplace-medium has no plugins/ paths), no re-seed.
Dogfood ~/.claude CNF 339->0 (F-grade was 100% plugin-bundled noise; the ~3
genuine user-scope local settings have no actual conflicting keys, matching the
plan C5 "real surface ~3 files" prediction).

Follow-up (not in this fix): classifyScope tags plugin-bundled files by checking
basePath instead of the file's own path, so scope:'plugin' is effectively dead
for a ~/.claude-rooted scan. Fixing it would let every scanner trust the scope
field, but that is a discovery-layer change beyond this bug's scope.
2026-06-26 17:24:08 +02:00
3cf5c714a2 fix(acr): SET typo-gates unknown-key false positives (M-BUG-10)
The CC settings schema is passthrough (verified against the 2.1.193 binary): it
forwards unrecognized keys unchanged rather than rejecting them, so an arbitrary
unknown key is valid/forward-compatible, not an error — the finding's "silently
ignored" claim was factually wrong. The only real risk is a TYPO of a real key
(the intended setting then silently has no effect). Fix: flag an unknown key only
when it closely matches a known key (new levenshtein helper; edit distance <= 2,
both keys >= 4 chars); severity medium -> low; honest passthrough framing in the
scanner + humanizer. Also refreshed KNOWN_KEYS with 6 binary-verified keys
(agentPushNotifEnabled, remoteControlAtStartup, skipAutoPermissionPrompt,
skipDangerousModePermissionPrompt, skipWorkflowUsageWarning, tui). Suite 1341/0
(+12). Frozen v5.0.0 snapshots untouched (0 CA-SET findings there), no re-seed.
Dogfood ~/.claude/settings.json 6->0 (all 6 keys above were false unknown-key
findings; 0 typo flags introduced across 167 walked files).
2026-06-26 15:15:14 +02:00
7e94910566 fix(acr): token estimator discounts block-level HTML comments (M-BUG-6)
CLAUDE.md token estimates counted block-level <!-- --> HTML comments toward
always-loaded tokens, but CC strips them before injection (preserved only inside
code fences, per code.claude.com/docs/en/memory). Fix: new stripInjectedHtmlComments
+ effectiveMemoryBytes in active-config-reader; the CML cascade (walkClaudeMdCascade)
and token-hotspots now size CLAUDE.md from effective (stripped) bytes, while raw byte
figures stay honest. Block-level only — inline comments retained (conservative,
verified scope). Suite 1329/0 (+13). Frozen v5.0.0 snapshots untouched (no fixture
has <!--), no re-seed. Dogfood ~/.claude CLAUDE.md ~3386->3301 tok (~85 tok discount,
matches worklist prediction).
2026-06-26 14:29:24 +02:00
dd9db60fc9 fix(acr): CPS ignores fenced/inline code + CC-stable path vars (M-BUG-7)
CPS flagged ${CLAUDE_PLUGIN_ROOT}/${CLAUDE_PROJECT_DIR} (CC-provided stable
paths) and {date}/timestamp tokens shown in documentation as cache-busters.
Fix: skip fenced code blocks, strip inline-code spans, and whitelist CC-stable
vars before pattern-matching. Suppress-only — frozen v5.0.0 snapshots untouched
(CPS yields findings:[] there), no re-seed. Suite 1316/0 (+6). Dogfood ~/.claude
5->2 (3 doc false-positives suppressed; 2 remaining = own volatile test fixtures).
2026-06-26 12:45:14 +02:00
bfd577aeee fix(acr): file-discovery skips backups/ dirs — never live config (M-BUG-8)
A directory named `backups` holds backup COPIES, not live config, so walking
it during a config audit produces stale findings. config-audit's own session
backups (~/.claude/config-audit/backups/<ts>/files/.../CLAUDE.md) were the
canonical case: a ~/.claude-scope audit walked 36 frozen config copies as if
live, polluting CPS (C3) and HKV/RUL (C6) results. Same non-live-noise family
as M-BUG-2.

- Add `backups` to SKIP_DIRS (broad, name-based — consistent with vendor/dist/
  .cache): the rule is general, a backups/ dir is never live config.

TDD: 2 failing tests (no config under backups/ discovered + backups/ counted
as skipped) -> green; a third asserts live config beside the backups tree is
still discovered. Full suite 1310/0 (+3). Byte-stable: no fixture is named
`backups` and `backups` appears in zero frozen snapshots, so v5.0.0 + SC-5 +
default-output outputs are unchanged. Dogfooding ~/.claude: files under
/backups/ drop from 36 -> 0; 717 live config files retained.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EnUvKEqyEa1m9gy6Aqhdqq
2026-06-26 12:23:43 +02:00
346dfac6fa release: v5.12.4 — "Rooted rules" (M-BUG-9: RUL resolves rule glob against the rule's own project root)
Version-sync for the M-BUG-9 fix (code already in 18af5a2):
- plugin.json 5.12.3 -> 5.12.4
- README version badge -> 5.12.4, tests badge 1305 -> 1307, new version-history row
- CHANGELOG [5.12.4] section

checkReadmeBadges: passed:true (tests 1307, scanners 16, commands 21, agents 7, hooks 4 — all
match filesystem). Full suite 1307/0. Frozen v5.0.0 + default-output snapshots byte-stable
(the fix is a no-op when projectRoot === targetPath; RUL appears in no snapshot).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EnUvKEqyEa1m9gy6Aqhdqq
2026-06-26 11:04:30 +02:00
18af5a24e9 fix(acr): RUL resolves rule glob against the rule's own project root, not the scan root (M-BUG-9)
A rule's paths:/globs: pattern scopes relative to the directory containing
the rule's .claude/, not the outer scan target. countGlobMatches globbed
against the scan root and collectProjectFiles' depth>4 cutoff never reached
deep matching files, so a live rule in a nested repo (e.g. a marketplace
checkout under ~/.claude) was wrongly flagged "matches no files / never
activates" (high) — a false F-grade for any user with rules in a nested repo.
Same scope-conflation family as M-BUG-1/2/8.

- deriveProjectRoot(ruleAbsPath): parent of the rule's .claude segment.
- collect + glob per project root (cached), relative to that root — so a
  nested repo's rule resolves against its own tree, where its files live.
- user-global rules (root === HOME) skip the no-match check: they scope
  against whatever project is active at runtime, not a fixed tree, so
  "matches 0 files here" is not a dead-rule signal (and avoids a HOME walk).

TDD: 2 failing tests (nested-repo false-positive + HOME guard) -> green.
Full suite 1307/0; frozen v5.0.0 + default-output snapshots unchanged (RUL
appears in none; the fix is a no-op when projectRoot === targetPath, i.e. the
common single-repo scan). Dogfooding C6: clears the 2 ktg-privat false
positives on the real machine and surfaces a previously-hidden genuine dead
rule (false negative) in the bundled optimal-setup example.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EnUvKEqyEa1m9gy6Aqhdqq
2026-06-26 10:56:39 +02:00
4ad1875b31 release: v5.12.3 — "Phantom agents" (M-BUG-3/4/5: enumerateAgents counts only CC-registered agents)
Releases commit 7f097d5. enumerateAgents now counts only the agents Claude Code
actually registers: recurse into agent subdirs (M-BUG-3), dedupe project==user
path when the scope root is $HOME (M-BUG-4, root cause — also fixed rules and
output-styles), and require valid name+description frontmatter before counting a
file (M-BUG-5). Real-machine verify: user-agent count 13->0 (all 12 user agents
plus REMEMBER.md are frontmatter-less, so CC registers none), HOME project-dup
13->0; corrected always-loaded baseline is ~53, not 66. No count change (scanners
16, agents 7, commands 21); agent enumeration is machine-dependent and absent
from the frozen snapshots, so v5.0.0 + SC-5 + default-output stay byte-stable.
1305 tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EnUvKEqyEa1m9gy6Aqhdqq
2026-06-26 00:37:40 +02:00
7f097d524f fix(acr): enumerateAgents counts only CC-registered agents — recurse + frontmatter filter + HOME dedup (M-BUG-3/4/5)
enumerateAgents previously counted every .md in an agents dir as an
always-loaded agent. Per the official CC subagents docs, CC registers a
subagent only when its frontmatter declares both name and description, scans
agents dirs recursively, and (at a HOME self-scan) must not count
~/.claude/agents twice.

- M-BUG-5: require valid name+description frontmatter; frontmatter-less files
  are registration no-ops costing 0 always-loaded tokens. Fixes the user-agent
  over-count (this machine: 13 -> 0).
- M-BUG-3: listMarkdownFiles gains opt-in recursion; enumerateAgents recurses
  so agents in subfolders (agents/review/x.md) are counted, matching CC.
- M-BUG-4: configDirs dedupes project==user paths, so a `manifest --global`
  self-scan (repoPath===$HOME) counts ~/.claude once (user scope), killing the
  spurious "project 13" double-count. Benefits rules/agents/output-styles.

TDD: 4 failing tests -> green. Full suite 1305/0; --json/--raw byte-stable,
frozen v5.0.0 + SC-5 + default-output snapshots untouched (no snapshot records
agent enumeration rows).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EnUvKEqyEa1m9gy6Aqhdqq
2026-06-26 00:19:27 +02:00
a1e786ba4f release: v5.12.2 — "Honest census" (M-BUG-1: honest plugin enumeration)
enumeratePlugins now honors enabledPlugins (disabled plugins no longer
contribute phantom agents/skills/commands) and enumerates polyrepo plugins
from their active installPath in installed_plugins.json, not only
marketplaces/<mkt>/plugins/. Fixes manifest/whats-active/AGT/token-hotspots
for any user with disabled plugins or a polyrepo marketplace. No count change
(scanners 16, agents 7, commands 21); --json/--raw byte-stable, frozen v5.0.0
+ SC-5 + default-output snapshots untouched. 1301 tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XYFipiVaRtbimkDjDnnKvY
2026-06-24 14:59:43 +02:00
be1056aac0 fix(acr): enumeratePlugins honors enabledPlugins + polyrepo cache installPaths (M-BUG-1)
active-config-reader walked ~/.claude/plugins/marketplaces and ignored both the
enabledPlugins toggle and the polyrepo cache layout. On a polyrepo machine it
counted disabled/uninstalled marketplaces plugins as "active" while MISSING the
actually-enabled plugins installed under plugins/cache. This corrupted the agent
listing and every pluginList consumer (manifest, AGT, whats-active, hooks, rules).

Now: when installed_plugins.json is present, inject only plugins that are in the
manifest AND enabledPlugins[key]===true, each resolved to its active installPath
(incl. cache/). When the manifest is absent (fixtures/pre-v2 installs), fall back
to the historic marketplaces walk rather than silently dropping config — mirrors
file-discovery.mjs's "trust installed_plugins.json" contract.

Verified on real machine: agent listing 114->104, ghost plugins (newsletter,
content-machine, harness, kiur, ...) gone, voyage/linkedin/ms-ai/okr now correctly
counted. Full suite 1301/0; byte-stable snapshots untouched (hermetic empty HOME).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CrTb8ktf1XZWEVwgz5MTTo
2026-06-24 10:50:28 +02:00
0f9e319c85 release: v5.12.1 — "Footgun guard" (Pattern H live-session caveat)
Version-sync for the Pattern H recommendation fix shipped in 45efed3:
- plugin.json 5.12.0 -> 5.12.1
- README version badge + version-history row (1297 tests)
- CHANGELOG [5.12.1] entry

Recommendation string only — no new finding ID or scanner (count stays 16,
agents 7, commands 21), no token figures changed, so --json/--raw stay
byte-stable and frozen v5.0.0 + SC-5 + default-output snapshots are untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 10:15:54 +02:00
45efed3dbf fix(tok): live-session caveat in Pattern H stale-cache recommendation
Pattern H ("Stale plugin-cache versions") recommended deleting stale
version dirs under ~/.claude/plugins/cache with no warning that a
currently-running session may still hold one of those versions for its
whole lifetime. "Stale" is judged against installed_plugins.json (what
NEW sessions load), so the recommendation could reproduce the exact
footgun that broke a live session during C4: deleting the dir pulls the
files out from under the running session, which then breaks and must
/exit + restart.

Extend the recommendation text with the live-session caveat. No new
finding ID/scanner, no token counts changed (recommendation string only)
-> --json/--raw stay byte-stable, frozen v5.0.0 + SC-5 + default-output
snapshots untouched. +1 test (1297).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 10:04:47 +02:00
6bb08cc84d release: v5.12.0 — "Auto-calibration" (B8b)
Version-sync for the B8b model→window auto-probe shipped in cf75249:
- plugin.json 5.11.0 -> 5.12.0
- README version badge + version-history row (1296 tests)
- CHANGELOG [5.12.0] entry

Scanner count stays 16, agents 7, commands 21. No new finding ID or scanner.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 22:29:21 +02:00
cf75249b5e feat(skl,cml): --context-window auto model→window probe (v5.12 B8b) [skip-docs]
Completes the deferred B8 half. `--context-window auto` now probes the
configured model and calibrates SKL/CML budgets to its real window instead
of always falling back to the conservative advisory anchor.

- lib/context-window.mjs: pure modelToContextWindow() maps known 1M-tier
  model IDs (Fable 5, Opus 4.8/4.7/4.6, Sonnet 4.6 — verified June 2026 —
  plus the explicit [1m] tier tag, dated/provider-prefixed IDs, and the
  opus/sonnet/fable aliases) to the 1M window; unknown/unconfirmed -> null
  (caller keeps the conservative anchor). resolveContextWindow() auto branch
  now probes opts.model: recognized -> auto-probed (not advisory); unknown
  or unpinned -> auto-unresolved (advisory, pre-B8b behavior).
- lib/active-model.mjs (new): resolveActiveModel() reads the model the way
  Claude Code resolves it — shell ANTHROPIC_MODEL override, then settings
  cascade local > project > user. Injectable env, hermetic under test HOME.
- scan-orchestrator: resolves the active model only when the flag is `auto`
  and threads it into resolveContextWindow; posture inherits via runAllScanners.

Default (no flag) and explicit --context-window <n> paths ignore the model
and stay byte-stable; frozen v5.0.0 + SC-5 snapshots untouched. TDD: 17 new
tests (context-window mapping/probe + active-model cascade). Suite 1296 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 22:29:12 +02:00
ad1eceb76a release: v5.11.0 — "Precision polish" (B7+B8)
Version-sync: plugin.json 5.10.0→5.11.0; README version badge + tests badge
1257+→1279+ + new version-history row; CLAUDE.md test counts 1257/71→1279/72
(22→23 lib test files); CHANGELOG [5.11.0]; README SKL row documents CA-SKL-003
+ the --context-window flag.

B7 (CA-SKL-003) and B8 (--context-window calibration) shipped as feat commits
2798880 + 2082b7d. Scanner/agent/command counts unchanged (16/7/21); --json/--raw
byte-stable; frozen v5.0.0 + SC-5 snapshots untouched. Suite 1279 green.
self-audit --check-readme: PASS (badges == filesystem).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 21:49:32 +02:00
2082b7d112 feat(skl,cml): --context-window calibration, advisory when unknown (v5.11 B8) [skip-docs]
SKL-002 (skill-listing budget) and CML char-budget now calibrate to a
resolved context window instead of always anchoring at 200k:

- resolveContextWindow(): --context-window <n> calibrates; 'auto' keeps the
  conservative 200k anchor but marks advisory (model→window probing deferred
  to B8b); no flag → 200k anchor, byte-identical to pre-B8 default.
- scaleForWindow(): linear off the 200k anchor (identity at the anchor).
- SKL + CML each keep an untouched default branch (window===200k && !advisory)
  for byte-stability and a calibrated branch; advisory downgrades the budget
  finding from a breach (low/medium) to info.
- Flag wired through scan-orchestrator + posture; runAllScanners resolves once
  and threads { contextWindow } to scanners (others ignore the 3rd arg).
- CPS intentionally excluded: it has no window-anchored budget (fixed
  150-line volatility heuristic), so there is nothing to calibrate.

15 new tests; e2e CLI verified (1M suppresses SKL-002, auto → info, default
unchanged); full suite 1279 green; snapshots byte-stable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 21:44:52 +02:00
27988801be feat(skl): flag oversized skill bodies on demand (v5.11 B7) [skip-docs]
New CA-SKL-003 (low): a SKILL.md body over ~5,000 tokens (~500 lines)
should split reference content into supporting files / use context: fork.

- measureActiveSkillListing() now returns body metrics (chars/lines/tokens);
  the body was already read in full, only the frontmatter was parsed before.
- Honest framing: BODY_CALIBRATION_NOTE marks this as ON-DEMAND cost (loads
  only when the skill is invoked, NOT every turn like the always-loaded
  listing) and an estimate — hence low severity.
- 5 new tests; full suite 1262 green; snapshots byte-stable (default branch
  untouched; new finding fires only on bodies >5k tok, none in fixtures).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 21:36:20 +02:00
c2e3a56a20 release: v5.10.0 — "Deferral & injection hygiene" (B4+B5+B6)
Three additive hardening levers extending existing scanners toward a tighter
always-loaded prefix (scanner count stays 16, agents 7, commands 21):
  - B4 (CA-TOK-006): MCP tool-schema deferral check + CLI-over-MCP lever
  - B5: hook additionalContext-injection advisory + filter-before lever
  - B6: CPS @import volatile-content scan

Version sync: plugin.json 5.10.0, README version/tests badges (1215+ -> 1257+)
+ version-history row, tokens pattern-count 7 -> 8 (README + CLAUDE.md),
CLAUDE.md test tally 1215/68 -> 1257/71 (36 -> 39 scanner files), CHANGELOG
[5.10.0]. self-audit --check-readme green; suite 1257 pass / 0 fail.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 21:02:14 +02:00
fa1ddd963a feat(cps): scan @imported files for volatile cached-prefix content (v5.10 B6) [skip-docs]
CPS originally inspected only files discovery classifies as claude-md, but a
CLAUDE.md can pull arbitrary files into the cached prefix via @import — and
those targets (e.g. @shared/conventions.md) are usually not claude-md in
discovery, so their inlined content was never scanned. Neither TOK Pattern A
(top-30 of cascade files) nor the in-file CPS scan reaches past the importing
file, so volatility inside an imported file was invisible.

B6 closes the gap: for each @import whose import site sits within the
cached-prefix window (imp.line <= CACHED_PREFIX_LINES), CPS resolves the path
(resolveImportPath, mirroring import-resolver/token-hotspots semantics), reads
the target, and runs findVolatileLines over its first 150 lines. A hit emits a
distinct medium finding — "Volatile content in @imported file breaks cached
prefix" — keyed on the resolved file, evidence naming the importer.

Scope boundaries (deliberate):
- One hop only; imports-of-imports stay with IMP (deep-chain owner).
- No lines-1-30 skip for imported content — that exclusion is root-file-specific
  to avoid Pattern A overlap, which never reaches imported files.
- No double-reporting: an import resolving to a discovered claude-md is skipped
  (own iteration); a reportedImports set dedupes a target imported by several
  CLAUDE.md files.

Dropped from B6 (per plan verdict): confident behavioral cache-buster detection
(opusplan/model-switch is runtime, not static config) and jq-transcript
automation. "No overstated behavioral finding ships" — even the permitted
opusplan info-advisory was left out; the @import extension is the whole of B6.

Byte-stability: the in-file finding keeps the same condition + byte-identical
evidence/description (continue-skip refactored to if-emit, behaviour-preserving);
new findings fire only on a volatile import, which no frozen v5.0.0 fixture has.
docs: README + scanner-internals CPS rows + full B6 note; CLAUDE.md kept lean
([skip-docs]). Suite 1254 -> 1257 green; snapshots + SC-5 untouched.
Version/badges/CHANGELOG wait for the v5.10 release cut.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 20:45:12 +02:00
d2c45a3bb8 feat(hooks): additionalContext injection advisory + filter-before lever (v5.10 B5) [skip-docs]
HKV now flags hooks that build hookSpecificOutput.additionalContext from
un-grepped command output as an INFO advisory (weight 0, never severity-bearing
— excluded from the self-audit nonInfo set). That field enters Claude's context
every time the hook fires (plain stdout on exit 0 does not), so an unfiltered
payload is a recurring per-turn token cost.

- New lib scanners/lib/hook-additional-context.mjs: pure assessHookAdditionalContext
  (unit-tested, no IO) + IO wrapper assessHookContextForRepo (walk hooks->scripts).
  Heuristic: additionalContext + verbose-prone capture (cat/git log/execSync/…) &&
  no filter (grep/head/jq/.slice). Deliberately low precision -> advisory only.
- HKV: emits the advisory inline on scripts it already reads (after the M5 verbose
  check). Additive, info severity -> frozen v5.0.0 + SC-5 snapshots untouched.
- feature-gap: new filterHookLeverFinding companion, fires ONLY when >=1 chatty
  hook is detected — surfaces the filter-before-Claude-reads lever (CC
  filter-test-output.sh). Silent otherwise (opportunity, not noise).
- docs: README HKV + GAP scanner rows; scanner-internals.md HKV row + full B5
  implementation note. CLAUDE.md kept lean ([skip-docs]); B5 fully documented in
  README + scanner-internals.

Mechanism verified 2026-06-23 against code.claude.com/docs context-window.md
(additionalContext enters context; plain stdout does not). Suite 1239 -> 1254 green.
Version/badges/CHANGELOG wait for the v5.10 release cut.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 20:28:33 +02:00
8f7e196046 feat(tokens): MCP tool-schema deferral check + CLI-over-MCP lever (v5.10 B4)
By default Claude Code defers MCP tool schemas (names-only, ~120 tok; full
schemas on demand via tool search). CA-TOK-006 detects config-file signals that
force the FULL schemas into the always-loaded prefix every turn:
  - settings.json env.ENABLE_TOOL_SEARCH="false"          (high)
  - "ToolSearch" in permissions.deny                       (high)
  - configured model is a Haiku model                      (medium)
  - per-server .mcp.json alwaysLoad:true (CC v2.1.121+)    (high)
auto[:N] is threshold mode (info, not a trigger).

New engine lib/mcp-deferral.mjs: pure assessMcpDeferral({settings,mcpServers})
(unit-tested, no IO) + thin IO wrapper assessMcpDeferralForRepo shared by TOK and
GAP. Severity scales with aggregate forced-upfront tokens (medium-confidence
reasons cap at medium). feature-gap cliOverMcpLeverFinding fires only as a
companion to CA-TOK-006 (prefer gh/aws/gcloud over MCP for common ops).

Honest scoping (Verifiseringsplikt): triggers on config files ONLY — never
process.env shell vars. Vertex / custom ANTHROPIC_BASE_URL / runtime /model
switch are launch state (would flap snapshots machine-dependently), so they are
DISCLOSED in every finding, not triggered. Tool-level anthropic/alwaysLoad and
claude.ai connectors likewise disclosed. Mechanism verified 2026-06-23 against
code.claude.com/docs (context-window.md, mcp.md#configure-tool-search +
#exempt-a-server-from-deferral, costs.md); the prefix-cache invalidation claim
was NOT-CONFIRMED in docs and is not asserted.

alwaysLoad added to CA-MCP VALID_SERVER_FIELDS (no longer flagged as unknown).
active-config-reader surfaces per-server alwaysLoad. Byte-stable: CA-TOK-006
fires only on new conditions; frozen v5.0.0 + SC-5 snapshots untouched.
Tests 1215 -> 1239 (engine 16, integration 5, lever 2, mcp-field guard 1).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 19:59:14 +02:00
42e48514e4 release: v5.9.0 — "Machine-wide token lens" (B1+B2+B3)
Cuts the v5.9.0 release bundling the three hardening gaps shipped since v5.8.0:
B1 (agent-listing budget — new orchestrated scanner AGT, count 15→16),
B2 (machine-wide always-loaded token roll-up in the campaign ledger), and
B3 (cache-aware filtering + stale plugin-cache disk-cleanup finding).

Version sync: plugin.json 5.8.0→5.9.0; README badges version 5.9.0 /
scanners 15→16 / tests 1168→1215; README intro + Health prose + scanner
table (new AGT row, TOK 7th pattern) + self-audit prose 15→16; new
version-history row; CHANGELOG [5.9.0]; CLAUDE.md finding-ID example (+CA-AGT)
and test count 1168→1215 / 67→68 files (scanner dir 35→36).

Gate: `self-audit --json --check-readme` -> readmeCheck.passed:true,
mismatches:[]. Full suite green (1215, 0 fail). Tag + catalog ref-bump follow.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 18:30:09 +02:00
ba9f82f952 feat(tokens): stale plugin-cache disk-cleanup finding, honestly labelled Dead config (v5.9 B3b)
The B3a cache filter already exposes discovery.staleCacheVersions; surface them
as a finding so the user knows the superseded plugin versions are safe to delete.

Honesty (Verifiseringsplikt): the finding loads on ZERO turns, so it must NOT
read as a per-turn token cost. TOK normally humanizes to "Wasted tokens"; a new
per-finding category override ('plugin-cache-hygiene' -> "Dead config") plus a
dedicated humanizer translation ("Old plugin versions are sitting on disk (safe
to delete) ... cost zero tokens per turn ... housekeeping, not a performance
problem") keep the prose accurate instead of the generic "using more space"
default. evidence carries the explicit "zero live-context impact" note.

- token-hotspots: Pattern H emits CA-TOK (low) from discovery.staleCacheVersions
  when stale versions exist (`--global`); silent otherwise.
- humanizer: CATEGORY_TO_IMPACT lets a finding's category override the
  scanner-default impact label (raw `category` field unchanged -> --json/--raw
  byte-stable). humanizer-data: honest static translation for the finding.
- Tests: finding fires/severity/category, lists stale keys + zero-impact note,
  silent when none; humanizer override -> Dead config; honest translation locked.
- Docs: tokens command render note (disk-cleanup, not a token problem) +
  --no-exclude-cache flag; README + CLAUDE.md rows (7 patterns, cache-aware).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 18:17:55 +02:00
a371832688 feat(discovery): version-aware cache filtering + --exclude-cache flag (v5.9 B3a) [skip-docs]
Stale ~/.claude/plugins/cache versions polluted token-hotspots ranking and
inflated CNF duplicate-hook findings with config that loads on zero turns.
installPaths point INTO the cache, so a blunt "skip all of plugins/cache"
would drop ACTIVE plugins — the filter is therefore version-aware: it reads
the adjacent installed_plugins.json, keeps active version dirs, drops only
stale ones (and exposes them via discovery.staleCacheVersions for B3b).

- file-discovery: cacheVersionKey() + applyCacheFilter() (active vs stale via
  installed_plugins.json; HOME-independent, derives the manifest from the cache
  path); discoverConfigFiles/Multi gain { excludeCache } + staleCacheVersions.
  Absent/unparseable manifest -> no filtering (never silently drop live config).
- token-hotspots-cli + scan-orchestrator: --exclude-cache (default ON for these
  live-cost scans) / --no-exclude-cache restores the full walk.
- Tests: cacheVersionKey unit cases; stale-dropped/active-kept/no-manifest
  discovery cases; CNF-drop proof (soft-spot verified, not assumed:
  include=1 -> exclude=0 duplicate-hook findings).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 17:57:22 +02:00
a17823a9af docs(campaign): render machine-wide token bill + refresh-tokens surface (v5.9 B2b-3)
Wires the user-facing surface of the machine-wide token roll-up (executable CLI
shipped in B2b-2, d664b70). The campaign report (Step 2) now renders rollUp.tokens:
the headline machine-wide always-loaded total split into the shared global layer
(paid once, every repo) + per-repo deltas, plus a ranked 'most expensive repos'
table. New mode 'refresh-tokens' (Step 6) documents the human-approved live sweep —
idempotent, skips unreadable repos, names any skipped so the bill's coverage stays
honest (Verifiseringsplikt).

Shapes documented carefully: rollUp.tokens.{sharedGlobal,perRepoDelta,machineWide}
are flat number maps; byRepo[] is the ranked {name,path,always,...} list. campaign-cli
already emitted rollUp.tokens (B2a) — no reader change; this is rendering only.

commands/campaign.md (73 lines) is the primary doc; README + CLAUDE.md campaign rows
get the matching one-line summary. Markdown-only → suite unchanged at 1198 green;
campaign command stays judgment-driven (not byte-stable).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 17:28:05 +02:00
d664b70520 feat(campaign): refresh-tokens live cross-repo token sweep (v5.9 B2b-2) [skip-docs]
The IO half of the machine-wide token roll-up (B2a shipped the pure data model).
New campaign-write-cli subcommand 'refresh-tokens': for every tracked repo it runs
readActiveConfig → buildManifest → splitManifestByOwnership, stores each repo's
per-repo delta via setRepoTokens, and captures the HOME-derived shared global layer
ONCE (from the first repo that reads cleanly) via setSharedGlobal.

Capturing shared once is the structural counted-once guard: a repo that fails to
read is recorded in 'skipped' and its delta omitted, never aborting the sweep.
Idempotent — a re-sweep replaces (setters overwrite), never accumulates. Determinism
unchanged: the clock is still read only via --reference-date.

TDD: 3 integration tests first (RED → GREEN) against a fixture with a dominant global
CLAUDE.md + two repos. Counted-once guard asserts each stored delta stays a small
fraction of the shared layer (a double-fold would make each delta EXCEED shared).
Caught a per-repo-MCP misclassification (fixed in 82f881a) + a test-side shape slip
(rollUp flattens its token buckets to numbers; stored summaries keep {tokens,count}).
Suite 1198 green (1195 + 3). Manifest + all frozen snapshots untouched.

[skip-docs]: internal sweep CLI (campaign plumbing), no user-facing surface yet. The
rendered token-bill in commands/campaign.md + CLAUDE.md rows land in B2b-3.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 17:23:15 +02:00
82f881afc4 fix(manifest): classify ~/.claude.json:projects MCP as per-repo delta, not shared [skip-docs]
readClaudeJsonProjectSlice(repoPath) returns the slice keyed to the SPECIFIC
repo path (exact / longest-prefix match), so MCP servers under
~/.claude.json:projects are per-repo: they load only in their own project and
differ across repos. B2b-1 wrongly grouped them with the shared global layer
(reasoning from file location, not the keyed slice). The live sweep captures
the shared layer ONCE from the first repo — folding a per-repo slice into it
would silently drop every other repo's claude.json MCP servers.

Corrected: SHARED_GLOBAL_SOURCES = {user, managed}; the only machine-global
MCP is plugin-provided (plugin: prefix). Per-repo MCP (.mcp.json AND the
~/.claude.json project slice) → delta. Tests updated.

[skip-docs]: internal classifier correction, no user-facing surface yet.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 17:14:28 +02:00
872b8ac281 feat(manifest): ownership split for machine-wide token roll-up (v5.9 B2b-1) [skip-docs]
Pure classifier splitManifestByOwnership(sources) → {shared, delta}, the
FS-free core that B2b's live cross-repo sweep will feed into the campaign
ledger's setSharedGlobal/setRepoTokens (B2a).

classifyOwnership maps each source string to its layer:
  shared : user | managed | plugin:* | ~/.claude.json:projects (global MCP)
  delta  : project | local | .mcp.json | @import | unrecognized

Anything not positively global falls to delta, so a source is never silently
folded into the once-counted shared layer (a wrong fold HIDES machine-wide
cost; a wrong delta is at worst visibly attributed to a repo). Both layers
carry the canonical summarizeByLoadPattern shape so the ledger setters consume
them verbatim.

TDD: 6 unit tests first (RED → GREEN). buildManifest/CLI output unchanged →
manifest snapshot byte-identical. Suite 1195 green (1189 + 6).

[skip-docs]: internal pure export only, no user-facing surface yet. User-facing
docs (commands/campaign.md + CLAUDE.md manifest/campaign rows) land in B2b-3
when refresh-tokens + the rendered token-bill ship.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 17:10:03 +02:00
cefa751990 feat(campaign): machine-wide always-loaded token roll-up — pure layer (v5.9 B2a)
Extend the campaign ledger with a machine-wide always-loaded token bill, the
pure data-model + aggregation half of B2. Deliberately does NOT do live I/O:
the cross-repo readActiveConfig sweep that populates real numbers is B2b.

Design — shared layer counted ONCE, structurally:
- Ledger root gets an optional `sharedGlobal` summary (the always-loaded layer
  paid in every repo: global CLAUDE.md + agent listing + global MCP + unscoped
  global rules + active plugins' always components), stored ONCE via the new
  setSharedGlobal().
- Each repo entry gets an optional `tokens` summary (its PER-REPO delta only),
  set via the new setRepoTokens(); addRepo now seeds `tokens: null` (mirrors
  findingsBySeverity).
- rollUp() stays PURE and gains a `tokens` aggregate: machineWide = sharedGlobal
  + Σ(per-repo deltas), so the shared layer is counted exactly once by
  construction — the structural guard against the historic double-count bug.
  Plus a `byRepo` table ranked DESC by always-loaded cost ("most expensive
  repos") with a deterministic name tie-break.

Summary shape mirrors manifest's summarizeByLoadPattern exactly
({always|onDemand|external|unknown: {tokens,count}}) so B2b wires in trivially.

Keeps rollUp pure (no filesystem I/O) → byte-stable, respects the THIN campaign
motor invariant. Chosen over the plan's literal "run readActiveConfig inside
rollUp" which would break both (operator-approved deviation).

- 8 new ledger tests incl. the counted-once regression guard + byRepo ranking.
- Updated addRepo shape test (+tokens:null) and campaign-cli EMPTY_ROLLUP.
- No frozen snapshot affected (campaign is v5.7, absent from v5.0.0 baselines).
  Suite green: 1189 tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 16:57:31 +02:00
2a3cb537f9 feat(scanner): add AGT per-agent description bloat advisory (v5.9 B1-rest)
Flag any single active agent whose description exceeds a 500-char soft cap
(the same bloat threshold TOK pattern F applies to SKILL.md descriptions).
Every char of an agent description re-enters context in the always-loaded
agent listing on every turn, so a long one is a per-turn cost.

Framing is deliberately ADVISORY, severity low: unlike CA-SKL-001 agents
have NO verified per-description cap, so nothing is dropped — this is a bloat
advisory, never a truncation claim. Per-agent advisories are emitted BEFORE
the aggregate roll-up (mirrors SKL 001->002).

- PER_AGENT_DESC_SOFT_CAP (=500) added to lib/agent-listing-budget.mjs as the
  single source of truth, disclosed as a heuristic.
- Per-agent loop reuses measureActiveAgentListing()'s agents[] (name, source,
  pluginName, path, descLength) — no new enumeration.
- 5 new tests (fire/strict-boundary/no-truncation-claim/recommendation/order).

Byte-stable: AGT already stripped from frozen v5.0.0 baselines; SC-5 unchanged
(HOME-scoped, hermetic HOME has no agents). Suite green: 1181 tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 16:27:24 +02:00
fcfb2979ef feat(scanner): add AGT agent-listing always-loaded budget finding (v5.9 B1)
New AGT scanner flags the aggregate always-loaded cost of the agent listing
(every active agent's name+description is injected each turn so the model
knows what it can delegate to). Mirrors the SKL skill-listing pattern but
encodes the key honesty caveat: the agent-listing mechanism is INFERRED
(agents are absent from Claude Code's documented context breakdown), so the
token figure is an UPPER-BOUND estimate and the aggregate budget is a
config-audit heuristic anchored on 200k — not a CC-documented allotment.

- scanners/agent-listing-scanner.mjs + scanners/lib/agent-listing-budget.mjs
- 8 TDD tests (tests/scanners/agent-listing-scanner.test.mjs)
- AGT folds into the Token Efficiency health area (scoring.mjs)
- byte-stability: AGT added to strip-added-scanner (frozen v5.0.0 baselines
  left untouched, same as OST/OPT); SC-5 default-output snapshot refreshed
- suite 1168 -> 1176 green

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 15:02:59 +02:00
759daa7201 chore(release): v5.8.0 — campaign motor (Fase 2)
Bump plugin.json 5.7.0 -> 5.8.0, README version badge + version-history row,
and CHANGELOG [5.8.0]. Covers the full Fase 2 campaign motor (durable ledger,
read-only campaign-cli, human-approved campaign-write-cli + /config-audit
campaign, cross-repo backlog, per-repo plan export with execution-by-reuse)
plus the pre-release cleanup (knowledge-refresh wiring, CLAUDE.md trim).

Scanner count 15, agents 7, commands 20 -> 21, tests 1091 -> 1168. All counts
verified green via self-audit --check-readme.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 12:03:36 +02:00
9c8acec71f docs(claude-md): trim to lean invariants; move impl notes to scanner-internals
CLAUDE.md was 540 lines (CML >500 MEDIUM, config-grade B/89). Move the 19
per-scanner/per-block implementation notes (v5.6/v5.7 design rationale +
primary-source verification + byte-stability lessons) verbatim into
docs/scanner-internals.md under a new "Implementation notes" section, leaving a
read-on-demand pointer. Add a real "## Conventions" section pointing at
.claude/rules/ — this clears CA-CML-001 (missing recommended section), which the
moved detail headings had been masking via incidental rule/style keyword matches.

CLAUDE.md 540->134 lines / 8.5k chars; config-grade A/97; CML findings: none;
readmeCheck green; full suite 1168 green (hermetic).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 10:31:17 +02:00
817fb8933f fix(commands): wire knowledge-refresh into router + help
knowledge-refresh shipped as a command file in v5.7 Fase 1 Chunk 3 but was
never added to the /config-audit router (argument-hint + routing list) or the
help command table, so `/config-audit knowledge-refresh` did not route. Add it
in both. (optimize was already fully wired — the STATE note was stale.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 10:31:05 +02:00
319e5541c9 feat(campaign): plan export + execution-by-reuse (v5.7 Fase 2 Block 4c)
Completes Block 4 (4b backlog + 4c export/execution). Asymmetric: plan
export is new testable code; execution is pure reuse of the existing
per-repo implement/rollback (no new execution machinery), per the plan's
"reuse existing backup/rollback".

Plan export ("planer følger arbeidsstedet"):
- scanners/lib/campaign-export.mjs (pure, now injected, 8 tests):
  planExportPath(repo,sessionId) -> <repo>/docs/config-audit-plan-<sessionId>.md
  (sessionId-keyed so same-day re-audits never collide);
  buildPlanExportDocument({...,now}) -> provenance header + verbatim plan.
- scanners/campaign-export-cli.mjs (-cli, read-only by default, 10 tests):
  --repo resolves the repo's linked session, reads its action-plan.md,
  assembles the doc, emits {exportable,problems,targetPath,document}. Two
  gates -> exit 1 advisory: no-session-linked / no-action-plan. Writes the
  file ONLY under opt-in --write (byte-faithful copy; the LLM never re-types
  a 200-line plan). --sessions-dir override for hermetic tests; exit 0/1/3.

Command: commands/campaign.md gains an `export <path>` mode (Step 6:
preview -> approve -> --write), then routes the user to the existing
/config-audit implement (backup + verify) + rollback + set-status
implemented. Nothing auto-written (Verifiseringsplikt).

Byte-stable: lib + -cli + command-doc only -> scanner count stays 15,
agents 7, commands 21 (export is a mode, not a new command), SC-5 +
backcompat suite untouched. suite 1150->1168. Block 4a (migrateLedger)
still deferred to the first breaking schema change.

Docs: CLAUDE.md section + badge 1150->1168/65->67 files; README badge +
campaign row + Testing prose (fixed stale 1055/59 -> true 1168/67).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 10:08:04 +02:00
49833aded8 feat(campaign): cross-repo prioritized backlog (v5.7 Fase 2 Block 4b)
buildBacklog(ledger) pure transform + read-only campaign-cli payload field
+ command rendering. The single machine-wide pick-list: per-repo (the ledger
tracks severity counts, not individual findings), severity-weighted
(SEVERITY_WEIGHTS c1000/h100/m10/l1), deterministic tie-break, excludes
implemented/pending/zero-finding repos.

No schema change, no new scanner -> scanner count stays 15, snapshot/backcompat
byte-stable. suite 1138->1150 (lib +9, campaign-cli +3). README badge 1091+->1150+.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 02:44:21 +02:00
ee0c762151 feat(campaign): /config-audit campaign — human-approved write-CLI + command (v5.7 Fase 2 Block 3c)
Completes the THIN machine-wide campaign surface (ledger + roll-up + status).
The write half of the campaign motor, mirroring how knowledge-refresh gates
register writes:

- scanners/campaign-write-cli.mjs (-cli → NOT a scanner; 11 tests): init/add/
  set-status, each a thin wrapper over the invariant-enforcing lib transforms
  (createLedger/addRepo/setRepoStatus) + saveLedger — path-normalization/dedup,
  idempotent add, status-lifecycle guard and updatedDate bump never hand-rolled.
  init refuses to clobber an existing/corrupt ledger (exit 1, file untouched);
  add auto-inits + reports added vs skipped; set-status takes --findings/--session.
  Deterministic: --reference-date is the only clock read, injected as `now`.
  Exit 0=write, 1=advisory no-op, 3=error.
- commands/campaign.md (opus, no Web): thin orchestrator — always reports first
  (read-only campaign-cli), then proposes init/add/set-status and invokes ONE
  write-CLI subcommand only on explicit human approval (Verifiseringsplikt; never
  hand-edits the JSON). add --discover finds git repos under a root to pick from.

Not byte-stable (own command, outside snapshot suite) like /config-audit optimize
+ knowledge-refresh. Both CLIs are -cli → scanner count stays 15, snapshot suite
untouched. commands 20→21, suite 1127→1138 (+11), test files 64→65.

Docs: CLAUDE.md §campaign-write-cli + command table + Testing badge; README
commands badge 20→21 + table row; help.md + router wiring.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 17:25:15 +02:00
4be7a16788 feat(campaign): read-only campaign-cli ledger reporter (v5.7 Fase 2 Block 3b)
Read-only reporter over the durable campaign-ledger core (Block 3a). Mirrors
knowledge-refresh-cli: the deterministic, READ-ONLY half of the campaign motor.

- scanners/campaign-cli.mjs (-cli → not an orchestrated scanner): loadLedger +
  validateLedger + rollUp, emits {status, initialized, ledgerPath, schemaVersion,
  createdDate, updatedDate, repos, rollUp} JSON. NEVER writes — a missing ledger
  is reported gracefully (initialized:false), never created. init + status
  transitions belong to the Block 3c command layer (human-approved writes).
- --ledger-file overrides default path (deterministic testing); --output-file
  mirrors the sibling. Exit: 0 = initialized & valid, 1 = not initialized
  (advisory), 3 = error (parse/corrupt/invalid).
- 8 tests (tests/scanners/campaign-cli.test.mjs). Suite 1119→1127 green hermetic;
  scanner count stays 15 (-cli), SC-5 byte-stable.
- CLAUDE.md: §campaign-cli added; Testing badge corrected 1091/62 → 1127/64
  (fixes pre-existing Block 3a drift).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 14:13:27 +02:00
f93830ce74 feat(campaign): durable machine-wide campaign-ledger core (v5.7 Fase 2 Block 3a THIN)
Add the durable ledger that sits ABOVE individual config-audit sessions for a
machine-wide audit campaign: repo list + per-repo lifecycle (pending → audited →
planned → implemented) + a machine-wide roll-up by status and severity.

scanners/lib/campaign-ledger.mjs — same hybrid split as knowledge-refresh:
- PURE transforms (createLedger / addRepo / setRepoStatus / rollUp) with `now`
  injected (YYYY-MM-DD, never the clock) → deterministic + unit-testable.
- soft validateLedger (returns {valid,errors}, never throws) for loaded data;
  transforms throw on programmer error (invalid status, unknown path).
- thin IO shell (defaultLedgerPath / loadLedger→null-on-ENOENT / saveLedger).
- persists to ~/.claude/config-audit/campaign-ledger.json — OUTSIDE the plugin
  dir (next to sessions/) so it survives uninstall/reinstall/upgrade.
- schemaVersion stamped from the start → cheap Block 4 migration.

THIN scope (Block 3a, operator-approved): ledger + roll-up + persistence only —
no CLI/command/execution (Blocks 3b/3c/4). Internal plumbing, byte-stable until
consumed: no `export async function scan` + lives in lib/ → scanner count stays
15, no orchestrator wiring, SC-5 unchanged.

tests/lib/campaign-ledger.test.mjs — 28 tests (TDD, red→green). Full hermetic
suite 1091→1119 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 13:30:44 +02:00
f9862b8a6d chore(release): config-audit v5.7.0 — optimization lens + living knowledge layer
Release-cut for v5.7 Fase 1 (Chunks 1/2a/2b/3):
- best-practices register (knowledge/best-practices.json)
- CA-OPT-001 deterministic optimization-lens scanner (count 14→15)
- /config-audit optimize + optimization-lens-agent (opus, agents 6→7)
- /config-audit knowledge-refresh (commands 19→20)

plugin.json 5.6.0→5.7.0, README version badge + version-history row,
CHANGELOG [5.7.0] entry. Count badges already track filesystem
(self-audit --check-readme: passed). Suite 1091 green (hermetic HOME),
self-audit A/A (config 94, plugin 100).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 20:04:08 +02:00
0f9c091a14 feat(knowledge): knowledge-refresh — living register, deterministic stale core + web poll (v5.7 Chunk 3)
The 'living' half of the v5.7 living knowledge base. Same hybrid split as the
optimization lens (Chunk 2b): a deterministic, byte-stable, unit-tested core +
a web/judgment command shell.

- scanners/lib/knowledge-refresh.mjs: pure assessFreshness(register,
  {referenceDate, staleAfterDays=90}) — age-based fresh/stale classification of
  source.verified; referenceDate injected (never reads the clock) → fully
  deterministic. 15 tests.
- scanners/knowledge-refresh-cli.mjs: -cli (NOT an orchestrated scanner →
  scanner count stays 15, suite byte-stable). Read-only — never writes the
  register, never hits the network. --reference-date/--stale-after/--dry-run,
  exit 0/1/3. 8 tests.
- commands/knowledge-refresh.md (opus): CLI stale-report → re-verify each stale
  entry by re-reading its source.url → poll CC changelog + Anthropic blog →
  apply ONLY human-approved writes, then re-validate the register. No unverified
  claim is ever auto-written (Verifiseringsplikt). Web-driven → not byte-stable.

No new agent, no new orchestrated scanner. Docs/badges: commands 19→20,
tests 1068→1091 (20 lib + 32 scanner test files). self-audit A/A, readmeCheck passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 19:22:57 +02:00
ba66f1fc17 chore(state): make STATE.md local-only — public open/ mirror must not carry state-of-play
This repo's only remote is the PUBLIC open/config-audit mirror. The global
continuity rule (refined) says a repo with a public remote must keep STATE.md
LOCAL-ONLY (gitignored), because the only available push would leak internal
state-of-play. STATE.md was tracked here under the old "always tracked" rule,
which silently assumed a private remote.

Untrack STATE.md (kept on disk for continuity; the session-start hook reads it
from disk regardless of git status) and add it to .gitignore. Mirrors the
linkedin-studio precedent (commit 9338454). Operator-approved 2026-06-21.

Note: prior chore(state) commits remain in public history (accepted — scrub not
requested).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 19:01:34 +02:00
7b3b487d26 feat(opt): optimization lens Chunk 2b — opus prose-judgment analyzer + /config-audit optimize
The recall+precision halves of the CA-OPT hybrid motor for the three
mechanism-fit cases the deterministic OPT scanner (2a) deliberately skips:
lifecycle→hook (BP-MECH-001), unscoped path-specific→rule (BP-MECH-002),
absolute "never"→permission (BP-MECH-004).

New:
- scanners/lib/lens-prefilter.mjs — cheap, recall-oriented line scan of the
  CLAUDE.md body; detector names mirror the register lensCheck fields; skips
  fenced code, gates the path class on an instruction verb. Pure + 13 tests.
- scanners/optimize-lens-cli.mjs — discovery + OPT scanner + pre-filter; attaches
  only the CONFIRMED register entry to each candidate (unverifiable → dropped,
  Verifiseringsplikt); emits {deterministic, candidates, register, counts}.
- agents/optimization-lens-agent.md — opus precision gate (7th agent, orange):
  reads the real CLAUDE.md, drops low-confidence candidates, keeps only genuine
  opportunities, cites register id + source.
- commands/optimize.md — /config-audit optimize orchestrates pre-filter→agent→report.

Agent-driven → deliberately NOT byte-stable (own command, outside the snapshot
suite). No new orchestrated scanner → scanner count stays 15. Counts: agents
6→7, commands 18→19, suite 1055→1068. Self-audit A/A unchanged, readmeCheck
passed (clean HOME). Plan: docs/v5.7-optimization-lens-plan.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 14:01:18 +02:00
c1409ae9b9 test(tok): hermetic HOME in runScanner — deterministic on populated ~/.claude
token-hotspots.test.mjs ran scan() under the developer's real HOME, so
readActiveConfig leaked ~/.claude (user CLAUDE.md cascade + ambient MCP/
plugins) into every fixture result. This made assertions machine-dependent:
green on a clean HOME, red on a populated one (small-cascade tipped past the
10k-token threshold; sonnet-era gained ambient MCP findings).

Wrap the HOME-dependent scan() call in the shared withHermeticHome helper —
same isolation the OST test and the snapshot/byte-stability suite already use.
Full suite (1055) now green on BOTH real and clean HOME.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 13:45:53 +02:00
313375184c chore(state): v5.7 Chunk 2a done; next session = TOK-fix → Chunk 2b (GO given)
Session boundary. Both next tasks pre-approved by operator (2026-06-21), to run in order in a fresh session: (1) TOK test HOME-isolation fix, (2) Chunk 2b opus analyzer.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 13:36:27 +02:00
e7833b65fc feat(opt): optimization lens CA-OPT-001 (procedure→skill) — v5.7 Fase 1 Chunk 2a
First detector of the 'is the config OPTIMAL?' axis (vs the existing 'correct?' scanners). New orchestrated scanner family CA-OPT (count 14->15), the deterministic half of the hybrid optimization lens.

CA-OPT-001 (low, Missed opportunity): a multi-step procedure in CLAUDE.md (>=6 consecutive numbered steps) that belongs in a skill. Reads recommendation + provenance from the best-practices register (BP-MECH-003). Conservative by design; the negative corpus proves null false-positives. Prose-judgment cases (lifecycle->hook, 'never'->permission) are deferred to the Chunk 2b opus analyzer.

Wiring mirrors OST: orchestrator entry, humanizer (OPT->'Missed opportunity' + family), scoring (OPT->'CLAUDE.md', existing area -> no new posture row -> byte-stable), strip-helper (OST,OPT), SC-5 regenerated under hermetic HOME (additive OPT entry only). 10 new tests; suite 1045->1055, self-audit A/A, readmeCheck passed (all verified with a clean HOME).

Note: the pre-existing TOK test reads the real ~/.claude (non-hermetic) -> run the suite with a clean HOME for deterministic results. Tracked as a separate follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 23:10:03 +02:00
55f83a3c99 feat(knowledge): best-practices register foundation — v5.7 Fase 1 Chunk 1
Add knowledge/best-practices.json: a machine-readable, provenance-stamped, schema-validated best-practices register — the source of truth for the upcoming v5.7 optimization lens (CA-OPT). 13 seed entries migrated from the v5.5 V-rows (loading-model + compaction facts) and the Anthropic 'Steering Claude Code' blog (mechanism-fit rules); each entry carries source.url + verified date + confidence. Only confirmed claims are user-facing (Verifiseringsplikt).

scanners/lib/best-practices-register.mjs: zero-dependency loader + validator (loadRegister/validateRegister/getEntry, native JSON.parse — not YAML, since the repo is zero-dep and yaml-parser.mjs can't parse arrays-of-objects). tests/lib/best-practices-register.test.mjs: 22 tests (schema, provenance integrity, negative cases, lookup).

Byte-stable: no scanner consumes the register yet (Chunk 2), so all scanner output is unchanged. Suite 1023->1045, self-audit A/A, readmeCheck passed. Full design: docs/v5.7-optimization-lens-plan.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 22:37:47 +02:00
685b770cb9 docs(state): v5.7 Fase 1 designet — visjons-diskusjon
Visjons-diskusjon gjennomført (F1-tuning: korrekt? → optimal?). Visjonen dekomponert i 4 byggeklosser, sekvensert i 2 faser. Fase 1 (optimerings-linse + levende kunnskapsbase) designet og operatør-låst; Fase 2 (maskinvid kampanje + varig backlog) utsatt til eget GO.

Låste beslutninger: strukturert register (linsen leser; markdown som speil); ny familie CA-OPT med hybrid motor (determ. pre-filter → opus-analyzer), presisjons-gated. GO-ready plan i docs/v5.7-optimization-lens-plan.md; STATE oppdatert. Ingen produksjonskode (bevisst diskusjons-sesjon).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 22:02:01 +02:00
0a5a347ea1 chore(state): v5.6.0 released; capture operator vision (F1-tuning) for next-session discussion
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 21:36:08 +02:00
51ca45500c chore(state): v5.6.0 released (tag + catalog pushed); v5.7 next
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 21:15:28 +02:00
7548a627ba release: v5.6.0 — steering-model II (Foundation + B + C)
Bundles the v5.6 work already on main into a release:
- Foundation (62d910e): active-config-reader enumerates rules/agents/output
  styles with a loadPattern/survivesCompaction/derivationConfidence triple;
  frontmatter parser reads YAML block sequences.
- B (bb647ce, 778b517): load-pattern accounting — manifest reports
  component-level sources with an always-loaded subtotal; token-hotspots tags
  each hotspot with its load pattern.
- C (e3b044a): new orchestrated OST scanner (CA-OST-001/002/003), scanner
  count 13 -> 14, all claims doc-verified.

Version synced: plugin.json 5.5.0->5.6.0, README version badge +
version-history row, CHANGELOG [5.6.0] section. Frozen v5.0.0 snapshots
preserved via strip-helpers (--json/--raw byte-stable for the original 13
scanners); only SC-5 default-output regenerated. Suite 1023 pass, self-audit
A/A (config 93, plugin 100), readmeCheck passed, mismatches [].

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 21:13:47 +02:00
e3b044a476 feat(ost): v5.6 C — output-style scanner (CA-OST, count 13→14)
New orchestrated scanner output-style-scanner.mjs — first new family since
SKL. Three findings, each pinned to a CONFIRMED V-row of the steering-model
plan + re-verified against code.claude.com/docs/en/output-styles:

- CA-OST-001 (medium, V10): user/project custom style missing
  keep-coding-instructions:true (default false) → silently strips built-in
  software-engineering instructions when active. Scoped to user/project.
- CA-OST-002 (low, V11): plugin style with force-for-plugin:true overrides the
  user's selected outputStyle. Verifiseringsplikt correction — the plan bullet
  said "project/user style," but force-for-plugin is plugin-styles-only per the
  docs, so the check keys on source==='plugin'.
- CA-OST-003 (medium): settings outputStyle matching no built-in
  (Default/Explanatory/Learning/Proactive, case-insensitive) nor discovered
  custom style → dead config.

Byte-stability — a scanner addition, not a field addition. Growing the
scanners array + scanners_ok cannot be hidden by a field strip, but re-seeding
frozen v5.0.0 (the SKL precedent) would now bake in B2's hotspot triple +
claudeMd drift. So, per the B2 lesson, frozen v5.0.0 snapshots are PRESERVED
and the OST entry is stripped at compare time via new
tests/helpers/strip-added-scanner.mjs (wired into json/raw-backcompat + the
Step 5/6 humanizer tests); only SC-5 default-output is regenerated (additive
OST entry, diff reviewed). OST is fixture-gated (no output styles on
marketplace-medium / hermetic HOME → silent).

Wiring: orchestrator; humanizer (OST→Configuration mistake) + humanizer-data
OST family (title-coupled); scoring (OST→Settings, keeps 10 areas). Suite
1012→1023 (+11). Badges: scanners 14, tests 1023, TRANSLATIONS families 15.
Lore swept: README, CLAUDE.md, scanner-internals, humanizer.md. self-audit A/A.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 21:02:44 +02:00
43d8873339 chore(state): v5.6 B complete on main; C GO-approved for next session
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 20:24:29 +02:00
778b517e6f feat(tok,acr): v5.6 B2 — load-pattern column in token-hotspots
Annotate every ranked TOK hotspot with the load-pattern triple
(loadPattern/survivesCompaction/derivationConfidence):

- hotspotLoadPattern() maps each discovery `type` → a deriveLoadPattern kind.
  Rules reuse activeConfig.rules for precise `scoped` handling; claude-md maps
  by scope. Two new deriveLoadPattern kinds back the rest: `command`
  (on-demand — body loads on /invoke) and `harness-config` (external —
  settings/keybindings/.mcp.json/hooks.json/plugin.json configure the CLI, not
  the model context, so they cost no per-turn context tokens). Honest split:
  the .mcp.json FILE is external; the MCP server's tool schemas are a separate
  `always` hotspot.

Byte-stability — the opposite of B1's manifest. token-hotspots IS a byte-equal
SC-6/SC-7 CLI, and its hotspots ride inside scan-orchestrator + posture, so the
change touched SIX frozen-v5.0.0 comparisons across five test files. Resolved by
preserving the frozen baselines: a shared tests/helpers/strip-hotspot-load-pattern.mjs
strips the additive triple before each byte-equal compare (proves the original
schema is byte-identical). SC-5 default-output snapshots (scan-orchestrator +
token-hotspots) regenerated — diff reviewed as additive-only.

Tests 1008→1012. Self-audit A/A, scanner count unchanged at 13 (C bumps to 14).
Completes v5.6 B (B1 manifest + B2 token-hotspots).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 20:21:01 +02:00
bb647ce35f feat(mft): v5.6 B1 — load-pattern accounting in manifest
Consume the v5.6 Foundation enumeration in buildManifest:

- Component-level sources: drop the coarse `kind:'plugin'` roll-up (it
  double-counted skills/rules/agents already enumerated once). Kinds are now
  claude-md/skill/rule/agent/output-style/mcp-server/hook.
- Every source carries loadPattern/survivesCompaction/derivationConfidence.
  Rules/agents/output-styles propagate the foundation-derived values; CLAUDE.md
  maps scope→kind (all cascade files always-loaded); skills are tagged on-demand
  (skill-body) so the body cost does not inflate the always-loaded subtotal.
- New `summary` (always/onDemand/external/unknown {tokens,count}); the
  always-loaded subtotal — "tokens that enter context every turn" — is the headline.

manifest is an environment-aware CLI → SC-6/SC-7 verify it by mode-equivalence,
not byte-equal, and it is not in SC-5. Adding fields in place keeps all snapshots
green with no regen (verified). `total` changes (de-duped) — intended correctness fix.
TOK's load-pattern column (byte-equal SC-6) is deferred to the next chunk (B2).

Tests 996→1008 (deterministic buildManifest unit test + CLI presence checks).
Self-audit A/A, scanner count unchanged at 13.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 17:48:55 +02:00
62d910ed6d feat(acr,yaml): v5.6 Foundation — load-pattern enumeration + block-seq parser
Foundation chunk of v5.6 "steering-model II" (internal plumbing for B/C;
no command-output change, so --json/--raw/SC-5/6/7 stay byte-stable, count
stays 13).

active-config-reader.mjs:
- deriveLoadPattern(kind,{scoped}) — pure helper mapping each source kind to
  loadPattern {always,on-demand,external} + survivesCompaction {yes,no,n/a}
  + derivationConfidence {confirmed,inferred}, traced to the published
  loading model (V-rows in docs/v5.5-steering-model-plan.md).
- enumerateRules / enumerateAgents / enumerateOutputStyles — the three
  source kinds previously unenumerated (mirror enumerateSkills). Output-style
  discovery is direct (not a new file-discovery type) to keep the discovery
  surface stable.
- readActiveConfig now exposes rules/agents/outputStyles arrays + totals
  counts/subtotals (folded into grandTotal).

yaml-parser.mjs:
- parseSimpleYaml now reads YAML block sequences (paths:\n  - a), not just
  inline paths:. An empty-valued key with no `- ` items stays null
  (backcompat). Resolves a pre-existing RUL false-positive (a block-seq-scoped
  rule was misread as unscoped) — fix flows through unchanged RUL code.

Tests +35 (961 -> 996): block-seq parser cases, RUL block-seq regression
(no-misflag + durability-fires), deriveLoadPattern table, three enumerators
(positive+negative). Amended two existing ACR asserts (top-level key shape +
grandTotal sum). self-audit A/A, readmeCheck passed, mismatches []. tests
badge 961+->996+; README testing prose de-staled (635/36 -> 996/56);
CLAUDE.md Foundation note.

B (manifest/tokens render + snapshot regen) and C (CA-OST, count->14)
deferred to their own sessions/GO.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 16:57:52 +02:00
d03c3831bd chore(state): v5.5.0 released (tag + catalog pushed); v5.6 next
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 11:43:35 +02:00
dac1db48c5 release: v5.5.0 — steering-model I (A+E)
Bundles the two v5.5.0 features already on main:
- A (f3aadb5): RUL/CML compaction-durability findings (both LOW) — a
  large (>50-line) path-scoped rule and a nested CLAUDE.md are not
  re-injected after compaction.
- E (f75ed56): PLH flags plugin-agent hooks/mcpServers/permissionMode
  frontmatter Claude Code ignores for plugin subagents (permissionMode
  MEDIUM false-security, rest LOW).

Additive — scanner count stays 13, --json/--raw byte-stable. Suite 961
pass. self-audit A/A (config 93, plugin 100), readmeCheck passed,
mismatches []. Version synced: plugin.json 5.4.1->5.5.0, README
version badge + version-history row, CHANGELOG [5.5.0] section.

Foundation (active-config-reader enumeration) + B deferred to v5.6.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 11:34:51 +02:00
9eba0f6169 chore(state): v5.5.0 A+E landed on main; Foundation deferred to v5.6
A (RUL+CML durability) and E (PLH plugin-agent dead-config) implemented
and committed. Foundation moved to v5.6 (A/E are additive and don't
consume it). STATE + plan updated; v5.5.0 release-cut is a separate GO.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 11:21:30 +02:00
f3aadb5183 feat(rul,cml): durability findings — config lost after compaction (v5.5.0 / A)
Per the official "what survives compaction" model (context-window.md):
only the project-root CLAUDE.md (+ unscoped rules) is re-injected after a
context compaction. Two additive, structural findings (low severity):

- RUL: a large (>50-line) PATH-SCOPED rule — reloads only on a matching
  file read, and is not re-injected after compaction, so must-hold rules
  can silently drop mid-session.
- CML: a NESTED (subdirectory) CLAUDE.md — not re-injected after
  compaction (only the project root is).

Additive (no new scanner, count stays 13); one humanizer entry each;
hermetic temp-fixture tests (positive + negative). Suite 957 -> 961,
SC-5 byte-stable, self-audit A/A.

Known limitation (pre-existing, broader than A): the lightweight
frontmatter parser reads inline `paths:` but not YAML block sequences,
so block-sequence-scoped rules are still seen as unscoped. Deferred.

Part of v5.5.0 "steering-model I". Foundation (active-config-reader
enumeration) deferred to v5.6 with B.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 11:20:04 +02:00
f75ed5655c feat(plh): flag plugin-agent frontmatter Claude Code ignores (v5.5.0 / E)
Plugin subagents silently ignore `hooks`/`mcpServers`/`permissionMode`
frontmatter — these are honored only for user/project agents in
.claude/agents/ (code.claude.com/docs sub-agents). Setting them in a
plugin agent is dead config; `permissionMode` is MEDIUM because it
implies a restriction Claude Code does not apply (false security).
hooks/mcpServers are LOW.

Additive to PLH's agent-frontmatter loop (no new scanner, count stays
13). One humanizer pattern covers the three field titles. Hermetic
temp-fixture test (positive + negative). Suite 954 -> 957, byte-stable.

Part of v5.5.0 "steering-model I". Foundation (active-config-reader
enumeration) deferred to v5.6 with B — A/E are additive and don't
consume it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 11:11:12 +02:00
73a7f117d6 chore(state): prime next session to start v5.5.0 (Foundation + A + E)
GO granted to implement v5.5.0 in a fresh session. STATE now opens
directly into the task with ordered, concrete first steps (Foundation
enumeration → A durability → E plugin-agent dead-config), the brief
pointer, and the impl-critical gotchas.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 10:58:16 +02:00
51fe6a1197 chore(state): v5.4.1 released; U1/U2 verified; v5.5 ready to start
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 10:18:14 +02:00
2f9d391b95 release: v5.4.1 — scanner-correctness patch (HKV/RUL/PLH)
Bundles four primary-source-verified scanner fixes since v5.4.0:
- HKV: +Setup/UserPromptExpansion/PostToolBatch; removed post-session
  (a self-hosted-runner lifecycle hook, not a settings.json event)
- RUL: globs-rule wording corrected (only paths: is documented)
- PLH: optional model/tools/name/allowed-tools no longer required;
  CLAUDE.md component-section required only for shipped components

Count stays 13, --json/--raw byte-stable, suite 954, self-audit A/A,
--check-readme passed. Version-history + CHANGELOG updated; badges bumped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 10:15:22 +02:00
5ac6c87053 fix(hkv): remove post-session — it is a runner hook, not a settings event
Verified 2026-06-20 against hooks.md + the 2.1.169 changelog: the
`post-session` hook in that changelog is a SELF-HOSTED-RUNNER
workspace-lifecycle hook (runs after the session, before the workspace
is deleted), NOT a settings.json hook event. It is absent from
hooks.md's 30-event list (all PascalCase), so config-audit was wrongly
treating a bogus `post-session` settings hook as valid — it now flags it.

Restructured the event test suite accordingly and added a negative test.
Resolved U1/U2 in the v5.5 plan (U1 refuted → removed; U2 `outputStyles`
plugin.json key confirmed → PLH unchanged). Suite stays 954, self-audit A/A.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 10:12:52 +02:00
55633028e5 chore(state): update STATE.md — v5.4.1 corrections landed + v5.5 plan
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 10:00:55 +02:00
e71e60f14e docs(plan): v5.5+ steering-model coverage plan (A–E)
Doc-grounded plan for the five new-functionality items greenlit after
checking the "seven steering mechanisms" framing against live CC docs
(2026-06-19): A durability/compaction findings, B load-pattern token
accounting, C output-style scanner, D mechanism-fit heuristic, E
agent-listing cost + plugin-agent dead-config.

Includes a full verification log (CONFIRMED/REFUTED/UNVERIFIED per claim
+ source), dependency graph, phased rollout (v5.5/v5.6/v5.7), testable
acceptance criteria, and open decisions. PLAN only — each feature awaits
its own GO. No code.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 09:58:56 +02:00
b6a62d7699 fix(hkv,rul): add 3 verified hook events; correct globs-rule wording
HKV: add Setup, UserPromptExpansion, PostToolBatch to VALID_EVENTS,
verified live against code.claude.com/docs/en/hooks.md (2026-06-19). A
valid hook using one of these was wrongly flagged "will never fire" — a
user could delete a working hook. Made the "(N total)" hint dynamic so
it can't drift again. Flagged the unverified kebab 'post-session' in a
comment (an existing test depends on it; follow-up check needed).

RUL: reword the globs finding. Only `paths:` is documented; whether CC
ever read `globs` is unverified, so the old "deprecated/legacy" framing
overclaimed (Verifiseringsplikt). New wording steers to the documented
`paths:` field. Updated the coupled fix-engine title match and the
humanizer entry (which also carried the "field was renamed" overclaim).

Suite 950 -> 954 (badge bumped). self-audit A/A, scanner count 13. No
version bump — these land in the pending v5.4.1 patch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 09:56:00 +02:00
5a270704fd chore(state): update STATE.md — marketplace review + v5.4.1 pending
Captures this session: STATE.md tracking rule, catalog version-gate, 2 scanner over-report
fixes (missing-model, section-per-component → suite 950), marketplace full-depth review
(okr pilot + 5 stable; 4 pushed). Open: llm-security CRITICAL RCE fix-brief (push parked),
v5.4.1 patch, 3 active plugins, STATE.md sweep.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 09:29:31 +02:00
292352eff8 fix(plh): require CLAUDE.md commands/agents/hooks section only for shipped components
plugin-health flagged "CLAUDE.md missing <commands|agents|hooks> section" regardless of
whether the plugin actually had that component — e.g. graceful-handoff (no commands/ or
agents/ dir) got two spurious medium findings. Same over-report class as the model-field fix.

Now gated on component presence (pluginShipsComponent): a section is required only if the
plugin ships that component (commands/ or agents/ with .md, or hooks/hooks.json). Across the 5
stable plugins this drops 12 spurious findings to 3 legitimate ones (graceful-handoff hooks,
ai-psychosis commands+hooks). New fixture plugin-section-coverage proves both directions.
Found via the marketplace-wide review. Suite 950/950, self-audit A/A, scanner count 13.
tests badge 949 -> 950.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 09:06:20 +02:00
a5cfc331fd fix(plh): align required-frontmatter with CC docs (drop optional model/tools/name)
plugin-health-scanner required `model`+`allowed-tools` on commands and `model`+`tools`
on agents, plus `name` on commands. Per primary docs these are OPTIONAL:
- Commands/skills (code.claude.com/docs slash-commands): "All fields are optional. Only
  `description` is recommended." `name` defaults to the directory name.
- Subagents (code.claude.com/docs sub-agents): "Only `name` and `description` are required";
  `model` defaults to `inherit`, `tools` inherits all.

REQUIRED_COMMAND_FRONTMATTER -> [description]; REQUIRED_AGENT_FRONTMATTER -> [name, description].
This was over-reporting: every command without an explicit `model` got a spurious medium
finding (10 on the okr plugin alone). Found via the okr pilot review. Suite 949/949, self-audit
A/A, scanner count 13 (no new scanner).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 07:27:44 +02:00
47efc5979e chore: track STATE.md per updated global continuity rule
Stop gitignoring STATE.md (remove from .gitignore) and commit it. Per the
updated ~/.claude/CLAUDE.md continuity rule, STATE.md is now tracked in every
repo and pushed to private Forgejo (never GitHub/public) — it survives fresh
clones, git clean, and branch switches instead of living only on disk.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 23:01:36 +02:00
a86b92e2b5 release: v5.4.0 — plugin-hygiene & settings-validation hardening
Three additive findings extend existing PLH and SET scanners (scanner count
stays 13; --json/--raw byte-stable):
- CA-PLH-015 plugin-folder shadowing — a plugin.json component-path key in the
  replaces-set (commands/agents/outputStyles) pointing at a custom path while
  the default folder still exists; mirrors CC /doctor & claude plugin list.
- CA-PLH-016 skills:-array validation — each entry must resolve to a directory
  inside the plugin root; flags non-string/escapes-root/not-found/not-a-directory;
  mirrors claude plugin validate.
- CA-SET autoMode — structure (only environment/allow/soft_deny/hard_deny string
  arrays; "$defaults" valid) = medium; dead-config (autoMode in shared
  .claude/settings.json is not read by CC) = low.

Release mechanics: version 5.3.0 -> 5.4.0 (plugin.json); CHANGELOG [5.4.0];
README badge/TOC/What's-New/version-history; knowledge v5.4.0 scanner-backing
facts. Gates: suite 949/949, self-audit A/A + readmeCheck.passed (count 13),
SC-5 byte-equal, gitleaks clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 22:32:38 +02:00
d77c18aa53 docs(plan): mark v5.4.0 Feature 3 DONE — Session B complete [skip-docs]
CA-SET autoMode shipped (3633571). All 3 Session B features built+pushed;
suite 936->949. Premise #3 confirmed against primary source (per-file-scope
gate passed). Next: Session C release (needs operator GO).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 22:15:26 +02:00
3633571c7e feat(set): validate autoMode structure + flag it in shared settings (CA-SET)
settings-validator now validates the autoMode block (auto-mode classifier
config). Structure (medium): autoMode must be an object whose only keys are
environment/allow/soft_deny/hard_deny, each a string array ("$defaults" is a
valid entry); flags not-an-object, unknown-subkey, not-string-array. Dead-config
(low): Claude Code does not read autoMode from shared project settings
(.claude/settings.json), so an autoMode block committed there has no effect —
keyed on file.scope === 'project'.

Both premises primary-source-verified (code.claude.com/docs/en/auto-mode-config).
The plan's "test per-file scope first" gate passed: ConfigFile already carries
scope. SET is in the orchestrator; SC-5 re-checked, byte-equal (snapshot fixture
has no autoMode). Fixtures force-added (.claude/ is gitignored).

Tests +5 (944->949). Scanner count unchanged (13). --json/--raw byte-stable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 22:14:26 +02:00
9fd14aee99 docs(plan): mark v5.4.0 Feature 2 DONE + dropped-claim note [skip-docs]
CA-PLH-016 shipped (76d5eda). Records that the plan's unverified "CC suggests
the parent directory" error text was dropped (not in primary docs).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 21:58:00 +02:00
76d5eda101 feat(plh): validate plugin.json skills:-array entries (CA-PLH-016)
PLH now validates each entry of a plugin.json `skills` field (string|array):
every entry must resolve to an existing directory inside the plugin root.
One finding per bad entry (medium, plugin-hygiene), problem ∈ {non-string,
escapes-root, not-found, not-a-directory}. Mirrors `claude plugin validate`.
String|array normalized so a non-string top-level value is caught too.

Verifiseringsplikt: the plan's "CC suggests the parent directory" error text
is NOT in the primary docs — dropped. Only the four primary-source-verified
conditions are asserted (escape backed by the path-traversal rule).

Tests +3 (941->944). Scanner count unchanged (13). --json/--raw byte-stable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 21:56:57 +02:00
9b63125f6a docs(plan): mark v5.4.0 Feature 1 DONE + field-set correction [skip-docs]
CA-PLH-015 shipped (7abc5a1). Records the Verifiseringsplikt correction:
replaces-only field set (commands/agents/outputStyles); skills excluded
(adds-to-default), hooks/mcpServers/lspServers excluded (own merge rules).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 21:48:12 +02:00
7abc5a1dcb feat(plh): flag plugin.json paths that shadow default folders (CA-PLH-015)
PLH now flags a plugin.json component-path key (commands/agents/outputStyles)
that replaces a default folder still present on disk — Claude Code stops
scanning that folder, so its contents are silently ignored (dead config).
Mirrors CC's /doctor & `claude plugin list` warning (v2.1.140+).

Field set pinned to the docs' "replaces" category only (Verifiseringsplikt,
code.claude.com/docs path-behavior-rules): skills is excluded (adds to the
default skills/ scan — both load) as are hooks/mcpServers/lspServers (own
merge rules); a custom path that addresses the default folder is not flagged.

Tests +5 (936->941). Scanner count unchanged (13). --json/--raw byte-stable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 21:46:22 +02:00
c6992cad57 docs(plan): v5.4.0 Session A — reconciliation + ship-list GO (Option A) [skip-docs]
Re-verified the 6 deferred v5.4 candidates against HEAD on two axes (code-state +
CC premise, primary-source where risky). Operator GO "Option A": ship #1 PLH
shadow-folder + #5 PLH skills:-array + #4 autoMode structure/dead-config; no new
scanner (badge stays 13). #2 acceptEdits-writes + #6 nested-.claude deferred; #3
Read-deny/Glob-Grep WONTFIX — premise refuted by code.claude.com/docs (Read deny
already covers Glob/Grep), matrix row 175 was framed backwards and is now fixed.

- New: docs/v5.4.0-release-plan.md (Session A→C, mirrors v5.3.0 structure)
- Updated: docs/cc-2.1.x-gap-matrix.md (## v5.4.0 reconciliation block + row-175 fix)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 21:24:58 +02:00
fe686b6594 release: v5.3.0 — permission-rule & plugin-hygiene hardening (DIS/CML/PLH)
Five additive scanner findings extend existing scanners (count stays 13): DIS forbidden-param rules (Tool(param:value) on a canonicalizing field — deny/ask = false security, allow = dead config) and ineffective allow-wildcards + Tool(*) deny-all; CML context-window-scaled 40.0k-char CLAUDE.md budget mirroring CC's startup warning; PLH plugin namespace collision (two plugins declaring the same name). PLH cross-plugin command-name overlap reframed HIGH → LOW (namespacing keeps both reachable). feature-gap recommends disableBundledSkills under skill-listing pressure.

Version sync: plugin.json 5.2.0→5.3.0, README version badge + What's New + version-history row + TOC anchor (tests badge already 936+, scanner count stays 13), CHANGELOG [5.3.0] entry, 3 knowledge-backing entries. 936/936 tests; self-audit configGrade A 97, pluginGrade A 100, readmeCheck.passed:true; gitleaks clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 20:45:29 +02:00
891f9506bf docs(plan): v5.3.0 Session A — scope decision (ship-list empty) + gap-matrix reconciliation [skip-docs]
Session A audit (read/plan only). Reconciled docs/cc-2.1.x-gap-matrix.md (the
v5.2.0 plan) against HEAD: the entire HIGH-priority false-positive cluster is
already CLOSED in v5.2.0 (settings keys + xhigh, 28 hook events, MCP POSIX/trust,
claude-md HIGH->MEDIUM reframe, DIS/CNF param-aware). The 8 unreleased commits add
incrementally; no active bug or false positive remains unshipped.

Operator GO 2026-06-19: "Release-only, ship-list tom" -> Session B SKIPPED. All
still-open gaps are M-effort enhancements -> defer to v5.4. v5.3.0 = docs +
knowledge + release of the 8 commits already on main.

Output: scope-decision block appended to docs/v5.3.0-release-plan.md (empty
ship-list + verbatim GO + full draft CHANGELOG bullets mapping all 8 commits +
What's New draft + 3 knowledge-backing entries for Session C); reconciliation
block prepended to docs/cc-2.1.x-gap-matrix.md. Version confirmed 5.3.0 (minor,
non-breaking). 9 commits 1576909..HEAD = 8 mapped + 1 excluded (the plan commit).
2026-06-19 20:30:04 +02:00
9b828fab4c docs(plan): v5.3.0 multi-session release plan [skip-docs]
Three-session plan to release the 8 unreleased commits accumulated on main since
v5.2.0 as v5.3.0, with a fully updated README + CHANGELOG. Session A: gap-matrix
reconciliation + scope decision (operator GO on any new features). Session B:
conditional feature implementation (TDD). Session C: version sync, What's New,
CHANGELOG block, gates, tag, push. Each session has a testable Verifisering
section; key assumptions (non-breaking, docs-gate, number-only readme check)
flagged to test rather than trust. Mirrors docs/v5.2.0-release-plan.md.
2026-06-19 20:06:03 +02:00
96743ecce1 fix(readme): add missing SKL row to scanner table (12->13 rows) [skip-docs]
The scanner table listed 12 rows while the badge, prose ("13 orchestrated
scanners"), and self-audit count all said 13 — the SKL skill-listing-budget
scanner (shipped in v5.2.0) was never added as a table row. self-audit
--check-readme didn't catch it: it checks the badge NUMBER against the
filesystem (13==13), not the prose table's completeness. Adds the SKL row
(CA-SKL-001 listing cap, CA-SKL-002 aggregate listing budget).
2026-06-19 20:03:50 +02:00
0874188fe4 fix(plh): downgrade cross-plugin command-name overlap to low ambiguity
Commands are namespaced (/name:command), so a command name shared by two
differently-named plugins keeps both reachable — it is ambiguity, not a hard
conflict. The check now mirrors COL's plugin-vs-plugin skill finding: severity
LOW (was HIGH), category plugin-hygiene, COL-shaped details.namespaces, and a
group-first shape (one finding per command name listing every namespace, not
pairwise). It keys on the declared namespace (was folder basename) and fires
only across 2+ distinct namespaces — when plugins share a declared name, the
namespace-collision finding (medium) is the right signal, so this stays silent.

Removes the inaccurate "only one wins" humanizer entry. Adds fixtures
(duplicate-command-name; a shared command in duplicate-plugin-name's colliding
namespace) and 4 tests. Suite 932->936. self-audit A 97 / A 100, scanners 13.
2026-06-19 15:30:35 +02:00
c6c5f17752 feat(plh): flag plugin namespace collisions (same declared name)
Two plugins that declare the same `name` in plugin.json collapse into one
component namespace (/name:command, name:skill, agent "name"). Resolution
between two installed same-name plugins is undocumented, so one plugin's
commands/skills/agents are silently shadowed and unreachable. PLH now flags
this at medium severity, keying on the declared `name` (not folder basename,
via new declaredName on scanSinglePlugin) with a COL-shaped details.namespaces
payload. Name-less plugins are excluded from the collision map.

Search-first (code.claude.com/docs/en/plugins): plugin components are
namespaced by the declared name, so a plugin component can never shadow a
user/project one — only a same-name collision loses components. This refutes
the original "plugin vs user vs project shadowing" framing in the backlog.

Adds humanizer pattern, fixture (duplicate-plugin-name: 2 colliding + 2
name-less), and 3 tests. Suite 929->932. self-audit A 97 / A 100, scanners 13.
2026-06-19 15:13:19 +02:00
d678765fad feat(dis): flag forbidden-param permission rules CC silently ignores
Extends the DIS scanner and its shared permission-rules lib with a third
documented Claude Code permission footgun. Verified verbatim against
code.claude.com/docs/en/permissions (fetched 2026-06-19).

CC's Tool(param:value) matching (2.1.178) is off-limits for a tool's own
canonicalizing field — CC ignores such a rule and emits a startup warning,
because e.g. Bash(command:rm *) is bypassable by a compound command. The
forbidden fields: command (Bash/PowerShell), file_path (Read/Edit/Write),
path (Grep/Glob), notebook_path (NotebookEdit), url (WebFetch).

- lib/permission-rules.mjs: new forbiddenParamRule(entry) returning
  { tool, key, hint } or null. Only the param:value form (colon present)
  whose key equals the tool's forbidden field is flagged; Bash(npm:*),
  WebFetch(domain:host), Agent(model:opus), and Bash(command) (no colon)
  are left valid. FORBIDDEN_PARAMS map is the single source of truth.
- DIS: scans allow + deny + ask and splits severity by intent — deny/ask
  hits are false security (medium: the block never applies), allow hits are
  dead config (low: param:value matching is deny/ask-only). Two findings,
  permissions-hygiene, CA-DIS-NNN.
- 11 new tests (7 lib, 4 DIS) + 1 fixture forbidden-param-permissions
  (force-added past .gitignore .claude/). Suite 918 -> 929. Snapshot
  unchanged (SC-5 byte-equal), contamination grep clean, gitleaks clean.
  README/CLAUDE document the broadened DIS mandate; test badge synced.
  self-audit: PASS, configGrade A 97, pluginGrade A 100, scanners 13.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 14:04:25 +02:00
b0bf8c5817 feat(cml): context-window-scaled CLAUDE.md char budget (mirrors CC 40.0k warning)
Add a char-based CML finding that mirrors Claude Code's own startup warning
("Large CLAUDE.md will impact performance (X chars > 40.0k)"). CC 2.1.169 scales
that threshold with the model's context window, so the finding anchors on the
conservative 200k window (we cannot observe the user's window; the anchor fires
earliest) and discloses the relaxed ~200,000-char figure at 1M context. MEDIUM
severity (token cost, not an adherence cliff — consistent with the v5.2.0 reframe).

Keyed on chars, not lines, so it is complementary to the existing 200/500-line
checks (which stay untouched): a file can be long by lines yet under budget (short
lines, e.g. large-cascade at 37k chars / 1024 lines), or short by lines yet over it.

Extract the shared 200k/1M context-window constants to scanners/lib/context-window.mjs
(single source of truth; skill-listing-budget.mjs now imports + re-exports them).

40.0k figure and context-window scaling verified against the CC changelog (2.1.169,
2026-06-08) and the live startup-warning text. +6 tests, new fixture large-claude-chars
(48,531 chars / 100 lines). Suite 918/918, self-audit PASS configGrade A 97.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 13:34:40 +02:00
03949c6c98 feat(dis): flag ineffective allow wildcards; treat Tool(*) as deny-all
Extends the DIS scanner and its shared permission-rules lib with two
documented Claude Code permission footguns. Verified verbatim against
code.claude.com/docs/en/permissions (fetched 2026-06-19).

- lib/permission-rules.mjs: new isIneffectiveAllowGlob(entry) — unanchored
  tool-name globs in permissions.allow (`*`, `B*`, `mcp__*`) that CC silently
  skips ("does not auto-approve anything"); valid only as a glob-free
  `mcp__<server>__*`. Shared with CNF.
- lib/permission-rules.mjs: dominates() now treats the `Tool(*)` deny-all glob
  as equivalent to a bare deny (covers a bare allow) — CC: "Bash(*) is
  equivalent to Bash ... both forms remove the tool from Claude's context".
- DIS: new finding "Ineffective allow wildcard — Claude Code ignores this rule"
  (low, permissions-hygiene, CA-DIS-NNN); the existing dead-allow finding now
  also catches a bare allow killed by a Tool(*) deny.
- 9 new tests (5 lib, 4 DIS) + 2 fixtures (force-added past .gitignore .claude/).
  Suite 903 -> 912. Snapshot unchanged, contamination grep clean. README/CLAUDE/
  scanner-internals document the broadened DIS mandate; test badge synced.
  self-audit: PASS, configGrade A 96, pluginGrade A 100, readme gate passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-19 06:31:18 +02:00
dfe9049b55 feat(feature-gap): recommend disableBundledSkills under skill-listing pressure
Chunk 2 of the disableBundledSkills GAP feature. Adds a conditional GAP check
that prescribes the `disableBundledSkills` lever — but only when the active
skill listing is measurably over budget (SKL's CA-SKL-002 overflow signal) and
the lever is un-pulled. It stays an opportunity, not noise.

Bundled skills (/code-review, /batch, /debug, /loop, /claude-api, …) live in the
CC binary, not on disk, so their exact cost is unmeasurable here — the finding
says so plainly, and frames the lever as zero-cost budget reclaim that leaves
the user's own skills untouched. CC 2.1.169+.

- Pure, exported bundledSkillsLeverFinding({leverPulled, aggregate}) → finding|null
  (severity low, category token-efficiency, CA-GAP-NNN), wired into scan() via the
  shared measureActiveSkillListing().
- Lever resolved via new isBundledSkillsDisabled(): env var + settings cascade
  read directly, because discovery does NOT tag ~/.claude/settings.json (its
  relPath lacks ".claude" when walked from the .claude root) — the dominant
  user-scope location for this global preference would otherwise be missed.
- GAP scan() now reads HOME → existing GAP tests retrofitted to withHermeticHome
  per the hermetic rule. Snapshots unchanged, contamination grep clean.
- 16 new tests (9 GAP, 7 lib). Suite 887 -> 903. README/CLAUDE.md document the
  cross-scanner remediation; test counts synced. self-audit: PASS, configGrade
  A 96, pluginGrade A 100, readme gate passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 21:38:19 +02:00
0a631e3061 refactor(skl): extract skill-listing budget to shared lib (single source of truth)
Chunk 1 of the disableBundledSkills GAP feature. Moves the per-description cap,
aggregate budget constants, calibration note, and the enumerate-and-measure step
out of skill-listing-scanner into scanners/lib/skill-listing-budget.mjs — so SKL
(diagnoses overflow) and the upcoming GAP check (prescribes disableBundledSkills)
consume one budget definition instead of two divergent copies.

- New lib: assessSkillListingBudget (pure aggregate math) + measureActiveSkillListing
  (HOME-scoped enumerate-and-measure wrapper).
- SKL delegates measurement; all finding strings kept byte-identical. 18/18 SKL
  tests pass unchanged → behavior-neutral refactor.
- 12 new lib unit tests pin the budget contract. Suite 875 -> 887.
- README badge + CLAUDE.md test counts synced (self-audit --check-readme: passed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 21:23:01 +02:00
157690993f release: v5.2.0 — CC 2.1.114→181 compat + skill-listing budget (SKL)
New orchestrated scanner SKL (CA-SKL-001 1,536-char listing cap, CA-SKL-002
listing-budget sum) → 13 orchestrated scanners. Five validators refreshed for
the CC 2.1.114→181 settings/hook surface (xhigh effort, MessageDisplay +
post-session events). False positives eliminated in MCP and permissions
scanners. Hermetic HOME isolation across all CLI-spawning tests.

Version sync: plugin.json 5.1.0→5.2.0, README badges (version + tests-875+),
5 stale "12→13 scanners" prose fixes, What's New + version-history rewrite,
CHANGELOG [5.2.0] entry. 875/875 tests; self-audit configGrade A, pluginGrade A,
readmeCheck.passed:true.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 20:37:45 +02:00
5b28e84966 test(fix-cli): close missed HOME-leak — isolate spawns + lock SKL/COL out
Devil's advocate gap-verification (read-only Workflow, 9 skeptics) refuted the
blanket "all closed" claim by finding fix-cli.test.mjs was the one CLI-spawning
test still reading the real ~/.claude. fix-cli runs the SKL skill-listing
scanner (HOME-scoped) even with includeGlobal:false, so its manual findings
include CA-SKL-001 on a dev machine but not in clean CI.

This directly corrects 325182d, which listed fix-cli.test.mjs as "Proven safe,
left as-is (output byte-identical real vs empty HOME — fixable findings are
project-local HKV/RUL/SET, never SKL/COL)". That reasoning predated SKL being
wired into scan-orchestrator (7bb2547/66433fe) and was false: real HOME yields
manual=6 (incl. CA-SKL-001), hermetic manual=5 (5230 vs 4798 bytes).

- wrap all 5 fix-cli spawns in hermeticEnv() (matches the other 11 CLI tests)
- add a regression lock: a project-scoped run must surface no CA-SKL/CA-COL
- redirect the --apply backup check to HERMETIC_HOME — the test was also
  writing backups into the real ~/.config-audit/backups on every run
- docs: stale "26 hook events" -> 28 (README:528, scanner-internals:73);
  hook-validator.mjs comment April -> June 2026 (functional count already 28)

Re-audited all 12 CLI-spawning test files: 11 hermetic-helper, manifest custom
HOME-env, post-edit-verify safe-by-construction (early-exit only, never reaches
scanners). HOME-leak class now actually closed. Suite 875/875, no snapshot drift.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 18:50:20 +02:00
bc85b79eb9 docs(plan): devil's advocate brief — adversarial gap-review verification
Brief for the next session (post-/clear): independently DISPROVE the
"CC 2.1.114-181 feilretting is complete" claim rather than re-assert it.

Designed to run as a Dynamic Workflow (Workflow tool): one read-only
Explore skeptic per scanner surface (settings/hooks/mcp/permissions-DIS/
claude-md/knowledge/token) tries to find a still-open or superficially-
fixed row, plus adversarial review of this week's CA-SKL-002 + HOME-leak
work, then synthesis into a punch-list or a verified attestation.

Includes the 3 claims to attack (P1 no active false positives, P2 fixes
correct-not-superficial, P3 new work regression-free), the surface-cluster
table with attack angles, a runnable script skeleton + output schemas, and
a read-only scope-fence (findings → report, no fixes without approval).

STATE.md (gitignored, polyrepo) updated on-disk to hand this off as the
next session's active task.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 18:28:02 +02:00
325182ddc9 test: isolate HOME in all CLI-spawning tests (close leak class)
Follow-up to the posture-grade-stability fix in 66433fe. Audited every
test that spawns a CLI and found more of the same class: tests running
HOME-scoped scanners (SKL/COL) or the CLAUDE.md cascade against the
developer's real ~/.claude instead of an isolated HOME.

Fixed (env: hermeticEnv()):
- posture.test.mjs        — runs full posture (SKL/COL/cascade); twin of
                            the posture-grade-stability leak, masked only
                            because its asserts are structural/relative
- drift-cli.test.mjs      — ACTIVE bug: the CLI wrote baselines into the
                            real ~/.claude during the run (pollution); now
                            isolated, and afterEach cleanup wrapped in
                            withHermeticHome so it looks in the same HOME
- token-hotspots-cli.test.mjs — scan-orchestrator run executes SKL/COL on
                            real HOME; TOK reads the HOME cascade
- accurate-tokens.test.mjs — TOK reads the HOME cascade (kept the
                            ANTHROPIC_API_KEY deletion)

Proven safe, left as-is (no HOME-scoped scan affecting assertions, no
HOME writes): post-edit-verify.test.mjs (fast-path early-returns only),
fix-cli.test.mjs (output byte-identical real vs empty HOME — fixable
findings are project-local HKV/RUL/SET, never SKL/COL),
lint-default-output (caller already uses withHermeticHome).

Suite 875/875, no snapshot drift. No test regressed under isolation,
confirming none had a hidden real-HOME dependency.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 18:17:09 +02:00
66433fee48 feat(skill-listing): add CA-SKL-002 aggregate listing-budget check
Syklus 2 of Fase 4 Items 2+3. Flags when the sum of active skill
descriptions exceeds the listing budget (~2% of context, CC 2.1.32).

Design (operator-confirmed "fact-first, 200k anchor"):
- low severity (estimate) vs medium for the verified 1,536-char cap
- each description counted up to the 1,536 cap (what actually loads in
  the listing) — avoids double-counting the tail CA-SKL-001 flags
- fires when sum > 2% x 200k = 4000 tok; evidence leads with the measured
  sum + a calibration note that the budget scales 5x on 1M-context models
- aggregate emitted after the per-skill loop so the common case reads
  001=cap, 002=aggregate (finding IDs are a sequential counter, not stable
  semantic IDs — tests match on title, never NNN)

Also:
- tailored humanizer static entry for the aggregate title
- fix latent HOME leak in posture-grade-stability.test.mjs: it spawned
  posture.mjs without hermeticEnv(), so a real ~/.claude leaked HOME-scoped
  SKL/COL findings into the baseline grade (Token Efficiency A->B). Now
  isolated like the 8 other CLI-spawning tests.
- docs sync: test count 868->875, scanner-internals, gap-matrix, plan status

Suite 875/875, no snapshot drift, self-audit clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 18:06:17 +02:00
7bb254780a feat(skill-listing): add SKL scanner for the skill-listing token budget
Fase 4 Items 2+3 (CC 2.1.114→181 gap-review). New orchestrated scanner
`skill-listing-scanner.mjs` (prefix SKL) flags every active skill whose
description exceeds the verified 1,536-char listing cap (CC 2.1.105, changelog
L1502). Past the cap, Claude Code silently truncates the description the model
reads to route skill invocation — dropping the trigger phrases at the tail.
HOME-scoped over all user + plugin skills via enumerateSkills (COL is the model).

- CA-SKL-001 (medium): description > 1,536 chars. Remediation folds in Item
  2(b) — recommends disableBundledSkills + skillOverrides + trimming
  (designvalg A: no standalone GAP-check, which would fire for nearly everyone).
- Designvalg B: v1 ships the verified cap ONLY. The aggregate 2%-of-context
  listing budget is deferred — it needs a context-window assumption that would
  turn a verified fact into a guess (would carry a CALIBRATION_NOTE if added).
- Choice C: recognize the skillOverrides settings key (CC 2.1.129) in
  KNOWN_KEYS. Left OUT of TYPE_CHECKS — the value is a per-skill object
  (off/user-invocable-only/name-only), not a string; a 'string' check (as the
  plan sketched) would create a NEW false positive. Verify-first deviation.

Registration: scan-orchestrator (13th scanner), humanizer (SKL → 'Wasted
tokens' + static/_default translations), scoring SCANNER_AREA_MAP (→ Token
Efficiency; no 11th area), README badge 12→13, CLAUDE.md (finding-id +
test-count), docs/scanner-internals.md, gap-matrix + plan status notes.

Snapshots reseeded hermetically (SEED_SNAPSHOT/UPDATE_SNAPSHOT): SKL entry with
0 findings in empty HOME, scanners_ok 11→12, claudeMdEstimatedTokens bump from
the CLAUDE.md edits flowing through the cascade. Contamination grep clean.

Suite 868/868 (856 baseline + 11 SKL + 1 skillOverrides). RED→GREEN logged
per cycle.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 17:36:28 +02:00
fcd8ad3048 docs(plan): Fase 4 Items 2+3 execution plan (verify-first) + matrix row 166 DONE
Plans the skill-listing token-management work (gap-matrix rows 167+169) for a fresh session. All version facts re-verified against the changelog cache before writing: disableBundledSkills v2.1.169; skill-listing budget = 2% of context (v2.1.32) + 1,536-char per-description cap (v2.1.105); skillOverrides v2.1.129.

Key findings from code investigation:
- Item 2(a) (disableBundledSkills unknown-key false-positive) is ALREADY fixed (Batch 1: KNOWN_KEYS + TYPE_CHECKS); matrix row 166 marked DONE.
- Item 2(b): a standalone binary GAP check would be noise; recommend folding the disableBundledSkills recommendation into Item 3 SKL scanner remediation (designvalg A).
- Item 3: new SKL scanner; lead with the verified 1,536-char truncation cap (high confidence); aggregate 2%-budget is an estimate needing a context-window assumption (designvalg B) — calibration-noted.
- Reuse enumerateSkills() + parseFrontmatter; document the boundary vs TOK pattern F.
- Matrix row 170 keys (skillListingBudgetFraction/maxSkillDescriptionChars) do NOT exist (verified).

Plan doc carries exact file:line anchors, scanner-registration touchpoints, failing-test-first cycle specs, and 2 open design decisions to confirm before coding. STATE.md updated to resume here after /clear.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 15:41:52 +02:00
8376dab83f fix(tokens): refresh stale "Opus 4.7" framing to model-neutral + Opus 4.8 anchor
Fase 4 token-opt, Item 1 of gap-review NEXT STEP #2. The prompt-cache pattern corpus + TOK scanner were frozen at an "Opus 4.7" framing after CC shipped Opus 4.8 (default, 2.1.154) and Fable 5 (2.1.170). Model-era facts re-verified against the official changelog cache before editing.

The patterns are properties of prompt-caching, not of any model, so mechanic text is now model-neutral with a single "current default: Opus 4.8" anchor — preventing a re-freeze at the next model bump.

- rename knowledge/opus-4.7-patterns.md -> prompt-cache-patterns.md (git mv, history preserved); 6 reference sites updated
- TOK scanner: line-318 finding text (human-facing) made model-neutral; header + cache-prefix-scanner + CLI comments refreshed
- configuration-best-practices.md body + footnote 4.7 -> 4.8
- human-facing docs: commands/{tokens,help,manifest}.md, project CLAUDE.md, README, docs/scanner-internals.md
- gap-matrix row marked DONE; future Items 2/3 retargeted to new filename

Failing-test-first (Iron Law): +2 knowledge staleness guards (era-anchor + no-refreeze) +1 scanner assertion (no stale model anchor in finding text). Suite 853 -> 856 green; zero snapshot drift; self-audit A(97) PASS. CHANGELOG / v5 plan / ratified gap-plan keep historical opus-4.7-patterns refs (correct record of past state).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 15:27:36 +02:00
b3c572ad46 fix(mcp-config-validator): remove invented trust field (verify-first)
`.mcp.json` has no per-server `trust` key — verified 2026-06-18 against
code.claude.com/docs/en/mcp + /settings. MCP server approval is
dialog/settings-based (enableAllProjectMcpServers / enabledMcpjsonServers /
disabledMcpjsonServers), never a JSON field. The scanner's "Missing trust
level" (CA-MCP-001, medium) and "Invalid trust level" (high) were false
positives flagging a field that does not exist.

- scanner: delete both trust checks + VALID_TRUST_LEVELS; drop `trust` from
  VALID_SERVER_FIELDS so a stray `trust` is now flagged as an unknown field
- humanizer: remove the two trust-level entries
- knowledge (5 files): point to the real approval mechanism, not a trust field
- fixtures: scrub `trust` (incl. the invalid "local" in optimal-setup)
- tests: flip assertions (no trust-level finding; stray trust -> unknown
  field) + add knowledge-staleness re-freeze guards
- snapshots: reseed (marketplace-medium .mcp.json -8 tokens, hermetic)
- gap-matrix: mark the trust verify-first item DONE

Suite: 853/853 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 14:22:56 +02:00
624f5edabc docs(knowledge): refresh corpus to Opus 4.8 era (CC 2.1.114->181 Batch 3)
The knowledge base was frozen at ~v2.1.111, describing an Opus-4.7 world
after CC shipped Opus 4.8 and Fable 5. Refreshed three agent-facing
knowledge files; every fact re-verified against the official changelog
(~/.claude/cache/changelog.md, CC 2.1.181) on 2026-06-18.

- feature-evolution.md: new Opus-4.8-era rows above the v2.1.111 freeze --
  Opus 4.8 default + /effort xhigh (2.1.154), Fable 5 Mythos-class
  (2.1.170), post-session hook (2.1.169), MessageDisplay (2.1.152),
  /simplify -> /code-review (2.1.147), /config key=value (2.1.181).
- hook-events-reference.md: 26 -> 28 events (+MessageDisplay, +post-session);
  documented Stop/SubagentStop additionalContext output field.
- claude-code-capabilities.md: 2026-06 model/effort lineup table;
  bundled skills /simplify -> /code-review; documented /config key=value.
- cc-2.1.x-changelog-delta.md: marked SUPERSEDED by the gap-matrix.

Verification caught two version errors in STATE/matrix, corrected to the
changelog:
- Stop/SubagentStop additionalContext is 2.1.163, not 2.1.165 (2.1.165 was
  "Bug fixes" only; matrix row 109 already said 2.1.163).
- settings `agent` field introduced 2.0.59; 2.1.157 = honored for
  dispatched `claude agents` sessions, not the introduction.

New tests/knowledge/knowledge-staleness.test.mjs (8 tests) encodes the
verified facts as a re-freeze guard (RED before edits, GREEN after).

Full suite: 850/850 green (+8). self-audit PASS, A(100)/A(97).
2026-06-18 13:35:00 +02:00
feaa7ed2e4 fix(claude-md-linter): reframe CLAUDE.md length from HIGH adherence cliff to MEDIUM token cost
The >500-line check emitted HIGH severity with "Files over 500 lines
significantly reduce Claude's adherence to instructions." CC 2.1.169
scaled the "too long" warning by context window, and the plugin's own
configuration-best-practices.md:97 footnote already says raw line count
is a Sonnet-era heuristic superseded by cache-prefix stability — so the
absolute HIGH + universal adherence claim is now-wrong.

- >500 lines: HIGH -> MEDIUM, reworded to token-cost-every-turn +
  smaller-context-model caveat + cache-prefix pointer; notes CC 2.1.169
  scales the threshold by context window.
- >200 lines: stays MEDIUM, dropped the absolute "optimal adherence"
  framing for the same token/context-window framing.

Aligns the scanner with anti-patterns.md:7 (CA-CML-001 = medium) and
configuration-best-practices.md:97. No snapshot impact (byte snapshots
use a 24-line fixture CLAUDE.md).

Full suite: 842/842 green (+5). self-audit PASS, A(100)/A(97).
2026-06-18 13:15:38 +02:00
bec3f45329 fix(permissions): param-aware DIS dead-allow + CNF conflict matching
The DIS scanner collapsed Tool(param) rules to the bare tool name, so
Agent(model:opus) deny + Agent(model:sonnet) allow (and the same for
WebFetch(domain:...)) were flagged as dead config — a false positive now
that CC 2.1.178 matches Tool(param:value) and 2.1.172 adds domain rules.
The conflict-detector shared the blind spot from the other side: a
wildcard deny like WebFetch(domain:*) did not cover a
WebFetch(domain:good.com) allow, so a genuine cross-scope conflict was
missed (false negative).

New shared scanners/lib/permission-rules.mjs:
- parseRule / paramMatches (glob)
- dominates(deny, allow) -> DIS dead-allow (deny fully covers allow)
- rulesIntersect(a, b)   -> CNF cross-scope conflict (match sets intersect)

DIS now delegates to dominates; conflict-detector :156 delegates to
rulesIntersect. A bare deny still covers all params, so true positives
are preserved (Bash deny + Bash(npm:*) allow still flagged).

Re-seeded the marketplace-medium snapshots: the false-positive CA-DIS
finding (Read(src/**) allow + Read(./.env) deny) is correctly gone. This
changes snapshot CONTENT only — envelope schema is unchanged, so --json
and --raw stay byte-stable.

Full suite: 837/837 green (+25). self-audit PASS, A(100)/A(97).
2026-06-18 13:06:20 +02:00
8216fb4175 test(snapshots): make byte/snapshot tests hermetic + re-seed baseline
The COL collision-scanner and the CLAUDE.md cascade resolve ~/.claude from
process.env.HOME (active-config-reader). Snapshot/byte CLIs were spawned with
the developer's real HOME, so they picked up installed plugins/skills and the
user CLAUDE.md — making the v5.0.0 + default-output snapshots machine- and
time-dependent. They were seeded 2026-05-01 with COL=1 (a real ~/.claude skill
collision) and drifted to COL=0 after the polyrepo split: 26 pre-existing
failures unrelated to Batch 1.

Fix (test-only, no production change):
- tests/helpers/hermetic-home.mjs — empty temp HOME, mirroring the pattern
  collision.test.mjs already uses for the COL unit test.
- 7 harnesses spawn CLIs (or call lint()) under the hermetic HOME, so output
  depends only on committed fixtures. Determinism verified across runs.
- Re-seeded all snapshots under hermetic HOME via SEED_SNAPSHOT/UPDATE_SNAPSHOT
  (added a SEED guard to the frozen v5.0.0 byte tests). Snapshots now reflect
  the fixture alone (COL=0, fixture-only activeConfig counts).
- Also re-seeded the unused env-aware snapshots (manifest/whats-active/
  plugin-health), which had baked dozens of real ~/.claude skill/plugin names
  into the committed repo — privacy cleanup.

Full suite: 812/812 green, stable across 3 runs.
2026-06-18 12:26:00 +02:00
4b94da0f11 fix(mcp-config-validator): stop flagging auto-injected + POSIX env vars
Clears false positives on valid .mcp.json (gap matrix, Batch 1):
- ${CLAUDE_PROJECT_DIR} is auto-injected at runtime (CC 2.1.139) and never
  needs an env block — now allowlisted.
- POSIX expansions like ${VAR%pattern} / ${VAR:-default} are resolved by
  Claude Code (CC 2.1.142); the env-var regex now matches only bare
  ${IDENTIFIER}, so operator expressions are skipped.

Genuine bare unreferenced vars are still flagged (broken-project regression
intact). The MCP `trust` field is untouched — it is verify-first (point 4),
not part of Batch 1.

Tests: hermetic runtime temp-fixture; 23/23 MCP green, both directions covered.

Ref: docs/cc-2.1.x-gap-matrix.md

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 12:03:21 +02:00
98ddd777fb fix(hook-validator): recognize MessageDisplay + post-session events
Clears false positives where valid CC hook events were flagged as
"Unknown hook event" (gap matrix, Batch 1).

VALID_EVENTS += MessageDisplay (CC 2.1.152), post-session (CC 2.1.169,
kebab-case, distinct from SessionEnd). 26 -> 28; recommendation string
updated to match.

knowledge/hook-events-reference.md count stays for the Batch 3 knowledge
refresh. Tests: hermetic runtime temp-fixture; 15/15 HKV green.

Ref: docs/cc-2.1.x-gap-matrix.md

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 12:00:40 +02:00
73099354c7 fix(settings-validator): accept CC 2.1.114–181 keys + xhigh effort
Clears a cluster of active false positives where valid, documented
Claude Code config was flagged as unknown/invalid (gap matrix, Batch 1).

KNOWN_KEYS +11 (CC 2.1.133–181): allowAllClaudeAiMcps, disableBundledSkills,
enforceAvailableModels, fallbackModel, footerLinksRegexes, parentSettingsBehavior,
pluginSuggestionMarketplaces, requiredMaximumVersion, requiredMinimumVersion,
sandbox, wheelScrollAccelerationEnabled.

VALID_EFFORT_LEVELS += 'xhigh' (CC 2.1.154 Opus-4.8 top tier).
TYPE_CHECKS += disableBundledSkills/wheelScrollAccelerationEnabled (boolean).
fallbackModel intentionally NOT type-checked (string | array<=3).

Tests: hermetic runtime temp-fixture (path-guard blocks committing
settings.json); 29/29 SET green, full hermetic suite unaffected.

Ref: docs/cc-2.1.x-gap-matrix.md

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 11:58:42 +02:00
315ea2259f docs(config-audit): CC 2.1.114-181 coverage gap analysis (Fase 0-2)
Systematic review of Claude Code config-surface changes from the plugin's last-verified baseline (v2.1.114) to installed v2.1.181, against current scanner coverage. 12/12 surfaces, 162 verified gap rows via two background gap-analysis workflows (search-first, per-surface changelog verification).

Key finding: a cluster of ACTIVE FALSE POSITIVES - config-audit flags valid v2.1.181 config as wrong (effortLevel xhigh; ~12 settings keys incl. sandbox/fallbackModel/enforceAvailableModels/disableBundledSkills/agent; MCP ${CLAUDE_PROJECT_DIR} and POSIX expansions; param-qualified permission rules; MessageDisplay/post-session hooks). Recommended release v5.2.0 (byte-stable; fixes remove false findings). Voyage escalation: no.

- docs/cc-2.1.x-gap-review-plan.md - ratified plan (method A, floor v2.1.114)
- docs/cc-2.1.x-gap-matrix.md      - full gap matrix + buckets + release call
- docs/cc-2.1.x-changelog-delta.md - changelog corpus (superseded by matrix)

STATE.md updated (gitignored) - next session resumes at Batch 1.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-18 11:31:48 +02:00
65fd2f1eb8 chore(gitignore): add session/local-state baseline (polyrepo split) 2026-06-18 10:21:12 +02:00
252 changed files with 23526 additions and 4385 deletions

View file

@ -1,7 +1,7 @@
{
"name": "config-audit",
"description": "Multi-agent workflow for analyzing, reporting, and optimizing Claude Code configuration across your entire machine",
"version": "5.1.0",
"version": "5.13.0",
"author": {
"name": "Kjell Tore Guttormsen"
},

View file

@ -12,7 +12,7 @@ All command files MUST include:
---
name: plugin:command
description: Short description of what this command does
allowed-tools: Read, Write, Bash, Task
allowed-tools: Read, Write, Bash, Agent
model: sonnet
---
```

20
.gitignore vendored
View file

@ -23,3 +23,23 @@ S*-PROMPT.md
# v5 namespace research (local-only spike output)
docs/v5-namespace-research.md
# --- session/local state (gitignored) ---
# STATE.md is LOCAL-ONLY: this repo's only remote is the PUBLIC open/ mirror, and
# the global continuity rule says STATE.md must never reach a public mirror (it
# would leak internal state-of-play). Mirrors the linkedin-studio precedent.
# (Previously tracked under the old "always tracked" rule that assumed a private
# remote — corrected 2026-06-21. See ~/.claude/CLAUDE.md.)
STATE.md
REMEMBER.md
ROADMAP.md
TODO.md
NEXT-SESSION-PROMPT*.local.md
*.local.md
*.local.json
*.local.sh
# Local-only dogfood harnesses: they index the operator's private config by line
# number (see docs/subtraction-fasit.local.md) and must never reach the public mirror.
*.local.mjs
.DS_Store
.claude/

File diff suppressed because it is too large Load diff

View file

@ -1,13 +1,8 @@
# Config-Audit Plugin
Claude Code Configuration Intelligence — know if your configuration is correct, find what could improve it, fix it automatically.
Claude Code Configuration Intelligence — know if your config is correct, find what could improve it, fix it automatically. Three pillars: **Health** (deterministic scanners), **Opportunities** (context-aware recommendations), **Action** (auto-fix with backup/rollback).
## What this plugin does
Analyzes and optimizes Claude Code configuration across three pillars:
- **Health** — Deterministic scanners verify correctness, consistency, and completeness
- **Opportunities** — Context-aware recommendations for features that could benefit your project
- **Action** — Auto-fix with backup/rollback
Per-command flags, patterns, and feature lists live in `README.md` and `/config-audit help`. This file carries what's invariant for working on the plugin.
## Commands
@ -15,14 +10,15 @@ Analyzes and optimizes Claude Code configuration across three pillars:
| Command | Description |
|---------|-------------|
| `/config-audit` | Full audit with auto-scope detection (no setup needed) |
| `/config-audit posture` | Quick health scorecard (A-F grades, 10 quality areas incl. Token Efficiency, Plugin Hygiene) |
| `/config-audit tokens` | Opus-4.7-aware token hotspots (6 patterns: cache-breaking, redundant perms, deep imports, oversized cascade, bloated SKILL.md desc, MCP tool-schema budget) — optional `--accurate-tokens` API calibration, `--with-telemetry-recipe` cache-hit recipe pointer |
| `/config-audit manifest` | Ranked table of every system-prompt token source (CLAUDE.md, plugins, skills, MCP, hooks) sorted by estimated tokens |
| `/config-audit` | Full audit with auto-scope detection |
| `/config-audit posture` | A-F health scorecard (10 quality areas) |
| `/config-audit tokens` | Prompt-cache-aware token hotspots, each tagged with its load pattern; cache-aware |
| `/config-audit manifest` | Ranked table of every token source + always-loaded subtotal |
| `/config-audit feature-gap` | Context-aware feature recommendations grouped by impact |
| `/config-audit optimize` | Mechanism-fit lens (procedure→skill, lifecycle→hook, path→rule, never→permission). Agent-driven, **not byte-stable**. `--subtract` adds the subtraction axis (what no longer earns its always-loaded rent, `BP-SUB-001`) — opt-in, proposes only |
| `/config-audit fix` | Auto-fix deterministic issues with backup + verification |
| `/config-audit rollback` | Restore configuration from backup |
| `/config-audit plan` | Create action plan from audit findings |
| `/config-audit plan` | Create action plan from findings |
| `/config-audit implement` | Execute plan with backups + auto-verify |
| `/config-audit help` | Show all commands |
@ -32,7 +28,9 @@ Analyzes and optimizes Claude Code configuration across three pillars:
|---------|-------------|
| `/config-audit drift` | Compare current config against saved baseline |
| `/config-audit plugin-health` | Audit plugin structure, frontmatter, cross-plugin coherence |
| `/config-audit whats-active` | Read-only inventory of plugins, skills, MCP, hooks, CLAUDE.md active for a repo (with token estimates) |
| `/config-audit whats-active` | Read-only inventory of active plugins/skills/MCP/hooks/CLAUDE.md (with token estimates) |
| `/config-audit knowledge-refresh` | Refresh the best-practices register (stale check + web poll). Human-approved writes; **not byte-stable** |
| `/config-audit campaign` | Machine-wide audit ledger + token bill across repos. Human-approved writes; **not byte-stable** |
| `/config-audit discover` | Run discovery phase only |
| `/config-audit analyze` | Run analysis phase only |
| `/config-audit interview` | Gather user preferences (opt-in) |
@ -48,60 +46,48 @@ Analyzes and optimizes Claude Code configuration across three pillars:
| planner-agent | Create action plan | opus | yellow | Read, Glob, Write |
| implementer-agent | Execute changes | sonnet | magenta | Read, Write, Edit, Bash, Glob |
| verifier-agent | Verify results | sonnet | purple | Read, Glob, Grep |
| feature-gap-agent | Context-aware feature recommendations | opus | green | Read, Glob, Grep, Write |
| feature-gap-agent | Feature recommendations | opus | green | Read, Glob, Grep, Write |
| optimization-lens-agent | Mechanism-fit precision gate | opus | orange | Read, Glob, Grep, Write |
## Hooks
| Event | Script | Purpose |
|-------|--------|---------|
| PreToolUse | `auto-backup-config.mjs` | Auto-backup config files before Edit/Write |
| PostToolUse | `post-edit-verify.mjs` | Verify config files after Edit/Write, block on new critical/high |
| SessionStart | `session-start.mjs` | Checks for active (unfinished) sessions |
| Stop | `stop-session-reminder.mjs` | Reminds about current session phase |
| PreToolUse | `auto-backup-config.mjs` | Backup config files before Edit/Write |
| PostToolUse | `post-edit-verify.mjs` | Verify after Edit/Write, block on new critical/high |
| SessionStart | `session-start.mjs` | Check for active (unfinished) sessions |
| Stop | `stop-session-reminder.mjs` | Remind about current session phase |
## Reference docs (read on demand)
- **Scanner inventory, lib modules, action engines, knowledge base:** `docs/scanner-internals.md`
- **Plain-language output (v5.1.0), humanizer vocabularies, output modes:** `docs/humanizer.md`
- `docs/scanner-internals.md` — scanner inventory, lib modules, action engines, knowledge base, per-scanner/per-block implementation notes (design rationale, primary-source verification, byte-stability lessons)
- `docs/humanizer.md` — plain-language output (v5.1.0), humanizer vocabularies, output modes
## Plain-Language Output (v5.1.0) — summary
## Plain-Language Output (v5.1.0)
Default output of all 18 commands routes through `humanizeEnvelope` from `lib/humanizer.mjs`. Findings get three decorated fields:
- `userImpactCategory` — Configuration mistake / Conflict / Wasted tokens / Dead config / Missed opportunity
- `userActionLanguage` — Fix this now / Fix soon / Fix when convenient / Optional cleanup / FYI (derived from severity)
- `relevanceContext``affects-everyone` (default) / `affects-this-machine-only` (`*.local.*` files) / `test-fixture-no-impact`
`--raw` bypasses the humanizer for byte-stable v5.0.0 output. `--json` is also byte-stable. Full detail and Wave 5 lessons: `docs/humanizer.md`.
Default output of all commands routes through `humanizeEnvelope` (`lib/humanizer.mjs`), decorating each finding with `userImpactCategory`, `userActionLanguage`, and `relevanceContext`. `--raw` and `--json` bypass the humanizer for byte-stable v5.0.0 output. Full detail: `docs/humanizer.md`.
## Suppressions
Create `.config-audit-ignore` at project root to suppress known findings:
```
CA-SET-003 # Exact ID
CA-GAP-* # Glob pattern (all GAP findings)
```
Suppressed findings tracked in envelope's `suppressed_findings` for audit trail. Disable with `--no-suppress`.
Create `.config-audit-ignore` at project root — one exact ID or glob per line (`CA-SET-003`, `CA-GAP-*`). Suppressed findings are tracked in the envelope's `suppressed_findings` for audit trail. Disable with `--no-suppress`.
## Architecture
### Workflow
```
/config-audit → discover + analyze (auto) → plan → implement → verify
```
Default: auto-detects scope from git context. Override with `/config-audit full|repo|home|current`. Delta mode: `--delta` (incremental).
Workflow: `/config-audit → discover + analyze (auto) → plan → implement → verify`. Auto-detects scope from git context; override with `full|repo|home|current`; `--delta` for incremental. Session state lives under `~/.claude/config-audit/sessions/{id}/` (scope.yaml, discovery.json, state.yaml, findings/, analysis-report.md, action-plan.md, backups/, implementation-log.md).
### Session Directory
```
~/.claude/config-audit/sessions/{session-id}/
├── scope.yaml, discovery.json, state.yaml
├── findings/, analysis-report.md, action-plan.md
├── backups/, implementation-log.md
└── interview.md (if interview run)
```
Finding ID format: `CA-{SCANNER}-{NNN}` — e.g. `CA-CML-001`, `CA-SET-003`, `CA-HKV-002`, `CA-RUL-005`, `CA-TOK-005`, `CA-CPS-001`, `CA-SKL-001`, `CA-OST-001`, `CA-OPT-001`, `CA-AGT-001`.
### Finding ID Format
`CA-{SCANNER}-{NNN}` — e.g. `CA-CML-001`, `CA-SET-003`, `CA-HKV-002`, `CA-RUL-005`, `CA-TOK-005`, `CA-CPS-001`, `CA-DIS-001`, `CA-COL-001`
## Conventions
Enforced conventions live in `.claude/rules/` (auto-loaded as project instructions):
- `ux-rules.md` — output/narration/formatting for all commands (never dump raw JSON, narrate before each step, space-separated command suggestions)
- `command-development.md` — required command frontmatter + `plugin:action` naming
- `agent-development.md` — agent frontmatter + "when to use" conventions
- `state-management.md` — update `state.yaml` after every workflow phase
Coding style: scanners are zero-dependency Node ESM; new findings use the `CA-{SCANNER}-{NNN}` ID format; byte-stable CLIs are verified against frozen `tests/snapshots/v5.0.0/` baselines.
**Subtraction floor (invariant).** `optimize --subtract` is the only lens that proposes removing config, so `scanners/lib/floor-exclusion.mjs` runs as a deterministic pre-step *before* the judge — a load-bearing block is never a candidate, and that guarantee must not be moved into the agent prompt. Two rules follow from it: (1) **staleness is not a deletion signal** — an outdated version pin inside a floor block is a `drift`/`CA-CML` dead-reference concern; (2) **tier 2 ≠ tier 3** — a compensatory block that keeps earning its place returns, and reporting it as dead weight is wrong even when the label matches. Norwegian keywords need the Unicode boundaries in `subtraction-prefilter.mjs`; JS `\b` is ASCII-only, so `/\bunngå\b/` silently never matches.
## Testing
@ -109,10 +95,10 @@ Default: auto-detects scope from git context. Override with `/config-audit full|
node --test 'tests/**/*.test.mjs'
```
792 tests across 52 test files (15 lib + 28 scanner + 1 hook + 1 agent + 3 commands + 4 top-level). Test fixtures in `tests/fixtures/`. Top-level humanizer tests: `json-backcompat.test.mjs`, `raw-backcompat.test.mjs`, `scenario-read-test.test.mjs`, `snapshot-default-output.test.mjs`.
Test fixtures in `tests/fixtures/`. Per-scanner and per-build-block implementation notes (design rationale, primary-source verification, byte-stability lessons) live in `docs/scanner-internals.md`**Implementation notes**.
## Gotchas
- Session directories accumulate — use `/config-audit cleanup` to manage
- Scanners run on Node.js >= 18 (uses node:test, node:fs/promises)
- Scanners run on Node.js 18 (uses node:test, node:fs/promises)
- Plugin CLAUDE.md files in node_modules should be excluded via scope

255
README.md
View file

@ -6,22 +6,22 @@
*AI-generated: all code produced by Claude Code through dialog-driven development. [Full disclosure →](../../README.md#ai-generated-code-disclosure)*
![Version](https://img.shields.io/badge/version-5.1.0-blue)
![Version](https://img.shields.io/badge/version-5.13.0-blue)
![Platform](https://img.shields.io/badge/platform-Claude_Code_Plugin-purple)
![Scanners](https://img.shields.io/badge/scanners-12-cyan)
![Commands](https://img.shields.io/badge/commands-18-green)
![Agents](https://img.shields.io/badge/agents-6-orange)
![Scanners](https://img.shields.io/badge/scanners-16-cyan)
![Commands](https://img.shields.io/badge/commands-21-green)
![Agents](https://img.shields.io/badge/agents-7-orange)
![Hooks](https://img.shields.io/badge/hooks-4-red)
![Tests](https://img.shields.io/badge/tests-792+-brightgreen)
![Tests](https://img.shields.io/badge/tests-1441-brightgreen)
![License](https://img.shields.io/badge/license-MIT-lightgrey)
A Claude Code plugin that checks configuration health, suggests context-aware improvements, and auto-fixes issues — `CLAUDE.md`, `settings.json`, hooks, rules, MCP servers, `@imports`, and plugins. 12 deterministic scanners across 10 quality areas, context-aware feature recommendations, auto-fix with backup/rollback, an Opus-4.7-aware Token Hotspots scanner with optional API-calibrated `--accurate-tokens` mode, plus cache-prefix stability, dead-tool, and cross-plugin collision detection. Zero external dependencies.
A Claude Code plugin that checks configuration health, suggests context-aware improvements, and auto-fixes issues — `CLAUDE.md`, `settings.json`, hooks, rules, MCP servers, `@imports`, and plugins. 16 deterministic scanners across 10 quality areas, context-aware feature recommendations, auto-fix with backup/rollback, a prompt-cache-aware Token Hotspots scanner with optional API-calibrated `--accurate-tokens` mode, plus cache-prefix stability, dead-tool, cross-plugin collision, output-style, and always-loaded agent-listing-budget detection. Zero external dependencies.
---
## Table of Contents
- [What's New in v5.1.0](#whats-new-in-v510)
- [What's New in v5.4.0](#whats-new-in-v540)
- [What Is This?](#what-is-this)
- [The Configuration Problem](#the-configuration-problem)
- [Quick Start](#quick-start)
@ -45,56 +45,28 @@ A Claude Code plugin that checks configuration health, suggests context-aware im
---
## What's New in v5.1.0
## What's New in v5.4.0
**Plain-language UX humanizer** — every command's default output now leads with prose. Findings are grouped by what they mean for the user (Configuration mistake, Conflict, Wasted tokens, Missed opportunity, Dead config) and led with an urgency phrase (Fix this now, Fix soon, Fix when convenient, Optional cleanup, FYI). Technical IDs (`CA-CML-001`, `CA-TOK-005`, …) still appear, but at end-of-line where they belong as references rather than headlines.
**Plugin-hygiene & settings-validation hardening.** Three additive findings extend the plugin and
settings surfaces — no new scanner, so the count stays **13**:
### Before / after
- **PLH plugin-folder shadowing** (`CA-PLH-015`) — flags a `plugin.json` component-path key in the
*replaces* set (`commands`/`agents`/`outputStyles`) that points at a custom path while the
default folder of that name still exists, so the folder is silently ignored (dead config).
Mirrors Claude Code's own warning in `/doctor`, `claude plugin list`, and the `/plugin` detail
view. `skills` is excluded (it *adds to* the default scan, never shadows), as are
`hooks`/`mcpServers`/`lspServers` (own merge rules); a custom path resolving *into* the default
folder is not flagged.
- **PLH `skills:`-array validation** (`CA-PLH-016`) — validates each `plugin.json` `skills` entry
(string or array) resolves to an existing directory inside the plugin root; flags `non-string`,
`escapes-root`, `not-found`, and `not-a-directory` entries. Mirrors `claude plugin validate`.
- **SET `autoMode` structure + dead-config** — checks that `autoMode` is an object whose only keys
are `environment`/`allow`/`soft_deny`/`hard_deny`, each a string array (the literal `"$defaults"`
is valid); unknown sub-keys and wrong types are flagged (medium). Separately, `autoMode` placed
in **shared** project settings (`.claude/settings.json`) is flagged as dead config (low) —
Claude Code's classifier does not read it there.
```
v5.0.0 default
- [low] CA-CNF-001: Hook duplicate event registration
v5.1.0 default
- [low] The same automation is set up more than once
v5.1.0 with --json (machine-readable, byte-stable)
{ "id": "CA-CNF-001", "title": "...", "userImpactCategory": "Conflict",
"userActionLanguage": "Optional cleanup", "relevanceContext": "affects-everyone" }
```
### Plain-language vocabulary
The toolchain uses these terms when describing findings:
| User-facing label | What it means |
|-------------------|---------------|
| Fix this now | Something is broken or risky and should be addressed immediately |
| Fix soon | High-priority issue worth scheduling this week |
| Fix when convenient | Real issue but not urgent |
| Optional cleanup | Tidy-up that improves polish but isn't required |
| FYI | Informational; no action expected |
| Configuration mistake | A configuration file has an error or omission |
| Conflict | Two configuration sources disagree |
| Wasted tokens | Configuration is loading content that costs tokens without payback |
| Missed opportunity | A Claude Code feature you aren't using that could help your project |
| Dead config | Configuration that has no effect (e.g., a permission that's also denied) |
### Backwards compatibility — the `--raw` flag
Every CLI accepts `--raw` for byte-stable v5.0.0 verbatim output (technical IDs, raw severity, no prose translation). `--json` is unchanged from v5.0.0 — already byte-stable for programmatic consumption. Use `--raw` only if you've built tooling against v5.0.0 stderr scrapes; for new automation, prefer `--json`.
```bash
node scanners/posture.mjs . # v5.1.0 plain-language default
node scanners/posture.mjs . --raw # v5.0.0 verbatim (byte-stable)
node scanners/posture.mjs . --json # unchanged JSON envelope
```
### What's not changed
- All scanner internals (12 scanners + standalone PLH) emit the same finding IDs and structural data — humanization happens at output-formatting time only
- `--json` envelope shape is byte-stable with v5.0.0 (humanizer fields are additive on findings only in default mode; the `--json` path bypasses humanization entirely)
- 635 tests grew to 792 (+157 covering humanizer module, scenario read-tests, forbidden-words lint, JSON / `--raw` backwards-compat, default-output snapshots, and command-template / agent-prompt shape)
All three extend existing PLH and SET scanners. `--json` and `--raw` output remain byte-stable.
---
@ -104,7 +76,7 @@ Claude Code reads instructions from at least 7 different file types across multi
This plugin provides three layers of configuration intelligence:
- **Health** — 12 deterministic scanners verify correctness across every configuration file, catching broken imports, deprecated settings, conflicting rules, format errors, permission contradictions, Opus-4.7-era token waste, cache-prefix instability, dead tool grants, and cross-plugin skill collisions
- **Health** — 16 deterministic scanners verify correctness across every configuration file, catching broken imports, deprecated settings, conflicting rules, format errors, permission contradictions, prompt-cache token waste, cache-prefix instability, dead tool grants, cross-plugin skill collisions, output styles that silently strip Claude Code's coding instructions, an oversized always-loaded agent listing, and procedures in CLAUDE.md that would fit better as a skill
- **Opportunities** — context-aware recommendations for Claude Code features that could benefit your specific project, backed by Anthropic's official guidance
- **Action** — auto-fix with mandatory backups, syntax validation, rollback support, and a human-in-the-loop workflow for anything non-trivial
@ -303,9 +275,11 @@ Your team configuration changes over time. Track it:
|---------|-------------|
| `/config-audit` | Full audit with auto-scope detection (no setup needed) |
| `/config-audit posture` | Quick health scorecard: A-F grades across 10 quality areas (incl. Token Efficiency, Plugin Hygiene) |
| `/config-audit tokens` | Opus-4.7-aware token hotspots — ranked by estimated waste; 6 patterns + optional `--accurate-tokens` API calibration |
| `/config-audit manifest` | Ranked table of every system-prompt token source (CLAUDE.md, plugins, skills, MCP, hooks) sorted by estimated tokens |
| `/config-audit tokens` | prompt-cache-aware token hotspots — ranked by estimated waste, each tagged with its load pattern (always / on-demand / external); 8 patterns + optional `--accurate-tokens` API calibration. **Cache-aware:** stale `~/.claude/plugins/cache` versions (superseded installs that load on zero turns) are excluded from the ranking by default — only each plugin's active version is counted; `--no-exclude-cache` restores the full walk. Stale versions surface as a separate **Dead config** disk-cleanup finding |
| `/config-audit manifest` | Ranked table of every token source (CLAUDE.md, rules, agents, skills, output styles, MCP, hooks) sorted by estimated tokens — each tagged with its **load pattern** (always-loaded / on-demand / external) plus an **always-loaded subtotal** ("≈X tokens enter context every turn before you type"). Component-level: no coarse plugin roll-up (it would double-count) |
| `/config-audit feature-gap` | Context-aware feature recommendations grouped by impact |
| `/config-audit optimize` | Optimization lens (mechanism-fit): config that works but fits a better mechanism — procedure→skill, lifecycle→hook, unscoped path→rule, "never"→permission. Hybrid motor (deterministic pre-filter + opus precision gate), every finding cites a best-practices-register rule |
| `/config-audit optimize --subtract` | **Subtraction lens** — the inverse question no other command asks: what no longer earns its always-loaded rent? Ranks CLAUDE.md blocks that correct general model *behaviour* rather than stating a local fact, split into **dead** (never missed) and **earned** (returns if the model stumbles), with the token payoff (`BP-SUB-001`). **Load-bearing local facts are excluded deterministically before the judge sees anything** — remotes, versions, paths, filenames, policy invariants and unresolvable entity names are never candidates, and an ordered list is treated as a contract. Opt-in, proposes only, never writes. Pair with `--global` to reach the user-level CLAUDE.md, where the always-loaded cost actually sits |
| `/config-audit fix` | Auto-fix deterministic issues with backup + verification |
| `/config-audit rollback` | Restore configuration from a previous backup |
| `/config-audit plan` | Generate prioritized action plan from audit findings |
@ -319,6 +293,8 @@ Your team configuration changes over time. Track it:
| `/config-audit drift` | Compare current config against a saved baseline |
| `/config-audit plugin-health` | Audit plugin structure, frontmatter, cross-plugin coherence |
| `/config-audit whats-active` | Read-only inventory of plugins, skills, MCP, hooks, CLAUDE.md active for a repo (with token estimates) |
| `/config-audit knowledge-refresh` | Keep the best-practices register fresh — flag stale entries (sources older than ~90d) + poll for new/changed Claude Code practices; **human-approved writes only** (Verifiseringsplikt). Deterministic stale core + web candidate poll |
| `/config-audit campaign` | Machine-wide audit campaign — durable ledger above sessions tracking each repo's lifecycle (pending → audited → planned → implemented) + a machine-wide roll-up by severity + a **machine-wide always-loaded token bill** (`refresh-tokens` live cross-repo sweep — the shared global layer counted once + per-repo deltas, ranked "most expensive repos") + a single **cross-repo prioritized backlog** to pick from (severity-weighted) + plan **export** (drop a planned repo's plan into its own `docs/`), resumable across sessions; **human-approved writes only** (read-only report + deterministic write/export CLIs). Execution reuses the existing `/config-audit implement` + `rollback` |
| `/config-audit discover` | Run discovery phase only |
| `/config-audit analyze` | Run analysis phase only |
| `/config-audit interview` | Set preferences for action plan _(optional)_ |
@ -333,24 +309,104 @@ By default, `/config-audit` auto-detects scope from your git context. Override w
## Deterministic Scanners
12 Node.js scanners that perform structural analysis an LLM cannot reliably do: schema validation, circular reference detection, import resolution, conflict detection across scopes, Opus-4.7-aware token-cost analysis, cache-prefix stability, dead-tool detection, and cross-plugin skill collisions. Plus a standalone plugin-health scanner. Zero external dependencies.
15 Node.js scanners that perform structural analysis an LLM cannot reliably do: schema validation, circular reference detection, import resolution, conflict detection across scopes, prompt-cache-aware token-cost analysis, cache-prefix stability, dead-tool detection, cross-plugin skill collisions, output-style validation, and a best-practice optimization lens (mechanism-fit). Plus a standalone plugin-health scanner. Zero external dependencies.
**Why deterministic?** LLMs are powerful at understanding intent and context. But they cannot reliably validate JSON schemas, detect circular `@import` chains, or catch that your global `settings.json` contradicts your project-level one. These scanners fill that gap — fast, repeatable, and zero false positives on structural issues.
| Scanner | Prefix | What It Catches |
|---------|--------|-----------------|
| `claude-md-linter.mjs` | CML | Oversized files, missing sections, broken @imports, duplicates, stale TODOs |
| `claude-md-linter.mjs` | CML | Oversized files (line count **plus** a context-window-scaled char budget mirroring Claude Code's ~40.0k-char startup warning), missing sections, broken @imports, duplicates, stale TODOs |
| `settings-validator.mjs` | SET | Schema violations, unknown/deprecated keys, type mismatches, permission issues |
| `hook-validator.mjs` | HKV | Invalid format, missing scripts, wrong event names, timeout risks |
| `hook-validator.mjs` | HKV | Invalid format, missing scripts, wrong event names, timeout risks, verbose-stdout scripts, and a low-precision **advisory** (info) when a hook injects un-grepped command output into `hookSpecificOutput.additionalContext` — that payload enters context on every fire (plain stdout does not) |
| `rules-validator.mjs` | RUL | Bad glob patterns, orphaned rules, deprecated fields, unscoped rules |
| `mcp-config-validator.mjs` | MCP | Invalid server types, missing trust levels, exposed env vars |
| `mcp-config-validator.mjs` | MCP | Invalid server types, exposed env vars, unknown fields |
| `import-resolver.mjs` | IMP | Broken @imports, circular references, deep chains, tilde path issues |
| `conflict-detector.mjs` | CNF | Settings contradictions across scopes, permission conflicts, hook duplicates |
| `feature-gap-scanner.mjs` | GAP | 25 feature checks shown as opportunities, not grades |
| `token-hotspots.mjs` | TOK | Cache-breaking volatile content, redundant tool permissions, deep import chains, oversized cascades, bloated skill descriptions, MCP tool-schema budget |
| `cache-prefix-scanner.mjs` | CPS | Volatile content in lines 31150 of the CLAUDE.md cascade — beyond the cache-prefix window but still re-loaded every turn |
| `disabled-in-schema-scanner.mjs` | DIS | Tools listed in BOTH `permissions.deny` and `permissions.allow` — deny wins, allow entries are dead config |
| `feature-gap-scanner.mjs` | GAP | 25 feature checks shown as opportunities, not grades — plus a conditional `disableBundledSkills` recommendation when the active skill listing is over budget, and a conditional **filter-before-Claude-reads** lever when a hook injects unfiltered output into `additionalContext` (companion to the HKV advisory; cites the documented `filter-test-output.sh` pattern) |
| `token-hotspots.mjs` | TOK | Cache-breaking volatile content, redundant tool permissions, deep import chains, oversized cascades, bloated skill descriptions, MCP tool-schema budget, and stale `~/.claude/plugins/cache` versions (disk-cleanup, zero live-context impact) — cache-aware ranking excludes superseded plugin versions by default (`--no-exclude-cache` to include) |
| `cache-prefix-scanner.mjs` | CPS | Volatile content in lines 31150 of the CLAUDE.md cascade — beyond Pattern A's top-30 window but still re-loaded every turn — **plus** volatile content inside `@import`-ed files (inlined into the cached prefix, one hop, otherwise invisible to per-file scans) |
| `disabled-in-schema-scanner.mjs` | DIS | Dead/ineffective permission entries: (1) tools in BOTH `permissions.deny` and `permissions.allow` — deny wins (incl. the `Tool(*)` deny-all glob, equivalent to a bare deny); (2) unanchored allow wildcards (`*`, `B*`, `mcp__*`) that Claude Code silently skips — valid only as `mcp__<server>__*`; (3) `Tool(param:value)` rules whose key is the tool's own canonicalizing field (`command`/`file_path`/`path`/`notebook_path`/`url`) — CC ignores these and emits a startup warning |
| `collision-scanner.mjs` | COL | Cross-plugin skill name collisions; user-vs-plugin overlaps |
| `skill-listing-scanner.mjs` | SKL | Skill-listing token budget: a single skill description over the ~1,536-char listing cap Claude Code truncates (`CA-SKL-001`), the summed active-skill descriptions exceeding the ~2%-of-context listing budget (`CA-SKL-002`), and an oversized SKILL.md **body** over ~5,000 tokens (`CA-SKL-003`, low — on-demand cost: the body loads only when the skill runs, not every turn; recommends supporting-file split + `context: fork`). The `CA-SKL-002` (and CML char-budget) findings accept `--context-window <n>` to calibrate to your real window instead of the conservative 200k anchor (`--context-window auto` keeps the anchor but downgrades to advisory) |
| `output-style-scanner.mjs` | OST | Output-style validation: a custom (user/project) style missing `keep-coding-instructions: true` that silently strips built-in software-engineering instructions (`CA-OST-001`), a plugin style with `force-for-plugin: true` overriding the user's selected `outputStyle` (`CA-OST-002`), and a settings `outputStyle` resolving to no built-in or custom style — dead config (`CA-OST-003`) |
| `optimization-lens-scanner.mjs` | OPT | Optimization lens (mechanism-fit): a multi-step procedure in CLAUDE.md that would fit better as a skill (`CA-OPT-001`) — reads the machine-readable best-practices register, framed as an opportunity, not a failure. The deterministic half of the lens; prose-judgment cases (lifecycle→hook, unscoped path→rule, "never"→permission) are judged by the opus `optimization-lens-agent` via `/config-audit optimize` |
| `agent-listing-scanner.mjs` | AGT | Always-loaded agent-listing budget: a per-agent description over the soft bloat cap (`CA-AGT-001`, advisory) and the summed active-agent name+description listing — re-sent every turn — exceeding the listing budget (`CA-AGT-002`). Both LOW and explicitly **inferred / upper-bound**: the agent-listing mechanism is undocumented, so the evidence discloses the estimate and heuristic budget rather than overstating certainty |
> **Cross-scanner remediation — diagnosis meets the fix.** SKL diagnoses an over-budget
> skill listing (`CA-SKL-002`); GAP prescribes the remedy. When the active skill listing
> exceeds its ~2%-of-context budget and `disableBundledSkills` is not already set (in the
> env var or the settings cascade), the feature-gap scanner recommends that lever — hiding
> Claude Code's bundled skills (`/code-review`, `/batch`, `/debug`, `/loop`, `/claude-api`, …)
> from the model to reclaim listing budget without touching your own skills (CC 2.1.169+).
> It fires only under measured pressure, so it stays an opportunity rather than noise. Both
> scanners share one budget definition (`scanners/lib/skill-listing-budget.mjs`).
> **CLAUDE.md size — two complementary signals.** CML checks line count (200/500, for
> readability) **and** a character budget that mirrors Claude Code's own startup warning
> — *"Large CLAUDE.md will impact performance (X chars > 40.0k)."* CC 2.1.169 scales that
> threshold with the model's context window, so the char finding anchors on a conservative
> 200k window and discloses the relaxed ~200,000-char figure at 1M context. A file can be
> long by lines yet under the char budget (short lines), or short by lines yet over it — so
> both signals earn their place. The 200k/1M window constants live in the shared
> `scanners/lib/context-window.mjs` (single source of truth with the skill-listing budget).
> **Permission rules CC silently ignores — severity follows intent.** `Tool(param:value)`
> matching is real (CC 2.1.178), but the tool's own canonicalizing fields are off-limits:
> `command` (Bash/PowerShell), `file_path` (Read/Edit/Write), `path` (Grep/Glob),
> `notebook_path` (NotebookEdit), `url` (WebFetch). CC ignores a rule keyed on its tool's
> field and emits a startup warning, because `Bash(command:rm *)` is bypassable by a compound
> command. DIS splits severity by where the rule lives: in **deny/ask** it is **false
> security** (medium — the block you intended never applies), in **allow** it is **dead
> config** (low — `param:value` matching is deny/ask-only, so the entry grants nothing). The
> predicate lives in `scanners/lib/permission-rules.mjs`; valid forms like `Bash(npm:*)`,
> `WebFetch(domain:host)`, and `Agent(model:opus)` are never flagged.
> **Plugin namespace collisions — the one shadow that actually loses components.**
> Claude Code namespaces every plugin component by the plugin's declared `name`
> (`/name:command`, `name:skill`, agent `name`), so a plugin component can never shadow a
> user- or project-level one — they live in separate namespaces. The real hazard is two
> plugins that declare the **same** `name` in `plugin.json`: their namespaces collapse into
> one, and because the resolution between two installed same-name plugins is undocumented,
> one plugin's commands, skills, and agents are silently shadowed and become unreachable.
> The standalone plugin-health scanner (PLH) flags this at **medium** severity, keying on the
> declared `name` field rather than the folder name (the folder name is irrelevant to the
> namespace). A *command* name shared by two **differently-named** plugins is a milder case —
> namespacing keeps both reachable as `/a:cmd` and `/b:cmd`, so it is only ambiguity in error
> messages, search results, and the command listing. PLH reports that at **low** severity
> (group-first, one finding per command name), mirroring the COL scanner, which owns the
> analogous skill-name overlaps across *different* namespaces.
> **Plugin-folder shadowing — when a manifest path silently buries a default folder.**
> A plugin's `plugin.json` can point a component type at a custom path — `commands`,
> `agents`, and `outputStyles` all *replace* their default folder when set. So if a plugin
> declares `"commands": "./custom/"` while a `commands/` folder still exists, Claude Code
> stops scanning `commands/` entirely and everything in it silently disappears (dead config).
> PLH flags this at **medium** severity (`CA-PLH-015`), mirroring Claude Code's own warning in
> `/doctor` and `claude plugin list` (v2.1.140+). It does **not** flag `skills` — that key
> *adds to* the default `skills/` scan rather than replacing it, so both load — nor does it
> flag a custom path that points back into the default folder (e.g.
> `"commands": ["./commands/x.md"]`), because the folder is then addressed explicitly.
> **`skills:`-array validation — every listed path must be a real skill folder.**
> A plugin's `plugin.json` may list custom skill directories in a `skills` array (each entry a
> path to a folder containing `SKILL.md`). PLH validates each entry (`CA-PLH-016`, **medium**)
> and flags four ways an entry can be broken: it isn't a string, it points at a path that
> **doesn't exist**, it points at a **file** instead of a directory, or it **escapes the plugin
> root** (`../…` — installed plugins can't reference files outside their own directory, so the
> skill never loads). A valid existing directory is never flagged. This mirrors
> `claude plugin validate`. Note `skills` *adds to* the default `skills/` scan, so a custom path
> here is never a shadow — it just has to resolve to a real folder.
> **`autoMode` validation — structure and the shared-settings blind spot.** The SET scanner
> checks the auto-mode classifier config two ways. **Structure:** `autoMode` must be an object
> whose only keys are `environment`, `allow`, `soft_deny`, and `hard_deny`, each a list of
> plain-text rule strings (the literal `"$defaults"` is allowed). An unknown sub-key (e.g. a
> typo'd `hard_denies`), a non-object value, or a sub-key that isn't a string array is flagged
> **medium** — a typo'd key silently drops those rules. **Scope:** Claude Code does **not** read
> `autoMode` from *shared* project settings (`.claude/settings.json`) — "a checked-in repo cannot
> inject its own allow rules" — so an `autoMode` block committed there is dead config (**low**);
> it only takes effect in user (`~/.claude/settings.json`), local (`.claude/settings.local.json`),
> or managed settings.
### CLI Tools
@ -361,8 +417,8 @@ All tools work standalone — no Claude Code session needed:
| **Posture** | `node scanners/posture.mjs <path> [--json] [--global] [--full-machine] [--output-file path]` |
| **Fix** | `node scanners/fix-cli.mjs <path> [--apply] [--json] [--global]` |
| **Drift** | `node scanners/drift-cli.mjs <path> [--save] [--baseline name] [--json]` |
| **Tokens** | `node scanners/token-hotspots-cli.mjs <path> [--json] [--global] [--output-file path] [--accurate-tokens] [--with-telemetry-recipe]` |
| **Manifest** | `node scanners/manifest.mjs <path> [--json]` — ranked system-prompt source table |
| **Tokens** | `node scanners/token-hotspots-cli.mjs <path> [--json] [--global] [--no-exclude-cache] [--output-file path] [--accurate-tokens] [--with-telemetry-recipe]` |
| **Manifest** | `node scanners/manifest.mjs <path> [--json]` — ranked component-level source table with per-source load pattern + always-loaded subtotal |
| **What's active** | `node scanners/whats-active.mjs <path> [--json] [--verbose] [--suggest-disables]` |
| **Self-audit** | `node scanners/self-audit.mjs [--json] [--fix] [--check-readme]` |
| **Full scan** | `node scanners/scan-orchestrator.mjs <path> [--global] [--full-machine] [--no-suppress]` |
@ -381,6 +437,7 @@ Six specialized agents collaborate through the audit workflow, each matched to a
| **implementer-agent** | Sonnet | Change execution with mandatory backups | Read, Write, Edit, Bash, Glob |
| **verifier-agent** | Sonnet | Post-implementation verification | Read, Glob, Grep |
| **feature-gap-agent** | Opus | Context-aware feature recommendations | Read, Glob, Grep, Write |
| **optimization-lens-agent** | Opus | Mechanism-fit precision gate — judges prose-judgment lens candidates (lifecycle→hook, path→rule, never→permission), cites the best-practices register | Read, Glob, Grep, Write |
### Orchestration Flow
@ -474,7 +531,7 @@ node scanners/posture.mjs examples/optimal-setup/
### Self-Audit: Scanning the Scanner
The plugin runs all 12 scanners + the standalone plugin-health scanner on itself via `self-audit.mjs`. Test fixtures and example files are automatically excluded from scoring — a configuration plugin that ships deliberately broken examples shouldn't fail its own audit. Use `--check-readme` to verify badge counts are in sync with the filesystem.
The plugin runs all 16 scanners + the standalone plugin-health scanner on itself via `self-audit.mjs`. Test fixtures and example files are automatically excluded from scoring — a configuration plugin that ships deliberately broken examples shouldn't fail its own audit. Use `--check-readme` to verify badge counts are in sync with the filesystem.
```bash
node scanners/self-audit.mjs
@ -510,7 +567,7 @@ Shared modules used by all scanners — useful if you're reading the source or e
| `rollback-engine.mjs` | `listBackups()`, `restoreBackup()`, `deleteBackup()` |
| `fix-cli.mjs` | CLI entry point for auto-fix |
| `drift-cli.mjs` | CLI entry point for drift detection |
| `manifest.mjs` | CLI: ranked system-prompt source table (v5 N2) |
| `manifest.mjs` | CLI: ranked component-level source table w/ load-pattern accounting (v5 N2; v5.6 B) |
| `whats-active.mjs` | CLI: read-only active-config inventory (v3.1.0+) |
| `token-hotspots-cli.mjs` | CLI: token hotspots ranking with optional `--accurate-tokens` |
@ -523,14 +580,48 @@ Reference documents that inform the feature-gap agent and context-aware recommen
| File | Content |
|------|---------|
| `claude-code-capabilities.md` | Feature register: 18 config surfaces, Anthropic guidance, relevance table |
| `configuration-best-practices.md` | Per-layer best practices (Opus 4.7 cache-stability guidance) |
| `configuration-best-practices.md` | Per-layer best practices (cache-stability guidance) |
| `anti-patterns.md` | Common mistakes mapped to scanner IDs |
| `hook-events-reference.md` | All 26 hook events with details |
| `hook-events-reference.md` | All 28 hook events with details |
| `feature-evolution.md` | Feature timeline for staleness detection |
| `gap-closure-templates.md` | Config-specific templates for closing gaps |
| `opus-4.7-patterns.md` | Token-cost dynamics for Opus 4.7 era — patterns powering the TOK scanner |
| `prompt-cache-patterns.md` | Token-cost dynamics (prompt-cache patterns) — patterns powering the TOK scanner |
| `cache-telemetry-recipe.md` | `jq` recipe for verifying prompt-cache hit rate from session transcripts |
**Machine-readable register (`best-practices.json`).** Alongside the human-readable documents
above, `knowledge/best-practices.json` is a provenance-stamped, schema-validated register of
best-practice claims and mechanism-fit rules — each entry carries `source.url`, a `verified`
date, and a `confidence`. It is the source of truth for the optimization lens (OPT scanner +
`/config-audit optimize`); the Markdown files remain the human-readable mirror. Loaded and validated by
`scanners/lib/best-practices-register.mjs` (zero-dependency, native JSON). See
`docs/v5.7-optimization-lens-plan.md`.
### The subtraction floor
`optimize --subtract` is the only lens that proposes *removing* configuration, so it carries a
guarantee the others do not need: **a load-bearing block is never a candidate.** Precision here
is asymmetric — a missed dead line costs a few tokens per turn, while a deleted load-bearing
line costs a wrong remote or a broken script — so the floor is decided in code
(`scanners/lib/floor-exclusion.mjs`), before the opus judge sees anything, rather than being
left to prose judgement.
A block is floored when it carries an underivable local literal (inline code span, rooted path,
domain, concrete filename, version pin), when it states a policy invariant (secrets,
credentials, production, prompt-injection boundaries — floor *by decision*, not by
classification), or when it names a capitalized entity the mechanism cannot resolve without a
dictionary. That last rule is a deliberate conservative default: it declines to decide and
keeps the block, paying in recall rather than risk.
Granularity is the **leaf block** — one list item including its wrapped continuation lines, or
one paragraph — with two structural exceptions: a paragraph ending in `:` merges with the list
it introduces, and an *ordered* list is treated as a contract whose steps inherit floor from
any sibling. Unordered lists deliberately do not inherit, so a load-bearing bullet and a
disposable one can coexist in the same list. Measured against a hand-built ground truth over
a real 250-line CLAUDE.md (48 classified blocks, 19 of them genuinely ambiguous, ~65 % floor):
**zero load-bearing blocks proposed**, 11 of 18 deletable line-ranges surfaced, ≈18 % of an
always-loaded file. On a well-maintained config this axis is mostly a no-op — which is itself
the finding, and the reason precision-over-recall is the only defensible tuning.
---
## Testing
@ -539,7 +630,7 @@ Reference documents that inform the feature-gap agent and context-aware recommen
node --test 'tests/**/*.test.mjs'
```
635 tests across 36 test files (12 lib + 23 scanner + 1 hook). Test fixtures in `tests/fixtures/`. Requires Node.js 18+ (`node:test`).
1168 tests across 67 test files (22 lib + 35 scanner + 1 hook + 1 agent + 3 commands + 1 knowledge + 4 top-level). Test fixtures in `tests/fixtures/`. Requires Node.js 18+ (`node:test`).
---
@ -598,6 +689,24 @@ This plugin is cautious by design — configuration files are important, and a b
| Version | Date | Highlights |
|---------|------|-----------|
| **5.13.0** | 2026-07-31 | "Pipeline hardening" — the batch release of everything found by dogfooding the plugin against the maintainer's real machine and by walking the `analyze → plan → implement → rollback` pipeline end-to-end on a throwaway repo copy: one new lens mode plus **14 bugs** (`M-BUG-11``M-BUG-20`, `M-BUG-22``M-BUG-25`). **Added — `optimize --subtract` (`BP-SUB-001`):** the subtraction axis, asking what no longer earns its always-loaded rent. Opt-in, proposes only, and the only lens that removes config — so a **load-bearing block is never a candidate**, decided in code (`scanners/lib/floor-exclusion.mjs`) *before* the judge runs, never in prose. Verified against a hand-built ground truth written before any classifier existed: **zero load-bearing blocks proposed**, 11/18 groups, ~756 tok ≈ 18% of a ~4300-token file. **Fixed — `rollback` (`M-BUG-22/23/24/25`):** nothing agreed on where a backup lives; `listBackups()` returned 9 phantom test backups and 0 of 4 real ones, and `restoreBackup` returned `{restored:[],failed:[]}` — a **success-shaped no-op** — because `parseManifest` knew only one of the two manifest spellings in use. Canonical root now, legacy kept readable, unparseable manifests **throw**. **`M-BUG-19`:** `globToRegex` corrupted mid-pattern `/**/`, flagging live rules dead. **`M-BUG-18`/`M-BUG-20`:** the subagent harness won't write report-shaped `.md` (analyze now persists the returned report), and parallel agents clobbered the shared log with `Write` (pinned to Bash `>>`). **`M-BUG-11`/`M-BUG-13`:** `optimize` and `feature-gap` scanned vendored plugin config a user cannot act on — `optimize` candidates **454→45**, `feature-gap` **~0 (masked) → 18** opportunities. **`M-BUG-12`/`M-BUG-14`/`M-BUG-15`/`M-BUG-16`/`M-BUG-17`:** plain-language output that contradicted its own evidence (posture's `--output-file` never humanized; four finding types with no humanizer entry). Known, deliberately unfixed: `rollback` cannot delete files `implement` *created* — it now reports them (`createdNotRemoved`) instead of failing silently; automatic deletion of user files gets its own design. No count change (scanners **16**, agents **7**, commands **21**, hooks **4**); frozen v5.0.0 untouched, SC-5 regenerated once for two humanized titles. **1398** tests (+54). |
| **5.12.5** | 2026-06-26 | "Dogfood denoise" — a samle-release of the Fase-3 scanner false-positive batch (`M-BUG-2/6/7/8/10`, all dogfooding finds on the maintainer's real machine). Five scanners stop counting non-user / non-live config as the user's: **CNF** (`M-BUG-2`) excludes files under `.claude/plugins/` from conflict analysis (`isPluginBundled`) — installed plugins' bundled settings/hooks/fixtures are not a user-resolvable cascade (dogfood **339→0**, Conflicts was an F on pure plugin noise). **file-discovery** (`M-BUG-8`) adds `backups` to `SKIP_DIRS` — a `backups/` tree holds frozen copies, never live config (dogfood files-under-`/backups/` **36→0**, 717 live retained). **token estimator** (`M-BUG-6`) strips block-level `<!-- -->` HTML comments from CLAUDE.md sizing — CC strips them before injection, so they were never always-loaded tokens (dogfood ~3386→3301, ~85 tok). **CPS** (`M-BUG-7`) skips fenced/inline code and whitelists CC-stable path vars (`${CLAUDE_PLUGIN_ROOT}`/`${CLAUDE_PROJECT_DIR}`) before cache-buster matching (dogfood **5→2**). **SET** (`M-BUG-10`) typo-gates the unknown-settings-key finding — the CC schema is passthrough (verified against the 2.1.193 binary), so an unknown key is forward-compatible, not an error; it now flags only a near-miss of a known key (levenshtein ≤2), severity medium→low, +6 binary-verified `KNOWN_KEYS` (dogfood **6→0**). No count change (scanners **16**, agents **7**, commands **21**, hooks **4**); all five are byte-stable — frozen v5.0.0 + SC-5 + default-output snapshots untouched, no re-seed (each fixture's findings are genuinely unchanged). **1344** tests (+37). |
| **5.12.4** | 2026-06-26 | "Rooted rules" — fixes `M-BUG-9` (dogfooding find) in `scanners/rules-validator.mjs`: the RUL "Rule path pattern matches no files" check now resolves a rule's `paths:`/`globs:` pattern against the rule's **own project root** (the dir containing its `.claude/`), not the outer scan root. Previously `countGlobMatches` globbed against the scan target and `collectProjectFiles`' `depth>4` cutoff never reached deep matching files, so a live rule in a **nested repo** (e.g. a marketplace checkout under `~/.claude`) was wrongly flagged "never activates" (high) — a false F-grade for anyone with rules in a nested repo. The fix derives each rule's project root, collects+globs per root (cached), and skips the check for user-global rules (`root === $HOME`), which scope against the active project at runtime. Same scope-conflation family as `M-BUG-1/2`. No count change (scanners **16**, agents **7**, commands **21**); the fix is a no-op when `projectRoot === targetPath` (the common single-repo scan), so frozen v5.0.0 + default-output snapshots stay byte-stable. **1307** tests (+2 TDD: nested-repo false-positive + HOME guard). |
| **5.12.3** | 2026-06-26 | "Phantom agents" — fixes `M-BUG-3/4/5` (dogfooding finds) in `scanners/lib/active-config-reader.mjs`: `enumerateAgents` now counts only the agents Claude Code actually **registers**. Per the official subagents doc, an agent needs valid `name`+`description` frontmatter, and CC scans recursively and silently skips frontmatter-less files. The reader previously (`M-BUG-5`) counted every `.md` regardless of frontmatter, (`M-BUG-3`) never recursed into agent subdirs, and (`M-BUG-4`) double-counted when the project dir equals the user dir (scanning `$HOME` — root cause, also affecting rules/output-styles). Real-machine verify: user-agent count **13→0** (all 12 user agents + `REMEMBER.md` are frontmatter-less → CC registers none), HOME `project`-dup **13→0**; corrected always-loaded baseline ≈ **53** (was 66). Agent enumeration is machine-dependent and absent from the frozen snapshots, so the v5.0.0 + SC-5 + default-output snapshots stay byte-stable; no count change (scanners **16**, agents **7**, commands **21**). **1305** tests. |
| **5.12.2** | 2026-06-24 | "Honest census" — fixes `M-BUG-1` (dogfooding find): `enumeratePlugins` walked `~/.claude/plugins/marketplaces/<mkt>/plugins/` and ignored both enable-state and the polyrepo cache layout, so it **over-counted phantom agents** from disabled plugins while **missing the entire enabled polyrepo set** (whose plugins live under `cache/`). It now gates on `installed_plugins.json` + `enabledPlugins` and enumerates each plugin from its active `installPath`, with the marketplaces walk as fallback. Fixes `manifest`/`whats-active`/AGT/`token-hotspots` for any user with disabled plugins or a polyrepo marketplace. No count change (scanners **16**, agents **7**, commands **21**); `--json`/`--raw` byte-stable, frozen v5.0.0 + SC-5 + default-output snapshots untouched. Real-machine verify: agent listing 114→104, ghosts gone. **1301** tests. |
| **5.12.1** | 2026-06-24 | "Footgun guard" — Pattern H (stale plugin-cache versions, `token-hotspots`) recommended deleting stale version dirs with **no warning** that a currently-running session may still hold one of those versions for its whole lifetime. "Stale" is judged against `installed_plugins.json` (what NEW sessions load), so the recommendation could reproduce the exact failure that breaks a live session: deleting the dir pulls the files out from under the running session, which then must `/exit` + restart. The `plugin-cache-hygiene` recommendation now carries the live-session caveat. **Recommendation string only** — no new finding ID or scanner (count stays **16**, agents **7**, commands **21**), no token figures changed, so `--json`/`--raw` stay byte-stable and the frozen v5.0.0 + SC-5 + default-output snapshots are untouched. **1297** tests. |
| **5.12.0** | 2026-06-23 | "Auto-calibration" — completes the deferred B8 half (**B8b**): `--context-window auto` now **probes the configured model** instead of always falling back to advisory. New pure `modelToContextWindow()` maps known 1M-tier model IDs (Fable 5, Opus 4.8/4.7/4.6, Sonnet 4.6 — verified June 2026 — plus the explicit `[1m]` tier tag, dated/provider-prefixed IDs, and the `opus`/`sonnet`/`fable` aliases) to the 1M window; new IO helper `lib/active-model.mjs` `resolveActiveModel()` reads the model the way Claude Code resolves it (shell `ANTHROPIC_MODEL` override, then the settings cascade local > project > user). When `auto` resolves a recognized model the budget calibrates to its window (`auto-probed`, not advisory); when no model is pinned or it is unrecognized it keeps the conservative anchor and stays advisory (`auto-unresolved`) — the honest fallback. No new finding ID or scanner (count stays **16**, agents **7**, commands **21**); the default and explicit `--context-window` paths are unchanged, so `--json`/`--raw` stay byte-stable and the frozen v5.0.0 + SC-5 snapshots are untouched. **1296** tests. |
| **5.11.0** | 2026-06-23 | "Precision polish" — the two LOW-priority calibration gaps, both additive (scanner count stays **16**, agents **7**, commands **21**; `--json`/`--raw` byte-stable, frozen v5.0.0 + SC-5 untouched). **B7 — oversized skill body (`CA-SKL-003`, low):** the SKL scanner now measures the SKILL.md **body** (it already read the file in full) and flags bodies over ~5,000 tokens, recommending a supporting-file split + `context: fork`. Honestly framed as an **on-demand** cost — the body loads only when the skill is invoked, **not** every turn like the always-loaded listing — hence low severity. **B8 — context-window calibration (`--context-window`):** `CA-SKL-002` (skill-listing budget) and the CML char-budget now calibrate to a real context window via `--context-window <n>` (e.g. `1000000` stops the 200k anchor crying wolf on a 1M host) instead of always anchoring at 200k; `--context-window auto` keeps the conservative anchor but **downgrades budget findings to info/advisory** rather than firing a breach (model→window auto-probing deferred to a later B8b). No flag → byte-identical to the pre-B8 200k default. CPS is intentionally excluded (no window-anchored budget to calibrate). 1279 tests |
| **5.10.0** | 2026-06-23 | "Deferral & injection hygiene" — three additive hardening levers that extend existing scanners toward a tighter always-loaded prefix (scanner count stays **16**, agents **7**, commands **21**; `--json`/`--raw` byte-stable, frozen v5.0.0 + SC-5 untouched). **B4 — MCP tool-schema deferral (`CA-TOK-006`; tokens patterns 7→8):** Claude Code defers MCP tool schemas (names-only, ~120 tok; full schemas load on demand) by default, so `CA-TOK-006` detects config-file signals that force the FULL schemas into the always-loaded prefix every turn — `env.ENABLE_TOOL_SEARCH="false"` (high), a `"ToolSearch"` deny (high), a configured Haiku model (medium), or per-server `alwaysLoad:true` (CC v2.1.121+, high); severity scales with the aggregate forced-upfront tokens. New pure engine `lib/mcp-deferral.mjs` shared by TOK + GAP, plus a feature-gap **CLI-over-MCP** companion lever (prefer `gh`/`aws`/`gcloud`). Triggers on config files ONLY — Vertex / custom `ANTHROPIC_BASE_URL` / a runtime `/model` switch are launch state and are disclosed, never triggered; the prefix-cache-invalidation claim was NOT-CONFIRMED in docs and is not asserted. **B5 — hook `additionalContext` advisory + filter-before lever:** HKV emits an info advisory when a hook injects unfiltered output into `additionalContext`, with a feature-gap **filter-before-Claude-reads** companion citing the documented `filter-test-output.sh` pattern. **B6 — CPS `@import` volatile scan:** the cache-prefix scanner now follows `@import`s (one hop) and flags volatile content in the imported file that breaks the cached prefix — a new medium finding, keyed on the resolved file. 1257 tests |
| **5.9.0** | 2026-06-23 | "Machine-wide token lens" — the three highest-impact hardening gaps toward whole-machine token tuning. **B1 — agent-listing budget (new orchestrated scanner AGT, count 15→16):** the always-loaded agent listing (name+description re-sent every turn) is now measured — `CA-AGT-001` per-agent description bloat (advisory), `CA-AGT-002` aggregate listing over budget; both LOW and explicitly **inferred / upper-bound** (the mechanism is undocumented — the evidence discloses it rather than overstating). **B2 — machine-wide always-loaded token roll-up:** the campaign ledger now carries a token bill — `campaign refresh-tokens` does a live cross-repo sweep that counts the **shared global always-loaded layer once** + per-repo deltas, with a ranked "most expensive repos" table (the `whats-active` double-count, avoided by construction). **B3 — cache-aware filtering (folds in B0):** `~/.claude/plugins/cache` holds *both* active and stale plugin versions (installPaths point INTO it), so token-hotspots + CNF are now **version-aware**`--exclude-cache` (default ON) keeps each plugin's active version and drops only stale ones (`installed_plugins.json`-driven), so stale versions stop polluting the hotspot ranking and inflating duplicate-hook conflicts; stale versions surface as a separate **Dead config** disk-cleanup finding (zero live-context impact). `--json`/`--raw` byte-stable; frozen v5.0.0 + SC-5 snapshots untouched. 1215 tests |
| **5.8.0** | 2026-06-23 | "Campaign motor" — a durable, machine-wide audit **campaign** that sits ABOVE individual sessions (one repo = one session; a fleet of repos = a campaign). **Ledger:** `~/.claude/config-audit/campaign-ledger.json` (outside the plugin dir → survives uninstall/upgrade) tracks a repo list + per-repo lifecycle (pending→audited→planned→implemented) + a machine-wide roll-up by status & severity; pure transforms with injected `now`. **`/config-audit campaign` (commands 20→21):** read-only report (`campaign-cli`) + human-approved writes (`campaign-write-cli`: init / add / set-status) — reports first, mutates only on explicit approval, never hand-edits the ledger. **Cross-repo backlog:** one severity-weighted prioritized pick-list (`buildBacklog`, `critical:1000/high:100/medium:10/low:1`). **Plan export + execution-by-reuse:** `campaign-export-cli --write` drops a planned repo's plan verbatim into its own `docs/`; execution reuses the existing `/config-audit implement` + `rollback` (no new execution machinery). All campaign code is `-cli`/lib → scanner count stays **15**, agents **7**, byte-stable. Plus pre-release cleanup: `knowledge-refresh` wired into the router + help; CLAUDE.md trimmed 540→134 lines (impl notes → `docs/scanner-internals.md`, config grade B→A). 1168 tests |
| **5.7.0** | 2026-06-21 | "Optimization lens" — first detector of the «optimally shaped?» axis (vs «correct?»), plus a living knowledge layer. **Register:** `knowledge/best-practices.json`, a provenance-stamped, schema-validated best-practices register (first runtime-consumed `knowledge/` file). **OPT scanner (count 14→15):** `CA-OPT-001` (LOW) a ≥6-step CLAUDE.md procedure that would fit better as a skill, citing register entry `BP-MECH-003`. **`/config-audit optimize` + `optimization-lens-agent` (opus, agents 6→7):** prose-judgment lens for lifecycle→hook (`BP-MECH-001`), unscoped path→rule (`BP-MECH-002`), "never"→permission (`BP-MECH-004`); pre-filter recall + opus precision gate. **`/config-audit knowledge-refresh` (commands 19→20):** deterministic stale-check (injected reference date, 90-day cadence) + web re-verify/poll, human-approved writes only. Last two are agent/web-driven (not byte-stable). 1091 tests |
| **5.6.0** | 2026-06-20 | "Steering-model II" — the load-pattern / compaction-survival model lands end-to-end. **Foundation:** `active-config-reader` now enumerates rules, agents, and output styles (alongside CLAUDE.md/plugins/skills/hooks/MCP), each tagged `loadPattern` (always / on-demand / external) + `survivesCompaction` from the published loading model; the frontmatter parser also reads YAML block sequences (`paths:` lists). **B (load-pattern accounting):** `manifest` reports component-level sources (the double-counting plugin roll-up is gone), tags every source with the load-pattern triple, and leads with an **always-loaded subtotal** ("tokens that enter context every turn"); `token-hotspots` annotates each ranked hotspot with its load pattern. **C (output styles):** new orchestrated **OST** scanner (count 13→**14**) — `CA-OST-001` a custom style stripping built-in coding instructions (missing `keep-coding-instructions: true`, V10), `CA-OST-002` a plugin style with `force-for-plugin: true` overriding the user's `outputStyle` (V11), `CA-OST-003` a settings `outputStyle` resolving to no known style (dead config). Doc-verified; frozen v5.0.0 snapshots preserved via strip-helpers, SC-5 regenerated. 1023 tests |
| **5.5.0** | 2026-06-20 | "Steering-model I" — two additive compaction-durability / dead-config findings (count stays **13**, `--json`/`--raw` byte-stable). Per the official "what survives compaction" model: RUL flags a large (>50-line) **path-scoped** rule not re-injected after compaction (LOW); CML flags a **nested** (subdir) CLAUDE.md not re-injected after compaction (LOW). PLH flags a plugin agent setting `hooks`/`mcpServers`/`permissionMode` — Claude Code ignores these for plugin subagents, so it's dead config (`permissionMode` = MEDIUM false-security, `hooks`/`mcpServers` = LOW). Known limitation: the frontmatter parser reads inline `paths:` but not YAML block sequences (deferred to v5.6 Foundation). 961 tests |
| **5.4.1** | 2026-06-20 | Scanner-correctness patch (count stays **13**, `--json`/`--raw` byte-stable). HKV: added `Setup`/`UserPromptExpansion`/`PostToolBatch` to the valid-event set (a valid hook using one was wrongly flagged "will never fire"), and **removed** `post-session` (the 2.1.169 `post-session` is a self-hosted-runner workspace-lifecycle hook, **not** a settings.json event — absent from `hooks.md`; verified 2026-06-20). RUL: globs-rule wording corrected — only `paths:` is documented, so the finding drops the unverified "deprecated/legacy" claim and steers to the documented field. PLH: optional `model`/`tools`/`name`/`allowed-tools` frontmatter no longer required; CLAUDE.md component-section required only for components the plugin actually ships. 954 tests |
| **5.4.0** | 2026-06-19 | Plugin-hygiene & settings-validation hardening. Three additive findings extend existing PLH and SET scanners (count stays **13**): PLH plugin-folder shadowing (`CA-PLH-015` — a `plugin.json` component-path key in the *replaces* set `commands`/`agents`/`outputStyles` pointing at a custom path while the default folder still exists) mirroring CC's `/doctor` & `claude plugin list` warning; PLH `skills:`-array validation (`CA-PLH-016` — each entry must resolve to a directory in the plugin root; flags `non-string`/`escapes-root`/`not-found`/`not-a-directory`) mirroring `claude plugin validate`; SET `autoMode` structure (only `environment`/`allow`/`soft_deny`/`hard_deny` string arrays) + dead-config (`autoMode` in shared `.claude/settings.json` is not read by CC). `--json`/`--raw` byte-stable. 949 tests |
| **5.3.0** | 2026-06-19 | Permission-rule & plugin-hygiene hardening. Five additive scanner findings extend existing scanners (count stays **13**): DIS forbidden-param rules (`Tool(param:value)` on a canonicalizing field — deny/ask = false security, allow = dead config) and ineffective allow-wildcards + `Tool(*)` deny-all; CML context-window-scaled 40.0k-char CLAUDE.md budget mirroring CC's startup warning; PLH plugin namespace collision (two plugins declaring the same `name`); feature-gap `disableBundledSkills` lever under skill-listing pressure. PLH cross-plugin command-name overlap reframed HIGH → LOW (namespacing keeps both reachable). `--json`/`--raw` byte-stable. 936 tests |
| **5.2.0** | 2026-06-18 | CC 2.1.114→181 compatibility + skill-listing budget. New orchestrated scanner **SKL** (`CA-SKL-001` 1,536-char listing cap, `CA-SKL-002` listing-budget sum) → 13 orchestrated scanners. Five validators refreshed for CC 2.1.114181 settings/hook surface (`xhigh` effort, `MessageDisplay` + post-session events, 28 hook events). False positives eliminated in MCP (auto-injected/POSIX env vars, invented `trust` field) and permissions (param-aware DIS/CNF). Hermetic HOME isolation across all CLI-spawning tests. 875 tests |
| **5.1.0** | 2026-05-01 | Plain-language UX humanizer. Default output of all 18 commands now leads with prose; findings grouped by user-impact category (Configuration mistake, Conflict, Wasted tokens, Missed opportunity, Dead config) and led by urgency phrase (Fix this now → FYI). New `--raw` flag preserves v5.0.0 verbatim output for tooling that scrapes stderr; `--json` is unchanged and byte-stable. New scanner-lib modules: `humanizer.mjs`, `humanizer-data.mjs` with TRANSLATIONS for 13 scanner prefixes. Self-audit terminal output also humanized. 792 tests (+157 humanizer-tester) |
| **5.0.0** | 2026-05-01 | Reality-based token-optimization. 3 new scanners (CPS cache-prefix, DIS dead tools, COL plugin collisions) → 12 deterministic scanners. New `/config-audit manifest` and `--accurate-tokens` API calibration. Severity-weighted scoring (`scoringVersion: 'v5'`). MCP token estimates 15 → 500+. Plugin Hygiene as 10th quality area. Knowledge: cache-stability replaces 200-line rule, cache-telemetry recipe. **Breaking:** F2 token magnitude jump, F3 severity weighting, F5 Pattern D removed, N1 `CA-TOK-*` glob now matches CA-TOK-005. 635 tests |
| **4.0.0** | 2026-04-19 | Opus 4.7 era: new TOK scanner (cache-breaking volatile content, redundant tool permissions, deep import chains, sonnet-era setups), `/config-audit tokens` command, Token Efficiency 8th quality area, scanner-agent + verifier-agent migrated haiku → sonnet. 543 tests |

View file

@ -51,11 +51,16 @@ In `--raw` mode, fall back to v5.0.0 severity prefiks and verbatim scanner title
5. **Identify optimizations**: Rules to globalize, missing configs, orphaned files
6. **Security scan**: Aggregate secret warnings, check for insecure patterns
7. **CLAUDE.md quality assessment**: Score each file against rubric, assign letter grades
8. **Generate report**: Write comprehensive markdown report — group findings by `userImpactCategory`, lead with `userActionLanguage`
8. **Generate report**: Compose the comprehensive markdown report — group findings by `userImpactCategory`, lead with `userActionLanguage`
## Output
Write to: `~/.claude/config-audit/sessions/{session-id}/analysis-report.md`
Return the complete report as your final message — do not write it to a file
yourself. The Claude Code subagent harness instructs agents not to write
report/analysis files; your text output IS the deliverable. The orchestrating
command saves your returned report verbatim to
`~/.claude/config-audit/sessions/{session-id}/analysis-report.md` for the
downstream plan/interview/status phases.
**Output MUST NOT exceed 300 lines.** Prioritize findings by severity. Use tables, not prose.
@ -183,4 +188,4 @@ Verify report: all findings referenced, recommendations actionable, severity lev
- Process findings in memory (typically < 1MB total)
- Generate report in single pass
- No file modifications (read-only except report output)
- No file modifications (read-only; the report is returned as your final message)

View file

@ -146,6 +146,11 @@ Move content from one file to another.
Append to: `~/.claude/config-audit/sessions/{session-id}/implementation-log.md`
**Append discipline (shared log):** other implementer agents may be writing this
log concurrently. ALWAYS append your entry with a Bash `>>` heredoc;
NEVER use the Write or Edit tool on the log file — a full-file Write silently
clobbers entries other agents appended after you read the file.
### Success
```markdown

View file

@ -0,0 +1,166 @@
---
name: optimization-lens-agent
description: |
Judges CLAUDE.md mechanism-fit for the v5.7 optimization lens (CA-OPT). Reads
deterministic pre-filter candidates and decides, with prose judgement, whether
each is a genuine "you use mechanism X, but Y fits this better" opportunity —
lifecycle phrasing → hook, unscoped path-specific instruction → path-scoped
rule, absolute "never" prohibition → permission. Precision-gated: cites the
best-practices register rule + source, and stays silent when unsure.
model: opus
color: orange
tools: ["Read", "Glob", "Grep", "Write"]
---
# Optimization Lens Agent
You are the **precision gate** of the optimization lens's hybrid motor. A cheap
deterministic pre-filter (`lens-prefilter`) has already surfaced candidate lines
in CLAUDE.md that *might* fit a better mechanism. Your job is to read each
candidate **in its real context** and keep only the genuine opportunities.
This is the "is the config **optimal?**" axis, not "is it **correct?**" — every
finding is a *Missed opportunity*, never a mistake. The config works as written;
you are pointing at a mechanism that would fit the content better.
## The judgement you make
For each candidate, the register rule names the better-fit mechanism. Decide
whether the line is *really* that kind of instruction:
| lensCheck | Register | Keep it ONLY if the line is… | Better mechanism |
|---|---|---|---|
| `claude-md-lifecycle-phrasing` | BP-MECH-001 | a recurring automation the model is *told* to perform ("after every commit, run X") — something that should happen deterministically, not at the model's discretion | a **hook** (PreToolUse / PostToolUse / Stop) |
| `unscoped-path-specific-instruction` | BP-MECH-002 | a constraint that only applies when a *specific* file/path/glob is touched, sitting in root CLAUDE.md where it loads every turn regardless | a **path-scoped rule** (`.claude/rules/` with `paths:` frontmatter) |
| `never-instruction` | BP-MECH-004 | an *absolute* prohibition — something that must NEVER happen, where relying on the model to remember is the wrong guarantee | a **permission deny rule** or PreToolUse hook |
| `compensatory-instruction` | BP-SUB-001 | **`--subtract` mode only.** an instruction that corrects general model *behaviour* rather than stating a local fact — so it pays an always-loaded token cost without telling the model anything it could not work out | **removal**, re-added only if the model actually stumbles |
## The subtraction lens (`--subtract` only)
Present only when the payload has a `subtract` block. It asks the inverse of
every other lens: *what is no longer earning its always-loaded rent?* Three
things make it different, and all three are non-negotiable.
**1. The floor is not yours to decide.** A deterministic pre-step has already
excluded every block carrying a local fact — a code span, path, domain, version
pin, policy invariant, or an unresolved capitalized entity — plus the steps of
any ordered list whose siblings carry one. You never see those blocks, and you
must not reason about whether some *other* block ought to be deleted. Judge only
what you are given. Precision is asymmetric: a missed dead line costs a few
tokens per turn; a deleted load-bearing line costs a wrong remote, a broken
script, or a lost afternoon.
**2. Staleness is NOT a deletion signal.** A block that pins an outdated version
("use Opus 4.8") is a *dead-reference* problem for `drift` / `CA-CML`, not a
subtraction finding. The instruction is still load-bearing — it encodes a
decision only the operator can make; it is merely out of date. Recommending
deletion because content looks stale is a category error. Say "this looks
outdated" if you must, but never as a removal candidate.
**3. Tier 2 is not tier 3.** Deletable splits into *earned* (compensatory, but
this model still stumbles on it, so it returns) and *dead* (never missed). Sort
every candidate into one of the two and say which. A block that has visibly
earned its place — its subject matter recurs in the repo's own history — is tier
2 even when its classification is "compensatory". Reporting it as dead weight is
wrong even though the label matches.
Rank kept candidates by always-loaded token cost, and state the total payoff.
Frame it as *rent*, never as a mistake: this config was correct when written.
## Input
You receive an `optimize-lens` payload (JSON) with:
- `target` — the repo path.
- `deterministic` — OPT scanner findings already confirmed (CA-OPT-001:
procedure → skill). Report these **as-is**; do not re-judge them.
- `candidates` — pre-filter candidates, each with `file`, `line`, `lensCheck`,
`mechanism`, `signalText`, and a `register` block (`id`, `claim`,
`recommendation`, `severity`, `source`). Only CONFIRMED register rules reach
you.
- `register` — the full confirmed prose-judgment entries, for reference.
- `subtract`**present only under `--subtract`.** `{ enabled, candidates,
register, detectors }`. Each candidate spans `line``endLine` (a whole leaf
block, not one line) and carries `signalText` plus the BP-SUB-001 register
block. Everything load-bearing was already removed before you saw this.
Always **Read the actual CLAUDE.md file(s)** named in the candidates before
judging — `signalText` is one line out of context; the surrounding lines decide
whether it is really lifecycle/path-specific/prohibition phrasing.
## Precision rules (non-negotiable)
1. **Keep only high-confidence opportunities.** When the line is ambiguous,
rhetorical, an example, a heading, or already correctly placed (e.g. it is
*inside* a path-scoped rule, or already references a hook) — **drop it**. A
missed suggestion is far cheaper than a wrong one (Verifiseringsplikt).
2. **Never invent a recommendation.** Use the `register.recommendation` and cite
`register.id` + `register.source.url`. If a candidate has no register block,
skip it.
3. **De-duplicate.** If one line yields two candidates (e.g. "never edit
src/config.ts"), pick the single mechanism that fits best and say why,
rather than emitting two findings for one line.
4. **No false urgency.** These are LOW-severity opportunities. Do not imply the
config is broken.
## Output
Write `optimization-lens-report.md` to the session directory (≤120 lines).
```markdown
# Optimization Lens — mechanism-fit
**Date:** YYYY-MM-DD | **Target:** {repo}
**Confirmed opportunities:** {N kept} · **Candidates reviewed:** {M} · **Dropped (low confidence):** {M-N}
> The config works as written. These are places where a different Claude Code
> mechanism would fit the content better — usually cheaper per turn or more
> reliable.
## Procedures → skills (deterministic)
{For each `deterministic` finding — render title/recommendation verbatim, cite CA-OPT-001 + BP-MECH-003.}
## Lifecycle → hooks
{Kept BP-MECH-001 findings. For each:}
**{file}:{line}** — {one-line restatement of the line}
Why: {register.claim, condensed}
Move to: {register.recommendation}
Source: {register.source.url}
## Path-specific → scoped rules
{Kept BP-MECH-002 findings, same shape.}
## Absolute prohibitions → permissions
{Kept BP-MECH-004 findings, same shape.}
## What I deliberately left alone
{Brief, honest: candidates you dropped and why — "line 22 mentions a path but is
a cross-reference, not an instruction." This is the precision gate showing its
work. Keep to a few lines.}
## No longer earning its rent (--subtract only)
{Omit entirely unless the payload has a `subtract` block. Two sub-lists —
**Dead** (tier 3, out and never missed) and **Earned** (tier 2, out but likely
to return) — ranked by token cost, with a payoff total. For each:}
**{file}:{line}-{endLine}** — {what the block says, in one line} · ~{N} tok/turn
Tier: {dead | earned — and why}
Source: {register.source.url}
```
Omit any section with zero kept findings (except keep the "left alone" note when
you dropped anything). If nothing survived the gate, say so plainly — a clean
CLAUDE.md is a good outcome, not a failure to find problems.
## Guidelines
- Frame everything as *opportunities*, never failures.
- Cite the register rule id + source URL on every finding — provenance is the
product.
- Be concrete: name the file and line, and what the replacement mechanism is.
- Prefer dropping a borderline candidate over stretching to keep it.
- Do not recommend a mechanism the project already uses for that exact content.

View file

@ -171,18 +171,10 @@ Total backup size: ~6.4 KB
**Rationale**:
Code style rules found in 3 projects are identical. Moving to global reduces duplication.
**Content**:
```markdown
# Code Style Rules
## Language Preferences
- TypeScript > JavaScript
- Explicit > implicit
- Lesbarhet > cleverness
## Commit Format
- Conventional Commits: `type(scope): description`
```
**Content outline** (describe it — do not inline the file):
Language preferences, then commit format. The implementer reads the source
files and writes the content itself; a full file body pasted here is what the
200-line budget above forbids.
**Validation**:
- File exists after creation

View file

@ -20,9 +20,16 @@ Generate comprehensive analysis report from discovery findings.
## Implementation
### Step 1: Verify session state
### Step 1: Resolve the session and verify its state
Read `~/.claude/config-audit/sessions/{session-id}/state.yaml` using the Read tool and verify discovery phase completed. If not, tell the user: "Discovery hasn't been run yet. Start with `/config-audit discover` or just run `/config-audit` for a full audit."
Find the session first — never guess which one `{session-id}` refers to:
```
Glob: ~/.claude/config-audit/sessions/*/state.yaml
Sort by modification time — the most recently modified session wins
```
Every `{session-id}` below is that session's id. Read its `state.yaml` using the Read tool and verify discovery phase completed. If the Glob returns nothing, or discovery hasn't completed, tell the user: "Discovery hasn't been run yet. Start with `/config-audit discover` or just run `/config-audit` for a full audit."
### Step 2: Tell the user what's happening
@ -37,17 +44,16 @@ This includes hierarchy mapping, conflict detection, and prioritized recommendat
Tell the user: **"Generating analysis (this takes about 30 seconds)..."**
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
```
Check whether `$ARGUMENTS` contains `--raw`. Carry the answer yourself: the agent
prompt below is **not** a shell, so a variable assigned in a bash block cannot be
referenced from it. Substitute `{mode}` literally with `--raw` or `humanized`.
```
Agent(subagent_type: "config-audit:analyzer-agent")
model: sonnet
prompt: |
Analyze all findings in: ~/.claude/config-audit/sessions/{session-id}/findings/
Mode: $RAW_FLAG (empty = humanized; "--raw" = v5.0.0 verbatim severity prefiks)
Mode: {mode} ("humanized" = humanized; "--raw" = v5.0.0 verbatim severity prefiks)
Generate comprehensive report covering:
1. Executive summary with key metrics, grouped by userImpactCategory
2. Hierarchy map visualization
@ -60,12 +66,21 @@ Agent(subagent_type: "config-audit:analyzer-agent")
raw severity. The humanizer already replaced jargon-heavy
title/description/recommendation strings with plain-language
equivalents — render them verbatim, do not paraphrase.
Output to: ~/.claude/config-audit/sessions/{session-id}/analysis-report.md
Return the complete report as your final message. Do not write it
to a file — the orchestrating command saves it to the session directory.
```
### Step 4: Present summary
### Step 4: Save the report
After the agent completes, read the generated report and show a brief summary:
The agent returns the complete report as its final message — the Claude Code
subagent harness instructs agents not to write report/analysis files themselves,
so the command must persist it. Write the returned report verbatim (no edits,
no truncation) to `~/.claude/config-audit/sessions/{session-id}/analysis-report.md`
using the Write tool. Downstream phases (`plan`, `interview`, `status`) read this file.
### Step 5: Present summary
After saving the report, show a brief summary:
```markdown
### Analysis Complete
@ -84,6 +99,6 @@ Full report: `~/.claude/config-audit/sessions/{session-id}/analysis-report.md`
- **`/config-audit fix`** — Auto-fix deterministic issues right away
```
### Step 5: Update state
### Step 6: Update state
Update `state.yaml` with `current_phase: "analyze"`, `next_phase: "plan"`.
Update `state.yaml` with all four fields `.claude/rules/state-management.md` requires: `current_phase: "analyze"`, `completed_phases` (append `analyze` to the existing array — read it first), `next_phase: "plan"`, and `updated_at`. A write that names only two of the four silently deletes the other two.

332
commands/campaign.md Normal file
View file

@ -0,0 +1,332 @@
---
name: config-audit:campaign
description: Machine-wide audit campaign — track which repos are pending/audited/planned/implemented across sessions, with a machine-wide roll-up. Human-approved writes only.
argument-hint: "[init | add <path>... | set-status <path> <status> | refresh-tokens | export <path>]"
allowed-tools: Read, Write, Edit, Bash, Glob
model: opus
---
# Config-Audit: Campaign
A single config-audit session audits **one** scope. A **campaign** sits above sessions: a
durable ledger of every repo you mean to bring up to standard, each repo's lifecycle status
(**pending → audited → planned → implemented**), and a machine-wide roll-up of findings by
severity. It persists to `~/.claude/config-audit/campaign-ledger.json`**outside** the
plugin dir, next to `sessions/` — so it survives plugin uninstall/reinstall/upgrade and
resumes across sessions.
**The Iron rule (Verifiseringsplikt): nothing is ever auto-written.** Reporting is read-only.
Every mutation — creating the ledger, adding a repo, changing a status — is proposed first
and applied **only on explicit approval**, by invoking one deterministic write-CLI subcommand.
The command never hand-edits the ledger JSON.
This is the **THIN** campaign surface (ledger + roll-up + status + a cross-repo prioritized
backlog to pick from + plan **export**). It does not run audits or apply fixes itself: it tracks
where each repo stands, shows what to tackle next, exports a planned repo's plan into that repo's
own `docs/`, and points at the **existing** per-repo `/config-audit implement` (backup + apply +
verify) and `/config-audit rollback` for execution — Block 4c reuses that machinery, it does not
reinvent it.
## Three CLIs back this command
- **Read (report):** `scanners/campaign-cli.mjs` — loads + validates the ledger, emits the
repo list + roll-up + backlog. Never writes.
- **Write (mutate):** `scanners/campaign-write-cli.mjs``init` / `add` / `set-status` /
`refresh-tokens`, each a thin wrapper over the invariant-enforcing lib transforms + save.
Invoked **only** after the user approves a specific action. `refresh-tokens` is the live
cross-repo token sweep: it runs the manifest's always-loaded accounting across every tracked
repo and folds the result into the machine-wide token bill (shared global layer counted once
+ per-repo deltas).
- **Export:** `scanners/campaign-export-cli.mjs``--repo <path>` resolves the repo's linked
session, reads its `action-plan.md`, and assembles a `docs/config-audit-plan-<sessionId>.md`.
Read-only (a preview) by default; it writes the file **only** under `--write`, which is
invoked **only** after the user approves. The CLI writes the file byte-faithfully — the plan is
never re-typed.
All take `--ledger-file <path>` (defaults to the durable path) and `--output-file <path>`; the
write-CLI + export-CLI also take `--reference-date <YYYY-MM-DD>` (the date stamp), and the
export-CLI takes `--sessions-dir <path>` (defaults to `~/.claude/config-audit/sessions`).
## Implementation
### Step 1: Parse arguments
From `$ARGUMENTS`, pick the mode:
- *(empty)* or `report`**report** (read-only). Default.
- `init` → initialize the ledger.
- `add <path>...` → add one or more repo paths.
- `add --discover <root>` → find git repos under `<root>` and let the user pick which to add.
- `set-status <path> <status>` → transition a tracked repo (`status` ∈ pending/audited/planned/implemented).
- `refresh-tokens` → live cross-repo token sweep: compute the machine-wide always-loaded bill.
- `export <path>` → export a planned repo's action plan into that repo's own `docs/`.
- `help` → show this surface and stop.
Every write step derives its own date stamp inside its own block — there is no shared one to
set here, because each fenced block runs as a separate process.
### Step 2: Always report current state first
Whatever the mode, start by showing where the campaign stands (read-only):
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/campaign-cli.mjs \
--output-file ~/.claude/config-audit/sessions/campaign-report.json 2>/dev/null; echo $?
```
Exit **0** = a campaign exists, **1** = not initialized yet (advisory — normal first run),
**3** = real error → "The campaign ledger couldn't be read — it may be corrupt." (Stop; do not
attempt a write over a corrupt ledger.)
Read `~/.claude/config-audit/sessions/campaign-report.json` with the Read tool (per the UX
rules — never show the raw JSON). It has `initialized`, `repos[]` (each: `path, name, status,
sessionId, findingsBySeverity, tokens, updatedDate`), `rollUp {totalRepos, byStatus, bySeverity,
reposWithFindings, tokens}`, and `backlog[]` — the single cross-repo prioritized work list (each:
`path, name, status, findingsBySeverity, totalFindings, weightedScore, rank`), already sorted
DESC by severity (most critical work first).
Present it as two short tables:
**Campaign roll-up**
| Status | Repos |
|--------|-------|
| pending / audited / planned / implemented | … |
…plus a one-line severity total across audited repos (e.g. "Findings so far: 3 critical,
8 high, 5 medium, 12 low across 4 audited repos").
**Repos**
| Repo | Status | Findings (C/H/M/L) | Last updated |
|------|--------|--------------------|--------------|
**Prioritized backlog** — the one cross-repo list to pick from, highest-severity work first.
Render `backlog[]` in `rank` order (it is already sorted); omit this table entirely when the
backlog is empty (nothing outstanding — say "Backlog clear — no outstanding findings across
tracked repos."). Implemented repos and repos with no known findings are deliberately absent.
| # | Repo | Status | Findings (C/H/M/L) | Total |
|---|------|--------|--------------------|-------|
After it, point the user at the top item: "Highest priority: **`<name>`** (`<status>`) — pick it
with `/config-audit` (audit), `/config-audit plan`, or `/config-audit implement` in that repo,
then record progress here with `set-status`." The backlog is a **pick-list**, not an executor —
this command does not run audits or fixes (that is the later execution block).
**Machine-wide token bill** — render from `rollUp.tokens` (the whole-machine always-loaded
accounting). Note the shape: `sharedGlobal`, `perRepoDelta`, and `machineWide` are flat
`{always, onDemand, external, unknown}` number maps; `byRepo[]` is `{name, path, always,
onDemand, external}` already sorted DESC by always-loaded cost; `reposWithTokens` is the count.
If `reposWithTokens` is `0`, no sweep has run yet — say: "No token bill yet — run
`/config-audit campaign refresh-tokens` to compute the machine-wide always-loaded cost." and
omit the table. Otherwise lead with the headline and the once-vs-delta split:
> **Always-loaded every turn, machine-wide: ~`machineWide.always` tokens**`sharedGlobal.always`
> paid once (global config + installed plugins, in *every* repo) + `perRepoDelta.always` across
> `reposWithTokens` repos' own project config.
Then the **most expensive repos** (their per-repo delta — what each adds beyond the shared layer):
| # | Repo | Always-loaded delta |
|---|------|---------------------|
| 1 | `<byRepo[0].name>` | `<byRepo[0].always>` |
Add one plain-language line so the number is actionable, e.g. "The shared global layer is the
biggest lever — trim `~/.claude/CLAUDE.md`, the global agent listing, or rarely-used plugins
to cut cost in every repo at once." The bill reflects the **last** sweep; re-run
`refresh-tokens` after config changes.
If `initialized` is false, say so plainly: "No campaign yet. Run `/config-audit campaign init`
to start one." Then — if the mode was `init` or `add` — continue to that step (those bootstrap
a campaign); for `report`/`set-status` on an uninitialized ledger, stop after this message.
### Step 3 (mode `init`): Initialize
If already initialized, say so and stop (no clobber). Otherwise tell the user what will happen,
then create it:
```bash
# Re-derive here: each fenced block is its own Bash call, so a TODAY set
# in an earlier block is empty by the time this one runs.
TODAY=$(date +%F)
node ${CLAUDE_PLUGIN_ROOT}/scanners/campaign-write-cli.mjs init \
--reference-date "$TODAY" \
--output-file ~/.claude/config-audit/sessions/campaign-write.json 2>/dev/null; echo $?
```
Exit **0** = created, **1** = already initialized (advisory). Confirm: "Campaign ledger created
at `~/.claude/config-audit/campaign-ledger.json`." Then suggest `add`.
### Step 4 (mode `add`): Add repos — propose, approve, write
**Gather candidates.**
- Explicit paths: use the paths given after `add`.
- `--discover <root>`: find git repos (depth-limited), e.g.
```bash
find "<root>" -maxdepth 3 -type d -name .git 2>/dev/null | sed 's:/\.git$::'
```
Present the discovered repos as a numbered list and ask **which** to add (and confirm any
that are already tracked will be skipped). Use Glob as a fallback if `find` is unavailable.
**Confirm, then write.** Show the final list and ask for explicit approval. On approval, add
them in one call (idempotent — already-tracked repos are skipped, not reset):
```bash
# Re-derive here: each fenced block is its own Bash call, so a TODAY set
# in an earlier block is empty by the time this one runs.
TODAY=$(date +%F)
node ${CLAUDE_PLUGIN_ROOT}/scanners/campaign-write-cli.mjs add <path1> <path2> ... \
--reference-date "$TODAY" \
--output-file ~/.claude/config-audit/sessions/campaign-write.json 2>/dev/null; echo $?
```
(For a single repo with a custom display name, add `--name "<name>"`.) Read the result file and
report what was `added`, `addedUnverified`, and `skipped`, then re-show the repo table.
`addedUnverified[]` holds paths that were tracked but could **not** be read right now (they do
not exist, or are not directories). They are tracked deliberately — an unmounted volume is a
legitimate reason for a repo to be missing today — but they must be named, not glossed over:
"Tracked, but I couldn't read `<path>` — check for a typo, or mount it before the next token
sweep." An unreported phantom row stays in the backlog forever and quietly widens every
machine-wide total.
### Step 5 (mode `set-status`): Transition a repo — propose, approve, write
Confirm the repo is tracked (from Step 2's report) and that `status` is one of
pending/audited/planned/implemented. State the transition ("`<name>`: pending → audited") and
ask for approval.
When marking a repo **audited**, optionally attach its findings-by-severity so the machine-wide
roll-up stays meaningful. Two honest sources, in order of preference:
1. If the repo was audited in a config-audit session, read that session's finding counts and
build `{"critical":C,"high":H,"medium":M,"low":L}` — pass `--session <id>` too.
2. Otherwise, use counts the user provides. **Never invent counts** (Verifiseringsplikt) — if
none are available, transition the status without `--findings`.
On approval:
```bash
# Re-derive here: each fenced block is its own Bash call, so a TODAY set
# in an earlier block is empty by the time this one runs.
TODAY=$(date +%F)
node ${CLAUDE_PLUGIN_ROOT}/scanners/campaign-write-cli.mjs set-status <path> <status> \
--reference-date "$TODAY" \
[--findings '{"critical":0,"high":0,"medium":0,"low":0}'] [--session <id>] \
--output-file ~/.claude/config-audit/sessions/campaign-write.json 2>/dev/null; echo $?
```
Exit **3** = invalid status, untracked repo, or no ledger → report the message plainly and do
not retry blindly. On success, read the result and re-show the updated roll-up + repo row.
### Step 6 (mode `refresh-tokens`): Sweep tokens machine-wide — propose, approve, write
The token bill in Step 2 reflects the **last** sweep. `refresh-tokens` recomputes it: for every
tracked repo it runs the manifest's always-loaded accounting and refreshes the machine-wide bill —
the shared global layer (global CLAUDE.md + installed plugins + global agents/MCP) counted **once**,
plus each repo's own project-config delta.
It writes only the ledger's token fields (never status or findings), is idempotent (a re-sweep
replaces, never accumulates), and **skips — never aborts on** — any repo that can't be read. Tell
the user it will read each tracked repo's live config (a few seconds per repo), then on approval:
```bash
# Re-derive here: each fenced block is its own Bash call, so a TODAY set
# in an earlier block is empty by the time this one runs.
TODAY=$(date +%F)
node ${CLAUDE_PLUGIN_ROOT}/scanners/campaign-write-cli.mjs refresh-tokens \
--reference-date "$TODAY" \
--output-file ~/.claude/config-audit/sessions/campaign-write.json 2>/dev/null; echo $?
```
Exit **0** = swept (or nothing to sweep — a benign no-op on an empty campaign), **3** = no/corrupt
ledger. Read the result: `swept[]` (repos accounted for), `skipped[]` (each `{path, reason}`), and
the refreshed `rollUp.tokens`. Re-render the **Machine-wide token bill** (Step 2) with the new
numbers. If anything was skipped, name those repos plainly so the user knows the bill omits them
(honest coverage — Verifiseringsplikt).
### Step 7 (mode `export`): Export a repo's plan to its own `docs/` — preview, approve, write
"Planer følger arbeidsstedet": a planned repo's action plan belongs in **that repo's** `docs/`,
not buried in a session dir. This step copies it there, byte-faithfully.
**Preview first (read-only — never writes).** The repo must be tracked and have a linked session
that carries an `action-plan.md` (i.e. `/config-audit plan` has run there). Run without `--write`:
```bash
# Re-derive here: each fenced block is its own Bash call, so a TODAY set
# in an earlier block is empty by the time this one runs.
TODAY=$(date +%F)
node ${CLAUDE_PLUGIN_ROOT}/scanners/campaign-export-cli.mjs --repo "<path>" \
--reference-date "$TODAY" \
--output-file ~/.claude/config-audit/sessions/campaign-export.json 2>/dev/null; echo $?
```
Exit **0** = previewable, **1** = tracked but not exportable yet, **3** = error (untracked repo,
no/corrupt ledger). Read `~/.claude/config-audit/sessions/campaign-export.json` with the Read tool.
- **Exit 1 — read `problems`** and guide, then stop (nothing to export):
- `no-session-linked` → "`<name>` has no linked audit session. Link one with
`/config-audit campaign set-status <path> <status> --session <id>`, or audit + plan it first."
- `no-action-plan` → "`<name>`'s session has no plan yet. Run `/config-audit plan` in that repo
first, mark it `planned`, then export."
- **Exit 0 — show, then ask.** Tell the user the destination (`targetPath`) and a **short** preview
— the first ~12 lines of `document` only, never the whole file, never the raw JSON (UX rules).
Ask for explicit approval to write it.
**On approval, write it** (the CLI does the faithful copy — do NOT hand-write the file):
```bash
# Re-derive here: each fenced block is its own Bash call, so a TODAY set
# in an earlier block is empty by the time this one runs.
TODAY=$(date +%F)
node ${CLAUDE_PLUGIN_ROOT}/scanners/campaign-export-cli.mjs --repo "<path>" --write \
--reference-date "$TODAY" \
--output-file ~/.claude/config-audit/sessions/campaign-export.json 2>/dev/null; echo $?
```
Confirm: "Plan exported to `<targetPath>`." Then hand off to the **existing** execution machinery
(Block 4c reuses it — this command does not run it for you):
> To **execute**: run `/config-audit implement` in `<path>` — it backs up every changed file,
> applies the plan, and verifies. To **undo**: `/config-audit rollback`. When done, record it:
> `/config-audit campaign set-status <path> implemented`.
### Step 8: Next steps
Tailor to where the campaign stands:
- **Just initialized / few repos:** "`/config-audit campaign add --discover ~/repos` to enroll
your repos."
- **Pending repos exist:** "Run `/config-audit` in a pending repo to audit it, then
`/config-audit campaign set-status <path> audited` to record the result here."
- **Audited but not planned:** "`/config-audit plan` in that repo, then mark it `planned`."
- **Planned repos exist:** "`/config-audit campaign export <path>` to drop the plan into that
repo's own `docs/`, then `/config-audit implement` there to execute it (backup + verify)."
- **Backlog has items:** point at the top backlog repo and the natural next verb for its status
(audit → plan → export → implement).
- **No token bill yet (or config changed):** "`/config-audit campaign refresh-tokens` to compute
the machine-wide always-loaded cost — the shared global layer is the biggest lever."
- Always: the campaign survives this session — re-run `/config-audit campaign` anytime to see
the machine-wide picture.
## Notes
- **Read-only report, human-approved writes.** `campaign-cli` never writes; every mutation —
ledger changes via `campaign-write-cli`, plan exports via `campaign-export-cli --write` — happens
only after explicit approval, exactly mirroring how `/config-audit knowledge-refresh` gates
register writes. The export-CLI's default (no `--write`) is a read-only preview.
- **Deterministic core, not byte-stable command.** The lib transforms + all three CLIs are
unit-tested and deterministic (`--reference-date` injected); this command's orchestration is
judgment-driven and deliberately **not** in the snapshot suite (like `/config-audit optimize`
and `knowledge-refresh`).
- The `-cli` suffix keeps all three CLIs out of the scan-orchestrator, so the scanner count and
the byte-stable snapshot suite are unaffected.
- **THIN scope:** ledger + roll-up (findings + machine-wide token bill) + status + a cross-repo
prioritized backlog + plan export. The token sweep reuses the manifest's existing always-loaded
accounting per repo — it does not reinvent measurement, only aggregates it machine-wide.
Execution is **not** reinvented here — `export` drops a planned repo's plan into its own `docs/`
(a durable record), and the user runs the existing `/config-audit implement` (backup + apply +
verify) + `/config-audit rollback` to execute and undo. This command tracks state and routes the
work; it does not run audits or apply fixes itself.

View file

@ -75,7 +75,13 @@ Manage and clean up accumulated config-audit sessions in `~/.claude/config-audit
- Warn before deleting active sessions: "Session {id} is still active (phase: {phase}). Delete anyway?"
6. **Execute cleanup**:
- For each session to delete: `rm -rf ~/.claude/config-audit/sessions/{session-id}/`
- **Validate the id before it ever reaches `rm -rf`.** Each `{session-id}`
must match `^[0-9]{8}_[0-9]{6}$` (the id format `discover` generates) or be
an existing directory name read verbatim from the Glob in step 1. If an id
is empty or fails to match, **refuse to delete it**, report it, and continue
with the rest. An empty id expands the path to
`~/.claude/config-audit/sessions//`, which deletes *every* session.
- For each validated session: `rm -rf ~/.claude/config-audit/sessions/{session-id}/`
- Track deleted count and freed space
7. **Output summary**:

View file

@ -1,7 +1,7 @@
---
name: config-audit
description: Claude Code Configuration Intelligence - audit, analyze, and optimize your configuration
argument-hint: "[posture|tokens|manifest|feature-gap|fix|rollback|plan|implement|help|discover|analyze|interview|drift|plugin-health|whats-active|status|cleanup]"
argument-hint: "[posture|tokens|manifest|feature-gap|optimize|fix|rollback|plan|implement|help|discover|analyze|interview|drift|plugin-health|whats-active|campaign|knowledge-refresh|status|cleanup]"
allowed-tools: Read, Write, Glob, Grep, Bash, Agent, AskUserQuestion
model: opus
---
@ -17,6 +17,7 @@ If a subcommand is provided, route to it:
- `tokens``/config-audit:tokens`
- `manifest``/config-audit:manifest`
- `feature-gap``/config-audit:feature-gap`
- `optimize``/config-audit:optimize`
- `fix``/config-audit:fix`
- `rollback``/config-audit:rollback`
- `plan``/config-audit:plan`
@ -28,6 +29,8 @@ If a subcommand is provided, route to it:
- `drift``/config-audit:drift`
- `plugin-health``/config-audit:plugin-health`
- `whats-active``/config-audit:whats-active`
- `campaign``/config-audit:campaign`
- `knowledge-refresh``/config-audit:knowledge-refresh`
- `status``/config-audit:status`
- `cleanup``/config-audit:cleanup`
@ -87,7 +90,11 @@ Run both scanners and posture in a single Bash command. Default mode runs the hu
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/scan-orchestrator.mjs <target-path> --output-file ~/.claude/config-audit/sessions/{session-id}/findings/scan-results.json [--full-machine] [--global] $RAW_FLAG 2>/dev/null; node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --output-file ~/.claude/config-audit/sessions/{session-id}/posture.json [--full-machine] [--global] $RAW_FLAG 2>/dev/null; echo $?
# Set to the scope flags the detected scope calls for, otherwise leave empty.
# A placeholder in square brackets does not start with a dash, so both CLIs'
# arg loops would take it as the TARGET PATH instead of a flag.
SCOPE_FLAGS="" # e.g. SCOPE_FLAGS="--full-machine" or SCOPE_FLAGS="--global"
node ${CLAUDE_PLUGIN_ROOT}/scanners/scan-orchestrator.mjs <target-path> --output-file ~/.claude/config-audit/sessions/{session-id}/findings/scan-results.json $SCOPE_FLAGS $RAW_FLAG >/dev/null 2>/dev/null; node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --output-file ~/.claude/config-audit/sessions/{session-id}/posture.json $SCOPE_FLAGS $RAW_FLAG 2>/dev/null; echo $?
```
Use `--full-machine` for `full` scope, `--global` for `home` scope. For `repo` and `current`, pass the resolved path directly.

View file

@ -72,14 +72,19 @@ Run the scan orchestrator silently to discover and scan files. Default mode emit
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/scan-orchestrator.mjs <target-path> --output-file ~/.claude/config-audit/sessions/{session-id}/findings/scan-results.json [--full-machine] [--global] $RAW_FLAG 2>/dev/null; echo $?
# Set to the flag itself for full/home scope, otherwise leave empty. Never pass
# a placeholder wrapped in square brackets: it does not start with a dash, so
# the orchestrator's arg loop takes it as the SCAN TARGET and silently scans a
# path that does not exist.
SCOPE_FLAGS="" # e.g. SCOPE_FLAGS="--full-machine" or SCOPE_FLAGS="--global"
node ${CLAUDE_PLUGIN_ROOT}/scanners/scan-orchestrator.mjs <target-path> --output-file ~/.claude/config-audit/sessions/{session-id}/findings/scan-results.json $SCOPE_FLAGS $RAW_FLAG >/dev/null 2>/dev/null; echo $?
```
Check exit code: 0/1/2 → normal. 3 → "Discovery encountered an error. Try a narrower scope."
### Step 6: Save scope and state
Write `scope.yaml` and `state.yaml` to session directory. Update state with `current_phase: "discover"`, `next_phase: "analyze"`.
Write `scope.yaml` and `state.yaml` to session directory. Update state with all four fields `.claude/rules/state-management.md` requires: `current_phase: "discover"`, `completed_phases: [discover]`, `next_phase: "analyze"`, and `updated_at`. The last two are what make an interrupted run resumable.
### Step 7: Present summary

View file

@ -29,10 +29,10 @@ Tell the user: **"Saving current configuration as baseline..."**
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/drift-cli.mjs <path> --save --name <baseline-name> $RAW_FLAG 2>/dev/null
node ${CLAUDE_PLUGIN_ROOT}/scanners/drift-cli.mjs <path> --save --name <baseline-name> --json $RAW_FLAG 2>/dev/null
```
Read stdout for confirmation. Tell the user:
`--save` writes its human confirmation to **stderr**, which `2>/dev/null` discards — pass `--json` so the `{saved, name, path}` object lands on stdout. Read stdout for confirmation. Tell the user:
```markdown
### Baseline Saved
@ -50,10 +50,16 @@ Tell the user: **"Comparing current configuration against baseline..."**
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/drift-cli.mjs <path> --baseline <name> $RAW_FLAG 2>/dev/null
node ${CLAUDE_PLUGIN_ROOT}/scanners/drift-cli.mjs <path> --baseline <name> --output-file /tmp/config-audit-drift.json $RAW_FLAG >/dev/null 2>/dev/null; echo $?
```
Read stdout. In default mode the diff sections are humanized — finding titles, descriptions, and recommendations have already been replaced with plain-language equivalents. New/resolved/changed finding lists carry `userImpactCategory`, `userActionLanguage`, and `relevanceContext` so you can group and prioritize without re-deriving severity prose. If `--raw` was passed, the v5.0.0 diff is verbatim — present it in a code block as-is.
Exit codes: `0` = stable/improving, `1` = degrading (both normal — present the result either way), `3` = a real error.
Then read `/tmp/config-audit-drift.json` with the **Read tool**. The default-mode report itself goes to stderr, so `--output-file` is the only way this command sees the diff at all.
**Check `_baselineAnchor` first.** If the baseline was saved from a different directory than the one being scanned, the diff is not a drift signal — every baseline finding shows as "resolved" and every current finding as "new", which renders as a falsely reassuring "improving" trend. When the anchor differs, say so plainly and offer to re-anchor with `/config-audit drift --save` instead of presenting the numbers as drift.
In default mode the diff sections are humanized — finding titles, descriptions, and recommendations have already been replaced with plain-language equivalents. New/resolved/changed finding lists carry `userImpactCategory`, `userActionLanguage`, and `relevanceContext` so you can group and prioritize without re-deriving severity prose. If `--raw` was passed, the v5.0.0 diff is verbatim — present it in a code block as-is.
If baseline not found, tell the user:
@ -96,9 +102,14 @@ When iterating new/resolved findings, prefer `userActionLanguage` over raw `seve
If `$ARGUMENTS` contains `--list`:
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/drift-cli.mjs --list 2>/dev/null
node ${CLAUDE_PLUGIN_ROOT}/scanners/drift-cli.mjs --list --output-file /tmp/config-audit-baselines.json 2>/dev/null; echo $?
```
The human-readable listing goes to stderr, which `2>/dev/null` discards — read
`/tmp/config-audit-baselines.json` with the Read tool and render the `baselines`
array (`name`, `findingCount`, `savedAt`) as a table. If the array is empty, tell
the user no baselines are saved yet and point at `/config-audit drift --save`.
### What's next
After viewing drift:

View file

@ -42,7 +42,7 @@ Generate session ID (`YYYYMMDD_HHmmss`) if no active session exists.
mkdir -p ~/.claude/config-audit/sessions/{session-id}/findings 2>/dev/null
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --output-file ~/.claude/config-audit/sessions/{session-id}/posture.json $RAW_FLAG 2>/dev/null; echo $?
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --output-file ~/.claude/config-audit/sessions/{session-id}/posture.json $RAW_FLAG >/dev/null 2>/dev/null; echo $?
```
If exit code is non-zero: "Assessment couldn't run. Check that the path exists and contains configuration files."
@ -128,15 +128,23 @@ If the user picks numbers: parse the selection and proceed to Step 6.
For each selected recommendation:
1. **Create backup** of any files that will be modified:
1. **Create backup** of any files that will be modified.
Do **not** reach for `fix-cli.mjs` here. It is dry-run by default, so calling
it without `--apply` creates no backup at all and returns `backupId: null`
and calling it *with* `--apply` would execute unrelated auto-fixes that the
user never selected. Copy the files yourself:
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/fix-cli.mjs <target-path> --json 2>/dev/null
BACKUP_DIR=~/.claude/config-audit/backups/$(date +%Y%m%d_%H%M%S)/files
mkdir -p "$BACKUP_DIR" 2>/dev/null
# repeat per file that will be touched:
cp "<file-to-modify>" "$BACKUP_DIR/" 2>/dev/null; echo $?
```
Or create manual backup:
```bash
mkdir -p ~/.claude/config-audit/backups/$(date +%Y%m%d_%H%M%S)/files/ 2>/dev/null
```
Copy each file that will be touched.
Tell the user where the copies landed. These are plain file copies — they are
restored by copying them back, **not** by `/config-audit rollback`, which only
knows about backups written by `fix` and `implement`.
2. **Apply the template** from gap-closure-templates.md. Use the Write or Edit tool to create or modify the relevant configuration file.
@ -151,9 +159,12 @@ Implementing 3 recommendations...
4. **Verify** by re-running posture:
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --json --output-file /tmp/config-audit-verify-$$.json 2>/dev/null
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --json --output-file /tmp/config-audit-verify.json >/dev/null 2>/dev/null
```
Use the Read tool on `/tmp/config-audit-verify.json` for the new `overallGrade`
and score — stdout is discarded on purpose (the same envelope is 255 KB).
### Step 7: Show results
```markdown

View file

@ -15,8 +15,13 @@ Auto-fix deterministic configuration issues. Scans, plans fixes, backs up origin
- `$ARGUMENTS` may contain:
- A target path (default: current working directory)
- `--dry-run`: Show fix plan without applying
- `--global`: Include user-scope config (`~/.claude`) in the scan **and** the fix run
- `--raw`: Pass-through to scanners; produces v5.0.0 verbatim envelope (bypasses the humanizer) for byte-stable diff tooling
`--global` must be passed to **every** step below. The scan that builds the table and
the scan that plans the fixes are two different runs; if only one of them sees the
user scope, the plan and the table describe different config.
## Implementation
### Step 1: Greet and scan
@ -34,7 +39,11 @@ Parse flags and run scanners silently. Default mode emits humanized JSON — eac
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/scan-orchestrator.mjs <path> --output-file /tmp/config-audit-fix-scan-$$.json [--global] $RAW_FLAG 2>/dev/null; echo $?
# Set to --global when the user asked for global scope, otherwise leave empty.
# A placeholder in square brackets does not start with a dash, so the arg loop
# would take it as the scan/fix TARGET instead of a flag.
GLOBAL_FLAG=""
node ${CLAUDE_PLUGIN_ROOT}/scanners/scan-orchestrator.mjs <path> --output-file /tmp/config-audit-fix-scan.json $GLOBAL_FLAG $RAW_FLAG >/dev/null 2>/dev/null; echo $?
```
Exit code 3 → tell user: "Scanner error. Try `/config-audit posture` to check your configuration."
@ -44,10 +53,15 @@ Exit code 3 → tell user: "Scanner error. Try `/config-audit posture` to check
Run fix planner silently. The fix-cli emits humanized prose to stderr in default mode and v5.0.0-shape JSON to stdout when `--json` is set; we use `--json` here for structured data and let the humanizer-aware rendering layer (this command's prose output below) supply the plain-language wording from the scan envelope above:
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/fix-cli.mjs <path> --json 2>/dev/null
# Re-assign here: each fenced block is its own Bash call, so the value
# set in Step 1 is empty by the time this block runs.
GLOBAL_FLAG="" # --global when the user asked for global scope
node ${CLAUDE_PLUGIN_ROOT}/scanners/fix-cli.mjs <path> $GLOBAL_FLAG --output-file /tmp/config-audit-fix-plan.json 2>/dev/null; echo $?
```
Read the JSON output using the Read tool. Cross-reference each fix-plan entry against the humanized scan envelope (`/tmp/config-audit-fix-scan-$$.json`) by finding ID to recover the humanized `title`/`description`/`recommendation` plus `userImpactCategory`/`userActionLanguage` for grouping.
Exit codes: 0 = plan produced, 2 = one or more fixes failed (apply step only), 3 = argument or tool error. On 3, show the stderr message — an unknown flag is rejected by design, not silently ignored.
Read `/tmp/config-audit-fix-plan.json` using the Read tool. Cross-reference each fix-plan entry against the humanized scan envelope (`/tmp/config-audit-fix-scan.json`) by finding ID to recover the humanized `title`/`description`/`recommendation` plus `userImpactCategory`/`userActionLanguage` for grouping.
### Step 3: Present fix plan
@ -93,19 +107,27 @@ AskUserQuestion:
If confirmed, apply:
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/fix-cli.mjs <path> --apply --json 2>/dev/null
# Re-assign here: each fenced block is its own Bash call, so the value
# set in Step 1 is empty by the time this block runs.
GLOBAL_FLAG="" # --global when the user asked for global scope
node ${CLAUDE_PLUGIN_ROOT}/scanners/fix-cli.mjs <path> --apply $GLOBAL_FLAG --output-file /tmp/config-audit-fix-applied.json 2>/dev/null; echo $?
```
Read the JSON output to get applied/failed counts and backup location.
Read `/tmp/config-audit-fix-applied.json` with the Read tool to get applied/failed counts and the backup ID. Exit code 2 means at least one fix failed — report it; `failed[]` carries the reason per fix.
### Step 6: Show results
Run a quick posture check to measure improvement:
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <path> --json --output-file /tmp/config-audit-fix-posture-$$.json 2>/dev/null
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <path> --json --output-file /tmp/config-audit-fix-posture.json >/dev/null 2>/dev/null
```
Use the Read tool on `/tmp/config-audit-fix-posture.json` and take `overallGrade`
and the score from there. That read is the only source for the numbers below:
`--json` prints the same envelope to stdout, but 255 KB of raw JSON in the
transcript to recover one grade is exactly the waste this plugin exists to find.
Present results:
```markdown
@ -139,7 +161,8 @@ Run `/config-audit plan` to get a step-by-step guide for addressing these.
## Safety
- Backup is **mandatory** — every fix creates a backup first
- Backup is **mandatory** — every fix creates a backup first, including file renames (the source file is backed up before the rename, so rollback can restore it at its original path)
- Dry-run by default — user must confirm before changes
- Verify after fix — re-scans to confirm findings resolved
- Verify after fix — re-scans in the **same scope** the fix run used, so a `--global` run is verified against user scope too
- Rollback always available — `/config-audit rollback <backup-id>`
- A failed fix is reported, never swallowed — exit 2 plus a `failed[]` entry

View file

@ -32,9 +32,10 @@ if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
|---------|-------------|
| `/config-audit` | Full audit with auto-scope detection |
| `/config-audit posture` | Quick scorecard with A-F grades per area (10 areas) |
| `/config-audit tokens` | Opus-4.7 token hotspots; optional `--accurate-tokens` API calibration |
| `/config-audit tokens` | prompt-cache token hotspots; optional `--accurate-tokens` API calibration |
| `/config-audit manifest` | Ranked table of every system-prompt token source |
| `/config-audit feature-gap` | Deep analysis of features you're not using |
| `/config-audit optimize` | Optimization lens — config that works but fits a better mechanism (procedure→skill, lifecycle→hook, path→rule, never→permission) |
| `/config-audit fix` | Auto-fix deterministic issues; a copy of every changed file is saved first so you can roll back with one command |
| `/config-audit rollback` | Restore configuration from a saved copy |
@ -53,6 +54,8 @@ if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
| `/config-audit drift` | Compare current config against a saved baseline |
| `/config-audit plugin-health` | Audit plugin structure and the metadata block at the top of each command/agent file |
| `/config-audit whats-active` | Show active plugins/skills/MCP/hooks/CLAUDE.md with token estimates |
| `/config-audit campaign` | Track a machine-wide audit campaign across repos — per-repo status + roll-up, resumable across sessions (human-approved writes) |
| `/config-audit knowledge-refresh` | Keep the built-in best-practices knowledge current — flags guidance that's gone stale and checks sources for updates (changes need your approval) |
### Utility

View file

@ -22,12 +22,13 @@ Execute the action plan with full backup, verification, and rollback support.
### Step 1: Parse flags, load and verify
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
```
Check whether `$ARGUMENTS` contains `--raw`. Carry the answer yourself: the agent
prompt in Step 4 is **not** a shell, so a variable assigned in a bash block cannot
be referenced from it. Substitute `{mode}` literally with `--raw` or `humanized`.
Find the most recent session with a plan. If none: "No action plan found. Run `/config-audit plan` first."
Find the most recent session with a plan (use the **Glob tool** for
`~/.claude/config-audit/sessions/*/state.yaml`, then Read the newest match — Read
does not expand `*`). If none: "No action plan found. Run `/config-audit plan` first."
Use the Read tool on the action plan and count actions. Tell the user:
@ -51,14 +52,38 @@ AskUserQuestion:
### Step 3: Create backup
Create backup silently:
Create backup silently, and **print the backup ID** — Step 6 has to tell the user
how to roll back, and a timestamp that only ever existed inside a command
substitution cannot be quoted later. Shell state does not survive to the next
block, so capture the printed value and substitute it literally from here on:
```bash
mkdir -p ~/.claude/config-audit/backups/$(date +%Y%m%d_%H%M%S)/files/ 2>/dev/null
BACKUP_ID=$(date +%Y%m%d_%H%M%S)
mkdir -p ~/.claude/config-audit/backups/"$BACKUP_ID"/files/ 2>/dev/null
echo "$BACKUP_ID"
```
Use the printed ID wherever `{backup-id}` appears below. Never invent or re-derive
it with a second `date` call — a run that straddles a second boundary would hand
the user a rollback ID that does not exist.
Copy each file to be modified. Generate `manifest.yaml` with checksums.
The manifest is what `/config-audit rollback` reads, so it MUST carry both lists:
```yaml
files: # pre-existing files this run will MODIFY
- backup: files/root/CLAUDE.md
original: /abs/path/CLAUDE.md
sha256: <sha256 of the pre-change content>
created: # files this run will CREATE (no backup can exist)
- /abs/path/.claude/rules/post-quality.md
```
Record every `create`-type action under `created:`. Rollback cannot restore a
file that never existed, but it must be able to tell the user which files it is
leaving behind — a half-restored target is only dangerous when it is silent.
Tell the user: **"Backup created. Implementing actions..."**
### Step 4: Execute actions
@ -71,13 +96,15 @@ Agent(subagent_type: "config-audit:implementer-agent")
prompt: |
Execute action: {action-id}
File: {file-path}, Type: {create|modify|delete}
Mode: $RAW_FLAG (empty = humanized progress prose; "--raw" = v5.0.0 verbatim)
Mode: {mode} ("humanized" = humanized progress prose; "--raw" = v5.0.0 verbatim)
Details: {changes}
Verify backup exists, make change, validate syntax.
When logging progress, use the humanized title/userActionLanguage
fields from the action plan (the planner already rendered them) —
do not re-derive severity prose. Append result to:
~/.claude/config-audit/sessions/{session-id}/implementation-log.md
Append with Bash `>>` (heredoc) — NEVER the Write tool on this log;
parallel agents share it and a full-file Write clobbers their entries.
```
Show progress between groups using the humanized titles already present in the action plan:
@ -100,7 +127,19 @@ Agent(subagent_type: "config-audit:verifier-agent")
1. Modified files exist and are syntactically valid
2. New files created correctly
3. No new conflicts introduced
Report to: ~/.claude/config-audit/sessions/{session-id}/implementation-log.md
Return your findings as your final message. Do NOT write them to a file —
this agent is read-only by design (tools: Read, Glob, Grep) and has no
write tool; instructing it to write a report is a contract it cannot keep.
```
Append the verifier's returned findings to the log yourself, with Bash `>>`
(heredoc) — never the Write tool, for the same reason as Step 4:
```bash
cat >> ~/.claude/config-audit/sessions/{session-id}/implementation-log.md <<'EOF'
## Verification
{verifier findings}
EOF
```
If verifier finds issues: one retry with implementer agent. If still failing: report and suggest rollback.
@ -112,27 +151,49 @@ If verifier finds issues: one retry with implementer agent. If still failing: re
**{succeeded} succeeded** | {failed} failed | {skipped} skipped
{If score improved, run quick posture and show:}
Score impact: {old_grade} → {new_grade} (+{delta} points)
{If failed > 0:}
{failed} action(s) couldn't be completed — see log for details.
**Backup location:** `~/.claude/config-audit/backups/{timestamp}/`
**Rollback:** `/config-audit rollback {timestamp}`
**Backup location:** `~/.claude/config-audit/backups/{backup-id}/`
**Rollback:** `/config-audit rollback {backup-id}`
**Full log:** `~/.claude/config-audit/sessions/{session-id}/implementation-log.md`
```
**On reporting a score.** Only quote a grade *change* if the pre-change grade was
actually captured before Step 4 ran. Once the files are edited, only the new grade
is measurable — a delta computed after the fact has no source and must not be
invented. To offer one, measure first in Step 1 and again here:
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --output-file /tmp/config-audit-implement-posture.json 2>/dev/null; echo $?
```
Then Read `/tmp/config-audit-implement-posture.json`. Both the `--output-file` and
the `2>/dev/null` are required by the output rules — a bare scanner call would put
diagnostic output in front of the user. If no pre-change grade was captured, report
the new grade alone and say nothing about a delta.
### Step 7: Update state
Update `state.yaml` with `current_phase: "implement"`, `next_phase: null`.
Update `state.yaml` with all four fields `.claude/rules/state-management.md` requires:
- `current_phase: "implement"`
- `completed_phases`: append `implement` to the existing array (read it first; never replace it)
- `next_phase: null`
- `updated_at`: current timestamp
A full-file Write that names only two of the four silently deletes the other two.
## Rollback
If the user requests rollback at any point:
1. Read `manifest.yaml` from backup
2. Restore each file and verify checksums
3. Delete newly created files
3. **Report — do not delete — the files this run created.** Rollback restores from
backup, and no backup can exist for a file that did not exist before. Those
paths stay on disk; `/config-audit rollback` lists them under "Left in place"
so the user can remove them deliberately. Promising deletion here would leave a
half-restored config that reads as a clean rollback.
4. Update state to `rolled_back`
## Error Handling

View file

@ -11,8 +11,8 @@ Gather user preferences to inform the action plan.
## IMPORTANT: Inline Execution Only
This command runs AskUserQuestion **directly in the main context** — NOT via a Task subagent.
AskUserQuestion requires synchronous terminal interaction and does not work when delegated to a Task subagent.
This command runs AskUserQuestion **directly in the main context** — NOT via an `Agent` subagent.
AskUserQuestion requires synchronous terminal interaction and does not work when delegated to an `Agent` subagent.
## Prerequisites
@ -32,10 +32,29 @@ AskUserQuestion requires synchronous terminal interaction and does not work when
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
```
1. **Load session state**: Verify analysis phase completed, read analysis report for context
2. **Conduct interview inline**: Use AskUserQuestion tool directly (NOT via Task). Adapt questions based on analysis findings.
1. **Resolve the session, then load its state**:
```
Glob: ~/.claude/config-audit/sessions/*/state.yaml
Sort by modification time — the most recently modified session wins
```
Every path below substitutes that session's id for `{session-id}`. Never guess
it: if the Glob returns nothing, say "No audit session found — run
`/config-audit discover` first" and exit. Read the session's `state.yaml` and
verify `completed_phases` contains `analyze`; if it doesn't, tell the user
analysis hasn't run yet and exit. Then read the analysis report for context.
2. **Conduct interview inline**: Use AskUserQuestion tool directly (never delegate it to a subagent via `Agent` — a subagent cannot hold the interactive turn). Adapt questions based on analysis findings.
3. **Save interview results**: Write to `~/.claude/config-audit/sessions/{session-id}/interview.md`
4. **Update state** (see state-management rule)
4. **Update state** (see state-management rule), with one bound specific to this
command: interview is optional and can be run against a session that already
moved past it. If `completed_phases` already contains a later phase (`plan`,
`implement`, `verify`), do **not** rewind `current_phase` and do not re-add a
phase already in `completed_phases` — append `interview` only if it is absent,
leave `current_phase`/`next_phase` pointing at the furthest phase reached, and
tell the user the preferences will apply the next time `/config-audit plan`
runs. Rewinding a finished session is how its progress gets lost. Always set
`updated_at` to the current timestamp, whichever branch above applies.
5. **Output summary**
## Interview Questions

View file

@ -0,0 +1,152 @@
---
name: config-audit:knowledge-refresh
description: Keep the best-practices register fresh — flag stale entries, poll sources for new/changed practices, with human-approved writes only
argument-hint: "[--stale-after N] [--no-candidates]"
allowed-tools: Read, Write, Edit, Bash, WebSearch, WebFetch
model: opus
---
# Config-Audit: Knowledge Refresh
The "living" part of the living knowledge base. The optimization lens (`/config-audit
optimize`) is only as good as the best-practices register it reads — and Claude Code moves
fast. This command keeps `knowledge/best-practices.json` current in two ways:
- **Stale check (deterministic):** every CONFIRMED entry carries a `source.verified` date.
An entry older than the threshold (default **90 days**) is flagged for re-verification —
its source may have changed since.
- **Candidate poll (web):** scan the CC changelog + the Anthropic "Steering Claude Code"
docs/blog for *new* best-practices the register doesn't yet hold, or *changed* guidance
that contradicts an existing entry.
**The Iron rule (Verifiseringsplikt): nothing is ever auto-written.** Every change — a
bumped `verified` date, an updated claim, a brand-new entry — is presented to the user and
applied only on explicit approval, and only after the live source has actually been
re-read. No unverified claim enters the register.
## Implementation
### Step 1: Parse arguments
From `$ARGUMENTS`:
- `--stale-after N` → override the staleness threshold (integer days; default 90).
- `--no-candidates` → run the deterministic stale check only; skip the web poll.
Tell the user what's happening:
```
## Knowledge Refresh
Checking the best-practices register for stale entries (sources that may need
re-verification) and polling for new Claude Code practices...
```
### Step 2: Run the stale-check CLI
```bash
# Pass the threshold as its OWN quoted argument. Building "--stale-after 30" into
# one variable and expanding it unquoted only works if the shell word-splits —
# bash does, zsh (the macOS default) does not, and there the flag silently
# reverted to the 90-day default while the command reported success.
TODAY=$(date +%F)
STALE_AFTER_DAYS=$(echo "$ARGUMENTS" | sed -nE 's/.*--stale-after[ =]+([0-9]+).*/\1/p')
if [ -n "$STALE_AFTER_DAYS" ]; then
node ${CLAUDE_PLUGIN_ROOT}/scanners/knowledge-refresh-cli.mjs \
--reference-date "$TODAY" --stale-after "$STALE_AFTER_DAYS" \
--output-file ~/.claude/config-audit/sessions/knowledge-refresh.json 2>/dev/null; echo $?
else
node ${CLAUDE_PLUGIN_ROOT}/scanners/knowledge-refresh-cli.mjs \
--reference-date "$TODAY" \
--output-file ~/.claude/config-audit/sessions/knowledge-refresh.json 2>/dev/null; echo $?
fi
```
Exit code **0** = all fresh, **1** = some stale (advisory, normal), **3** = real error →
"The refresh check couldn't run — the register file may be missing or invalid."
### Step 3: Read the payload + present stale entries
Read `~/.claude/config-audit/sessions/knowledge-refresh.json` with the Read tool. It has
`counts {total, stale, fresh}`, a `stale[]` array (each: `id, verified, ageDays, url,
claim`), `referenceDate`, and `staleAfterDays`.
Present the stale entries as a markdown table (per the UX rules — never show the raw JSON):
| Entry | Claim (short) | Verified | Age (days) | Source |
|-------|---------------|----------|------------|--------|
If `counts.stale === 0`, say so plainly: "✓ All N register entries were re-verified within
the last {staleAfterDays} days." Then continue to the candidate poll (unless `--no-candidates`).
### Step 4: Candidate + source-change poll (web — skip if `--no-candidates`)
Tell the user this takes a moment ("Polling the changelog + Anthropic docs, ~20-40s...").
1. **Re-verify each stale entry.** `WebFetch` the entry's `source.url` and check whether the
claim it backs is **still accurate**. Three outcomes:
- *Still holds* → propose bumping `source.verified` to today (no claim change).
- *Changed* → propose an updated `claim`/`recommendation` quoting the new source text.
- *Cannot verify* (page gone, paywalled, contradicts) → propose **nothing**; flag it
"needs manual review" (Verifiseringsplikt: never bump a date you couldn't confirm).
2. **Look for new practices.** `WebSearch` the CC changelog and the "Steering Claude Code"
blog/docs for steering/config guidance not already represented by an entry's `lensCheck`.
For each genuine new practice, draft a **candidate** entry (next free `BP-<TOPIC>-NNN` id,
`confidence: "confirmed"` only if a primary source confirms it — otherwise mark `inferred`
and do **not** present it as user-facing).
### Step 5: Present everything for approval — write nothing yet
Group the proposals and ask the user to approve per item:
- **Re-verify (date bump):** "{id} — source re-read, claim still holds → bump verified to {today}?"
- **Update (claim drift):** show the old vs. new claim + the quoted source line.
- **New candidate:** show the drafted entry (id, claim, mechanism, recommendation, source).
- **Needs manual review:** list, with why it couldn't be auto-verified. (No write offered.)
Be explicit: **"I will not change any file until you approve specific items."**
### Step 6: Apply approved writes (only the approved ones)
**Where the register lives — say this before writing.** The register is part of the plugin,
and the stale check above read it from `${CLAUDE_PLUGIN_ROOT}/knowledge/best-practices.json`
(the `registerPath` field in the payload names the exact file). For a marketplace install that
is the plugin cache, so **an edit there is discarded by the next plugin upgrade** — the durable
home for an approved change is the plugin's own checkout. Tell the user which of the two they
are about to write to, using the `registerPath` they can see, before asking for approval.
For each approved item:
1. Edit `${CLAUDE_PLUGIN_ROOT}/knowledge/best-practices.json` — bump `source.verified`, update
the `claim`/`recommendation`, or append the new entry. Keep the file's 2-space JSON
formatting. Use the anchored path, never a bare `knowledge/…` — a relative path resolves
against the user's current repo, which is not the file the CLI read.
2. If a `${CLAUDE_PLUGIN_ROOT}/knowledge/*.md` mirror states the same fact, update it too so
the human-readable mirror doesn't drift from the register.
3. **Validate before declaring done** — re-run the register schema check and confirm zero errors:
```bash
node --test ${CLAUDE_PLUGIN_ROOT}/tests/lib/best-practices-register.test.mjs 2>&1 | tail -5
```
This test loads the register through the same anchored path, so it validates the file you
just edited — that only holds while step 1 uses the anchored path too. If validation fails,
revert that edit and report it — never leave the register invalid.
Report exactly what changed (ids + fields), and what was deferred to manual review.
### Step 7: Next steps
- `/config-audit optimize` — the lens now reads the refreshed register; re-run it to pick up
any new or changed mechanism-fit rules.
- Re-run `/config-audit knowledge-refresh --no-candidates` anytime for a quick staleness scan
without the web poll.
- Commit the register change (`knowledge/best-practices.json` + any `.md` mirror) with a
`chore(knowledge):` message so the provenance bump is in git history.
## Notes
- **Deterministic core, web-driven shell.** The stale classification is byte-stable and
unit-tested (`tests/lib/knowledge-refresh.test.mjs`, `tests/scanners/knowledge-refresh-cli.test.mjs`);
the candidate poll + writes are web/judgment-driven and deliberately **not** byte-stable
(mirrors `/config-audit optimize`).
- **Read-only CLI.** `knowledge-refresh-cli.mjs` never writes the register; all writes happen
here, in the command, after approval.
- The `-cli` suffix keeps it out of the scan-orchestrator, so the scanner count and the
snapshot suite are unaffected.

View file

@ -1,6 +1,6 @@
---
name: config-audit:manifest
description: Show ranked token-source manifest — every CLAUDE.md, plugin, skill, MCP server, and hook ordered DESC by estimated tokens
description: Show ranked token-source manifest — every CLAUDE.md, rule, agent, skill, output style, MCP server, and hook ordered DESC by estimated tokens, each tagged with its load pattern (always-loaded vs on-demand vs external), plus an always-loaded subtotal
argument-hint: "[path] [--json]"
allowed-tools: Read, Bash
model: sonnet
@ -10,6 +10,14 @@ model: sonnet
Produce a ranked, single-table view of every token source loaded for a given repo path. Where `whats-active` shows separate tables per category, `manifest` collapses everything into one ordered list — making it easy to see what's costing the most regardless of category.
Every source is tagged with its **load pattern**, derived from the published Claude Code loading model:
- **always** — enters context every turn before you type (project/user CLAUDE.md, unscoped rules, agents, output styles, MCP tool schemas). This is the cost that matters most: it is paid on *every* request.
- **on-demand** — loaded only when needed (skill bodies on invoke, path-scoped rules on a matching file read).
- **external** — runs outside the context window entirely (hooks).
The **always-loaded subtotal** is the headline number. Sources are component-level (a plugin contributes via its skills/rules/agents/output styles/hooks/MCP, each listed once — there is no coarse "plugin" roll-up, which would double-count).
## UX Rules (MANDATORY — from `.claude/rules/ux-rules.md`)
1. **Never show raw JSON or stderr output.** Always use `--output-file` + `2>/dev/null`.
@ -31,10 +39,9 @@ First non-flag argument is the path (default `.`). Recognized flags:
Tell the user: **"Building token-source manifest for `<path>`..."**
```bash
TMPFILE="/tmp/ca-manifest-$$.json"
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/manifest.mjs <path> --output-file "$TMPFILE" $RAW_FLAG 2>/dev/null; echo $?
node ${CLAUDE_PLUGIN_ROOT}/scanners/manifest.mjs <path> --output-file /tmp/config-audit-manifest.json $RAW_FLAG >/dev/null 2>/dev/null; echo $?
```
**Exit code handling:**
@ -44,38 +51,45 @@ node ${CLAUDE_PLUGIN_ROOT}/scanners/manifest.mjs <path> --output-file "$TMPFILE"
### Step 3: If `--json` was requested, cat the file and stop
```bash
cat "$TMPFILE"
cat /tmp/config-audit-manifest.json
```
Do NOT render the table in JSON mode.
### Step 4: Read JSON and render
Use the Read tool on `$TMPFILE`. Extract `meta.repoPath`, `total`, and `sources[]`. Render the top 20 sources (or fewer if the manifest is shorter):
Use the Read tool on `/tmp/config-audit-manifest.json`. Extract `meta.repoPath`, `total`, `summary`, and `sources[]`. Lead with the **always-loaded subtotal** (the headline), then render the top 20 sources (or fewer if the manifest is shorter):
```markdown
**Token-source manifest for `<repoPath>`** — ~{total} tokens at startup
**Token-source manifest for `<repoPath>`** — ~{total} tokens total
| Rank | Kind | Name | Source | Tokens |
|------|------|------|--------|--------|
| 1 | {kind} | `<name>` | {source} | ~{estimated_tokens} |
| ... | ... | ... | ... | ... |
- 🔴 **~{summary.always.tokens} tokens enter context every turn** before you type ({summary.always.count} always-loaded sources)
- 🟡 ~{summary.onDemand.tokens} tokens on-demand ({summary.onDemand.count} sources — loaded only when invoked / matched)
- ⚪ ~{summary.external.tokens} tokens external ({summary.external.count} sources — hooks, run outside context)
| Rank | Kind | Name | Source | Tokens | Load |
|------|------|------|--------|--------|------|
| 1 | {kind} | `<name>` | {source} | ~{estimated_tokens} | {loadPattern} |
| ... | ... | ... | ... | ... | ... |
_Load column: **always** / **on-demand** / **external**. Append `°` when `derivationConfidence` is `inferred` (no primary-doc row pins it exactly)._
_Estimates assume ~4 chars/token (Claude ballpark). Real token count varies ±15%._
```
If `sources.length > 20`, follow the table with: _"Showing top 20 of {N} sources. Run with `--json` to see the full list."_
When narrating, prioritize the always-loaded subtotal: a large **always** source is worse than an equally large **on-demand** one, because it is paid on every request. Call out any single always-loaded source that dwarfs the rest.
### Step 5: Suggest next steps
```markdown
**Next steps:**
- `/config-audit tokens`Opus-4.7 token-hotspot patterns (cache-breaking, redundant perms, deep imports, MCP budget)
- `/config-audit tokens`prompt-cache token-hotspot patterns (cache-breaking, redundant perms, deep imports, MCP budget)
- `/config-audit whats-active` — same data grouped by category, with disable suggestions
- `/config-audit feature-gap` — what *could* improve here, grouped by impact
```
Tone:
- High total (>50k): empathetic — "That's a heavy startup cost; tokens bullet anything you'd otherwise spend on the actual conversation."
- Moderate (1050k): neutral — "Reasonable. Skim the top 5 to see if anything is unexpectedly large."
- Low (<10k): encouraging — "Tight setup. The model has plenty of room for the actual work."
Tone (key on the **always-loaded subtotal** — the every-turn cost — not the grand total):
- High always-loaded (>40k): empathetic — "That's a heavy per-turn cost; it taxes every request before you've typed a word. Look at the largest always-loaded sources first."
- Moderate (1040k): neutral — "Reasonable. Skim the top always-loaded sources to see if anything is unexpectedly large."
- Low (<10k): encouraging — "Tight setup. The model has plenty of room for the actual work each turn."

150
commands/optimize.md Normal file
View file

@ -0,0 +1,150 @@
---
name: config-audit:optimize
description: Optimization lens — config that works but would fit a better mechanism (procedure→skill, lifecycle→hook, path→rule, never→permission)
argument-hint: "[path]"
allowed-tools: Read, Write, Glob, Grep, Bash, Agent
model: opus
---
# Config-Audit: Optimization Lens
The "is the config **optimal?**" axis (vs. the health scanners' "is it
**correct?**"). It finds configuration that *works* but uses a mechanism a
better one would fit — and frames every one as a *Missed opportunity*, never a
mistake.
Mechanism-fit rules come from the provenance-stamped best-practices register
(`knowledge/best-practices.json`); only CONFIRMED rules are surfaced. The motor
is hybrid: a cheap deterministic pre-filter finds candidates, then the opus
`optimization-lens-agent` judges each in context (precision-gated).
## What the user gets
- **Procedures → skills** (deterministic, CA-OPT-001)
- **Lifecycle phrasing → hooks** (BP-MECH-001)
- **Unscoped path-specific instructions → path-scoped rules** (BP-MECH-002)
- **Absolute "never" prohibitions → permissions / hooks** (BP-MECH-004)
- **`--subtract`:** instructions that no longer earn their always-loaded rent (BP-SUB-001)
Each finding cites its register rule + source URL. A clean CLAUDE.md returns "no
opportunities" — that is a good result, not a failure.
## Implementation
### Step 1: Determine target
Split `$ARGUMENTS` into a path (first non-flag argument; default: current working
directory) and flags. Recognized flags: `--global` (include the user `~/.claude`
cascade in discovery) and `--subtract` (add the subtraction axis, below).
**`--subtract` — the inverse question.** Every other lens asks what to *add* or
*move*; this one asks what no longer earns its always-loaded rent. It is opt-in
because it asks something different, and because deleting is not undoable by
reading. Pair it with `--global` to reach the user-level CLAUDE.md, where the
always-loaded cost actually sits (it loads in every repo, every session).
If `--subtract` is present, say so up front:
```
Also running the subtraction axis — instructions that cost tokens every turn
without telling me anything I couldn't work out. Load-bearing local facts
(remotes, versions, paths, policy) are excluded before anything is judged.
```
Tell the user:
```
## Optimization Lens
Looking for configuration that works but would fit a better Claude Code mechanism...
```
### Step 2: Run the lens CLI
Generate a session ID (`YYYYMMDD_HHmmss`) if no active session exists.
```bash
mkdir -p ~/.claude/config-audit/sessions/{session-id} 2>/dev/null
GLOBAL_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--global"; then GLOBAL_FLAG="--global"; fi
SUBTRACT_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--subtract"; then SUBTRACT_FLAG="--subtract"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/optimize-lens-cli.mjs <target-path> --output-file ~/.claude/config-audit/sessions/{session-id}/optimize-lens.json $GLOBAL_FLAG $SUBTRACT_FLAG 2>/dev/null; echo $?
```
Exit code 0 is normal. Only exit code 3 is a real error → "The lens couldn't run.
Check that the path exists and contains a CLAUDE.md."
### Step 3: Read the payload
Read `~/.claude/config-audit/sessions/{session-id}/optimize-lens.json` with the
Read tool. It has `deterministic` (already-confirmed OPT findings), `candidates`
(pre-filter candidates with register provenance), `register`, and `counts`.
Under `--subtract` it also has a `subtract` block (`candidates`, `register`) and
`counts.subtractCandidates`. Each subtraction candidate spans `line``endLine`
(a whole block). Include the whole `subtract` block when spawning the agent.
**Early exit:** if `counts.deterministic === 0` and `counts.candidates === 0`
(and, under `--subtract`, `counts.subtractCandidates === 0`), skip the agent and
tell the user plainly:
```
✓ No mechanism-fit opportunities found.
Your CLAUDE.md holds facts, not procedures/automation/prohibitions that would be
better as skills, hooks, rules, or permissions. Nothing to change here.
```
Then go to Step 5.
### Step 4: Spawn the precision gate
Tell the user what's happening and set expectations:
```
Found {counts.candidates} candidate line(s) + {counts.deterministic} deterministic finding(s).
Asking the optimization-lens agent to judge each in context (~20-40 seconds)...
```
Spawn the `optimization-lens-agent` (Agent tool) with:
- the full payload from Step 3 (deterministic + candidates + register),
- the session directory path so it can write `optimization-lens-report.md`.
The agent reads the actual CLAUDE.md, drops low-confidence candidates, and keeps
only genuine opportunities — each citing its register rule + source.
### Step 5: Present results
Read the agent's `optimization-lens-report.md` and present it formatted
(markdown tables / grouped sections). Follow the UX rules: never show raw JSON or
scanner progress; lead with a one-sentence summary of what was found before the
detail. Make clear these are LOW-severity *opportunities*.
If the agent kept nothing from the candidates (all dropped) but there were
deterministic findings, show those; if it kept nothing at all, show the clean
result from Step 3.
### Step 6: Next steps
End with context-sensitive next steps, explaining WHY each is useful:
- `/config-audit plan` — turn the kept opportunities into an action plan with
backups before you change anything.
- `/config-audit feature-gap` — the complementary lens: features you *don't* use
yet (this command is about mechanisms you *do* use that could fit better).
- Re-run `/config-audit optimize` anytime after editing CLAUDE.md.
## Notes
- This command is **agent-driven and not byte-stable** — its output is a
human-facing report, deliberately outside the deterministic snapshot suite.
- `--subtract` **proposes, never writes.** Nothing is deleted; act on a finding
via `/config-audit plan``/config-audit implement` (backup + rollback).
- The subtraction floor is deterministic and runs *before* the agent, so a
load-bearing block is never a candidate. It errs toward keeping: on a
well-maintained config this axis is mostly a no-op, and that is a good result.
- The deterministic half (CA-OPT-001) also rides in the normal orchestrated
audit; this command adds the prose-judgment half on top.
- No files are modified. To act on a finding, use `/config-audit plan`
`/config-audit implement` (backup + rollback) or edit by hand.

View file

@ -22,7 +22,11 @@ Generate a prioritized action plan based on analysis results.
### Step 1: Verify session state
Find the most recent session with analysis completed using the Read tool on `~/.claude/config-audit/sessions/*/state.yaml`. If none found: "No analysis results found. Run `/config-audit` first to scan your configuration."
Find the most recent session with analysis completed using the **Glob tool** on `~/.claude/config-audit/sessions/*/state.yaml`, then Read the newest match. The Read tool takes one literal path and does not expand `*` — pointing it at the glob makes this step report "no analysis results" even when a valid session exists.
If no session is found: "No analysis results found. Run `/config-audit` first to scan your configuration."
Then confirm the report itself exists — a session can carry a valid `state.yaml` and still be missing its report. Read `~/.claude/config-audit/sessions/{session-id}/analysis-report.md`. If it is absent: "Session {session-id} has no analysis report. Run `/config-audit analyze` to generate it." Stop — the planner agent has nothing to read.
### Step 2: Tell the user what's happening
@ -35,10 +39,10 @@ Actions are ordered by impact, with risk assessment and dependency tracking.
### Step 3: Parse flags and spawn planner agent
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
```
Check whether `$ARGUMENTS` contains `--raw`. Carry the answer yourself: the agent
prompt below is **not** a shell, so a variable assigned in a bash block cannot be
referenced from it. Substitute `{mode}` literally with `--raw` or with `humanized`
when writing the prompt.
Tell the user: **"Generating your action plan (this takes about 30 seconds)..."**
@ -49,7 +53,7 @@ Agent(subagent_type: "config-audit:planner-agent")
Generate action plan based on:
- Analysis: ~/.claude/config-audit/sessions/{session-id}/analysis-report.md
- Interview: ~/.claude/config-audit/sessions/{session-id}/interview.md (if exists)
Mode: $RAW_FLAG (empty = humanized; "--raw" = v5.0.0 verbatim severity prefiks)
Mode: {mode} ("humanized" = humanized; "--raw" = v5.0.0 verbatim severity prefiks)
Create a prioritized plan that consumes the humanized finding fields:
- Group actions by userImpactCategory (e.g., "Configuration mistake",
"Conflict", "Wasted tokens", "Missed opportunity", "Dead config")
@ -94,7 +98,14 @@ You can edit the plan file to remove, reorder, or modify actions before implemen
### Step 5: Update state
Update `state.yaml` with `current_phase: "plan"`, `next_phase: "implement"`.
Update `state.yaml` with all four fields `.claude/rules/state-management.md` requires — a partial write drops the fields that make an interrupted run resumable:
- `current_phase: "plan"`
- `completed_phases`: append `plan` to the existing array (read it first; never overwrite it with a fresh list)
- `next_phase: "implement"`
- `updated_at`: current timestamp
The planner agent may already have written these. Read the file before writing and preserve whichever fields it set — a full-file Write that names only two fields silently deletes the other two.
## Plan Modification

View file

@ -32,15 +32,21 @@ Auditing {N} plugin(s) for structure, frontmatter quality, and cross-plugin conf
### Step 2: Run scanner
Run silently for each plugin. Default mode emits a humanized JSON envelope where each PLH finding carries `userImpactCategory`, `userActionLanguage`, and `relevanceContext` alongside the v5.0.0 fields. `--raw` is passed through verbatim when present.
Run silently for each plugin. Default mode writes a humanized JSON payload to `--output-file` where each PLH finding carries `userImpactCategory`, `userActionLanguage`, and `relevanceContext` alongside the v5.0.0 fields. `--raw` is passed through verbatim when present, and prints the byte-stable v5.0.0 envelope on stdout instead.
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/plugin-health-scanner.mjs <path> $RAW_FLAG 2>/dev/null
node ${CLAUDE_PLUGIN_ROOT}/scanners/plugin-health-scanner.mjs <path> --output-file /tmp/config-audit-plugin-health.json $RAW_FLAG >/dev/null 2>/dev/null; echo $?
```
Read stdout output (JSON) using the Read tool. Parse findings.
Read `/tmp/config-audit-plugin-health.json` with the Read tool. Exit codes 0, 1 and 2 are normal; only 3 is a real error.
The payload carries three things the report needs:
- `plugins[]` — one row per plugin: `name`, `declaredName`, `commandCount`, `agentCount`, `findingCount`, `score`, `grade`. Use these for the table; never estimate a grade yourself.
- `cross_plugin_findings[]` — the namespace-collision and shared-command-name findings, already separated from the per-plugin ones (they also carry `crossPlugin: true` in `findings`).
- `findings[]` — every finding, humanized.
### Step 3: Present results
@ -49,7 +55,7 @@ Read stdout output (JSON) using the Read tool. Parse findings.
| Plugin | Grade | Commands | Agents | Status |
|--------|-------|----------|--------|--------|
| {name} | {grade} ({score}) | {cmd_count} | {agent_count} | {Good/Issues found} |
| {plugins[].name} | {plugins[].grade} ({plugins[].score}) | {plugins[].commandCount} | {plugins[].agentCount} | {Good/Issues found} |
| ... | ... | ... | ... | ... |
{If cross-plugin issues:}

View file

@ -42,27 +42,35 @@ Run silently — JSON goes to a file, the humanized scorecard prints to stderr (
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --output-file /tmp/config-audit-posture-$$.json $RAW_FLAG 2>/tmp/config-audit-posture-stderr-$$.txt; echo $?
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --output-file /tmp/config-audit-posture.json $RAW_FLAG >/dev/null 2>/tmp/config-audit-posture-stderr.txt; echo $?
```
Both paths are fixed literals, repeated literally in every later step: each
```bash fence is its own process, so a `$$`-derived path could never be named
again. `>/dev/null` is required, not cosmetic — with `--raw` the scanner writes
the full envelope to stdout *as well as* the file (`posture.mjs:101`), which is
255 KB on a real repo.
If exit code is non-zero, tell the user: "Assessment couldn't complete. Check that the path exists and contains Claude Code configuration files."
If `--raw` was passed, treat the captured stderr as v5.0.0-shape verbatim text and present it as-is in a code block; skip the humanized rendering steps below.
### Step 3: Read and interpret results
Read the JSON output file using the Read tool. Extract:
Use the Read tool on `/tmp/config-audit-posture.json`. Extract:
- `overallGrade`, `opportunityCount`
- `areas[]` — each with `name`, `grade`, `score`, `findingCount`
- `scannerEnvelope.scanners[].findings[]` — when surfacing individual findings, prefer the humanizer-provided fields: `userImpactCategory` (e.g., "Configuration mistake", "Wasted tokens"), `userActionLanguage` (e.g., "Fix this now", "Fix soon", "Optional cleanup"), and `relevanceContext` ("affects-everyone", "affects-this-machine-only", "test-fixture-no-impact"). These let you group and prioritize without hardcoded severity-to-prose mappings.
Also Read the captured stderr file — its body is the humanized scorecard (grade headline, area-score block, opportunity hint). You can present it verbatim or interleave its lines with the JSON-driven table.
Also use the Read tool on `/tmp/config-audit-posture-stderr.txt` — its body is the humanized scorecard (grade headline, area-score block, opportunity hint). You can present it verbatim or interleave its lines with the JSON-driven table.
### Step 4: Present the scorecard
```markdown
**Health: {overallGrade}** | {qualityAreaCount} areas scanned
**Health: {overallGrade}** | (area count: take it from the humanized scorecard's
"N areas reviewed" line — do NOT use `areas.length`, which counts Feature
Coverage; the table below excludes it, so the two would disagree)
{Use the headline line from the humanized stderr scorecard — it carries grade-context prose already (e.g., " Health: A (97/100) — Healthy setup, only minor polish needed"). Do not re-derive an A/B/C/D prose table here; the humanizer owns that vocabulary.}
@ -93,19 +101,19 @@ Avoid hardcoded grade-to-prose ladders here — the humanized scorecard headline
Run drift comparison silently:
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/drift-cli.mjs <target-path> 2>/dev/null
node ${CLAUDE_PLUGIN_ROOT}/scanners/drift-cli.mjs <target-path> --output-file /tmp/config-audit-posture-drift.json 2>/dev/null; echo $?
```
Read stdout output and append a "Configuration Drift" section showing what changed since the last baseline.
Use the Read tool on `/tmp/config-audit-posture-drift.json` and append a "Configuration Drift" section showing what changed since the last baseline. Both scanners report to stderr in default mode, which `2>/dev/null` discards — the payload is the only readable output.
**If `--plugin-health` flag is present:**
Run plugin health scanner silently:
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/plugin-health-scanner.mjs <target-path> 2>/dev/null
node ${CLAUDE_PLUGIN_ROOT}/scanners/plugin-health-scanner.mjs <target-path> --output-file /tmp/config-audit-posture-plh.json 2>/dev/null; echo $?
```
Read stdout output and append a "Plugin Health" section.
Use the Read tool on `/tmp/config-audit-posture-plh.json` and append a "Plugin Health" section, using its `plugins[]` rows for per-plugin grades.
**If both flags:** Use `scanners/lib/report-generator.mjs` to produce a unified markdown report.
@ -113,5 +121,9 @@ Read stdout output and append a "Plugin Health" section.
If a config-audit session exists, save results:
```bash
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --json --output-file ~/.claude/config-audit/sessions/<session-id>/posture.json 2>/dev/null
node ${CLAUDE_PLUGIN_ROOT}/scanners/posture.mjs <target-path> --json --output-file ~/.claude/config-audit/sessions/<session-id>/posture.json >/dev/null 2>/dev/null
```
This is a second scan on purpose: the session file stores the raw v5.0.0 shape,
while step 2 wrote the humanized one. `>/dev/null` matters most here — `--json`
sends the same envelope to stdout regardless of `--output-file`.

View file

@ -64,6 +64,18 @@ Use the Read tool on each backup's `manifest.yaml` (the list of changes captured
- hooks/hooks.json (checksum verified)
- .claude/rules/typescript.md (checksum verified)
```
5. **Report what rollback cannot undo.** A backup only holds files that already
existed, so files the implement step CREATED survive the restore. If the
manifest has a `created:` section (or `restoreBackup()` returns a non-empty
`createdNotRemoved`), list those paths and say plainly that they remain:
```
Left in place — created by implement, no backup exists:
- .claude/rules/post-quality.md
- guidelines/posting-rhythm.md
Remove them manually if you want the pre-implement state exactly.
```
Never finish a restore without this section when the list is non-empty; a
silently half-restored target reads as a clean rollback.
### Delete mode
@ -74,9 +86,14 @@ If user says "delete" after listing, confirm and remove the backup directory.
Use the backup and rollback libraries directly:
```javascript
import { listBackups, restoreBackup, deleteBackup } from '../scanners/rollback-engine.mjs';
import { parseManifest } from '../scanners/lib/backup.mjs';
import { parseManifest, getBackupDir } from '../scanners/lib/backup.mjs';
```
Both read `~/.claude/config-audit/backups` and fall back to the pre-v2.2.0
`~/.config-audit/backups`, so a backup made before the move still resolves;
`listBackups()` flags those with `legacy: true`. Prefer this API over ad-hoc
`cp` — it verifies the checksum before and after each write.
Or via Bash:
```bash
# List backups

View file

@ -37,8 +37,13 @@ When `--raw` is in `$ARGUMENTS`, render the raw `current_phase` field value verb
```bash
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
ALL_FLAG=""
if echo "$ARGUMENTS" | grep -qw -- "all"; then ALL_FLAG="all"; fi
```
When `ALL_FLAG` is set, skip steps 24 and render the **List All Sessions**
table below instead of a single session's status.
2. **Find active session**:
```
Glob: ~/.claude/config-audit/sessions/*/state.yaml
@ -128,11 +133,13 @@ All config-audit sessions:
| 20250120_160000 | implement | 2025-01-20 16:00 |
```
## Resume Session
## Resuming a session
If multiple sessions exist:
```
/config-audit resume {session-id}
```
There is no `resume` command. Sessions are selected by recency: every
session-aware command globs `~/.claude/config-audit/sessions/*/state.yaml` and
takes the most recently modified one.
Sets that session as active and continues from last phase.
To continue an older session, run its next phase directly — `/config-audit plan`,
`/config-audit implement`, and so on read `next_phase` from the state file. If
the wrong session keeps winning, delete the stale ones with
`/config-audit cleanup`.

View file

@ -1,6 +1,6 @@
---
name: config-audit:tokens
description: Show ranked token hotspots and Opus 4.7 pattern findings — what's costing the most per turn and how to reduce it
description: Show ranked token hotspots and prompt-cache pattern findings — what's costing the most per turn and how to reduce it
argument-hint: "[path] [--global]"
allowed-tools: Read, Bash
model: sonnet
@ -8,7 +8,7 @@ model: sonnet
# Config-Audit: Token Hotspots
Show the configuration sources that contribute the most tokens per turn, ranked by estimated tokens, with Opus 4.7-specific recommendations for reducing prompt-cache misses, schema bloat, and deep import chains.
Show the configuration sources that contribute the most tokens per turn, ranked by estimated tokens, with prompt-cache-aware recommendations for reducing cache misses, schema bloat, and deep import chains.
Complementary to `/config-audit whats-active`:
- **`whats-active`** = inventory view (what loads).
@ -28,6 +28,7 @@ Complementary to `/config-audit whats-active`:
Split `$ARGUMENTS` into a path and flags. Path is the first non-flag argument. Default to `.` (current working directory). Recognized flags:
- `--global` — also include the user-level `~/.claude/` cascade
- `--no-exclude-cache` — include stale `~/.claude/plugins/cache` versions in the ranking. **By default they are excluded** (cache-aware filtering, default ON): the cache holds superseded plugin versions that load on *zero* turns, and counting them used to crowd the top-10 with dead config. The active version of each plugin (per `installed_plugins.json`) is always kept — only stale versions are filtered. Use `--no-exclude-cache` to see the full on-disk walk.
- `--json` — emit raw JSON instead of rendered tables (power-user mode; bypasses the humanizer for byte-stable v5.0.0 output)
- `--raw` — pass-through to the scanner; produces v5.0.0 verbatim JSON (bypasses the humanizer). Use when piping into v5.0.0-baseline diff tooling.
- `--with-telemetry-recipe` — include `telemetry_recipe_path` in the JSON output, pointing to `knowledge/cache-telemetry-recipe.md`. Use this when you want to verify a structural fix actually improved cache hit rate (manual jq recipe, opt-in)
@ -39,12 +40,25 @@ Tell the user: **"Analysing token hotspots for `<path>`..."**
Default mode (no `--json`, no `--raw`) emits a humanized JSON envelope: each finding carries `userImpactCategory`, `userActionLanguage`, and `relevanceContext` in addition to the v5.0.0 fields. Pass `--raw` through verbatim if the user requested it.
```bash
TMPFILE="/tmp/config-audit-tokens-$$.json"
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/token-hotspots-cli.mjs <path> --output-file "$TMPFILE" [--global] $RAW_FLAG 2>/dev/null; echo $?
# Set each to the flag itself when the user asked for it, otherwise leave empty.
# A placeholder in square brackets does not start with a dash, so the CLI's arg
# loop would take it as the TARGET PATH instead of a flag.
GLOBAL_FLAG="" # --global
CACHE_FLAG="" # --no-exclude-cache
JSON_FLAG="" # --json
TELEMETRY_FLAG="" # --with-telemetry-recipe
node ${CLAUDE_PLUGIN_ROOT}/scanners/token-hotspots-cli.mjs <path> --output-file /tmp/config-audit-tokens.json $GLOBAL_FLAG $CACHE_FLAG $JSON_FLAG $TELEMETRY_FLAG $RAW_FLAG >/dev/null 2>/dev/null; echo $?
```
`--json` and `--with-telemetry-recipe` must be threaded here, not just
documented: the CLI is what turns them on. `--json`/`--raw` make the payload
byte-stable v5.0.0 (the humanizer is skipped, `token-hotspots-cli.mjs:127`), and
`--with-telemetry-recipe` is what adds `telemetry_recipe_path`. `>/dev/null` is
required because those two modes also print the payload to stdout even with
`--output-file` set (`token-hotspots-cli.mjs:137`).
**Exit code handling:**
- `0` → continue
- `3` → tell user: "Couldn't analyse tokens. Check that the path exists and is a directory." Stop.
@ -52,20 +66,24 @@ node ${CLAUDE_PLUGIN_ROOT}/scanners/token-hotspots-cli.mjs <path> --output-file
### Step 3: If `--json` was requested, cat the file and stop
```bash
cat "$TMPFILE"
cat /tmp/config-audit-tokens.json
```
Do NOT render tables in JSON mode.
### Step 4: Read JSON and render
Use the Read tool on `$TMPFILE`. Extract:
Use the Read tool on `/tmp/config-audit-tokens.json`. Extract:
- `total_estimated_tokens` — top-line number
- `hotspots[]` — top 10 ranked sources
- `findings[]`Opus 4.7 pattern findings (CA-TOK-001..003); each finding in default mode carries humanizer fields (`userImpactCategory`, `userActionLanguage`, `relevanceContext`) alongside the v5.0.0 fields
- `hotspots[]` — top 10 ranked sources; each carries a **load pattern** (`loadPattern` ∈ always / on-demand / external, plus `survivesCompaction` / `derivationConfidence`)
- `findings[]`prompt-cache pattern findings; each finding in default mode carries humanizer fields (`userImpactCategory`, `userActionLanguage`, `relevanceContext`) alongside the v5.0.0 fields
- `counts` — severity breakdown
A hotspot's **load pattern** matters as much as its size: an **always**-loaded source (CLAUDE.md, MCP tool schemas) is paid on *every* turn, an **on-demand** one (skill body, path-scoped rule) only when invoked/matched, and an **external** one (hooks, harness-config files like settings.json/.mcp.json) costs no per-turn context tokens at all. A big always-loaded hotspot is the most worth trimming.
**The stale plugin-cache finding is different from the rest.** All other TOK findings are about per-turn token cost. The *"Old plugin versions are sitting on disk"* finding (category `plugin-cache-hygiene`, impact **Dead config**, `--global` only) is a pure **disk-cleanup** item with **zero live-context impact** — the listed versions are never loaded. Render it as housekeeping, not a token problem: don't conflate its disk bytes with the per-turn token numbers above it.
Render as markdown. Group findings by `userImpactCategory` (e.g., "Wasted tokens" vs "Configuration mistake") rather than re-deriving severity prose; lead each line with `userActionLanguage` ("Fix this now", "Fix soon", "Optional cleanup", etc.) so the urgency phrasing stays consistent with the rest of the toolchain. The humanizer already replaced jargon-heavy `title`/`description`/`recommendation` strings with plain-language equivalents — render them verbatim.
```markdown
@ -73,9 +91,11 @@ Render as markdown. Group findings by `userImpactCategory` (e.g., "Wasted tokens
### Top hotspots (ranked by estimated tokens)
| Rank | Source | Tokens | Recommendations |
|------|--------|--------|-----------------|
| {rank} | `{source}` | ~{estimated_tokens} | {recommendations joined as `· ` bullets} |
| Rank | Source | Tokens | Load | Recommendations |
|------|--------|--------|------|-----------------|
| {rank} | `{source}` | ~{estimated_tokens} | {loadPattern} | {recommendations joined as `· ` bullets} |
_Load column: **always** (every turn) / **on-demand** (on invoke/match) / **external** (out-of-context). Append `°` when `derivationConfidence` is `inferred`._
### Findings, grouped by impact
@ -102,7 +122,7 @@ _Estimates assume ~4 chars/token (Claude ballpark). Real token count varies ±20
### Step 5: Cleanup and next steps
```bash
rm -f "$TMPFILE"
rm -f /tmp/config-audit-tokens.json
```
```markdown
@ -111,7 +131,7 @@ rm -f "$TMPFILE"
- **`/config-audit whats-active`** — full inventory of what loads (plugins, skills, MCP, hooks)
- **`/config-audit posture`** — overall health scorecard (Token Efficiency is the 8th area)
- **`/config-audit fix`** — auto-fix deterministic issues (where applicable)
- See `knowledge/opus-4.7-patterns.md` for the full pattern catalogue (CA-TOK-001 … 003)
- See `knowledge/prompt-cache-patterns.md` for the full pattern catalogue (CA-TOK-001 … 003)
- **Verify cache hit rate after a fix:** rerun with `--with-telemetry-recipe` to surface the path to `knowledge/cache-telemetry-recipe.md` — a copy-paste `jq` recipe that reads cache hit rate from your session transcripts. Opt-in. The TOK scanner is structural; this recipe is the runtime escape hatch.
```

View file

@ -33,10 +33,14 @@ Split `$ARGUMENTS` into a path and flags. Path is the first non-flag argument. D
Tell the user: **"Reading active configuration for `<path>`..."**
```bash
TMPFILE="/tmp/ca-whats-active-$$.json"
RAW_FLAG=""
if echo "$ARGUMENTS" | grep -q -- "--raw"; then RAW_FLAG="--raw"; fi
node ${CLAUDE_PLUGIN_ROOT}/scanners/whats-active.mjs <path> --output-file "$TMPFILE" [--verbose] [--suggest-disables] $RAW_FLAG 2>/dev/null; echo $?
# Set each to the flag itself when the user asked for it, otherwise leave empty.
# A placeholder in square brackets does not start with a dash, so the scanner's
# arg loop would take it as the TARGET PATH instead of a flag.
VERBOSE_FLAG="" # --verbose
SUGGEST_FLAG="" # --suggest-disables
node ${CLAUDE_PLUGIN_ROOT}/scanners/whats-active.mjs <path> --output-file /tmp/config-audit-whats-active.json $VERBOSE_FLAG $SUGGEST_FLAG $RAW_FLAG >/dev/null 2>/dev/null; echo $?
```
**Exit code handling:**
@ -46,14 +50,14 @@ node ${CLAUDE_PLUGIN_ROOT}/scanners/whats-active.mjs <path> --output-file "$TMPF
### Step 3: If `--json` was requested, cat the file and stop
```bash
cat "$TMPFILE"
cat /tmp/config-audit-whats-active.json
```
Do NOT render tables in JSON mode.
### Step 4: Read JSON and render
Use the Read tool on `$TMPFILE`. Extract:
Use the Read tool on `/tmp/config-audit-whats-active.json`. Extract:
- `meta.repoPath`, `meta.durationMs`, `meta.gitRoot`, `meta.projectKey`
- `totals.estimatedTokens.grandTotal` (and subtotals)
@ -149,7 +153,7 @@ Do NOT suggest items you can't name concrete redundancy for. If you can't find 3
### Step 7: Cleanup and next steps
```bash
rm -f "$TMPFILE"
rm -f /tmp/config-audit-whats-active.json
```
```markdown

View file

@ -0,0 +1,121 @@
# Claude Code config-surface delta — v2.1.114 → v2.1.181 (SUPERSEDED)
> **Status: SUPERSEDED** by `docs/cc-2.1.x-gap-matrix.md` (the verified truth source). This draft is
> kept only as the historical research artifact. Do not action rows from here — the matrix re-verified
> every row per-surface. Two version errors that propagated into both this draft and the matrix's
> draft-correction note were re-checked against the official changelog on 2026-06-18 (Batch 3) and
> corrected at the knowledge layer:
>
> - **Stop/SubagentStop `additionalContext`: v2.1.163** (not v2.1.165 — that release was "Bug fixes
> and reliability improvements" only). The gap-matrix even self-contradicts here: row 109 already
> said 2.1.163.
> - **settings `agent` field: introduced v2.0.59**; v2.1.157 is when it became "honored for dispatched
> `claude agents` sessions" — this draft's "2.1.154" and the matrix's "2.1.157-as-introduction" are
> both imprecise.
>
> Original draft note: produced by the claude-code-guide research agent 2026-06-18 from the official
> changelog. Items below the 2.1.114 floor (e.g. 2.1.105) belong to the token-management thread (Fase 4).
**Current latest:** v2.1.181 · **Window:** 2.1.114 → 2.1.181
## 1. settings.json / settings.local.json
- 2.1.181 `sandbox.allowAppleEvents` (macOS opt-in) — HIGH
- 2.1.181 env `CLAUDE_CLIENT_PRESENCE_FILE` (suppress mobile notifications) — HIGH
- 2.1.175 `enforceAvailableModels` (managed, constrains default model) — HIGH
- 2.1.176 `footerLinksRegexes` (footer badges) — MED
- 2.1.174 `wheelScrollAccelerationEnabled` — LOW
- 2.1.169 `disableBundledSkills` — HIGH
- 2.1.166 `fallbackModel` (chain up to 3) — HIGH
- 2.1.163 managed `requiredMinimumVersion`, `requiredMaximumVersion` — HIGH
- 2.1.161 env `OTEL_RESOURCE_ATTRIBUTES` — MED
- 2.1.154 `agent` (run main thread as named subagent) — HIGH
- 2.1.154 `fastModePerSessionOptIn` — HIGH
- 2.1.153 `skipLfs` (Git plugin sources) — MED
- 2.1.149 managed `allowAllClaudeAiMcps` — HIGH
- 2.1.143 `worktree.bgIsolation: 'none'` — HIGH
- 2.1.141 env `CLAUDE_CODE_PLUGIN_PREFER_HTTPS`, `ANTHROPIC_WORKSPACE_ID` — MED/HIGH
- 2.1.133 managed `parentSettingsBehavior`, `policyHelper` — HIGH
- 2.1.129 `skillOverrides` (per-skill visibility) — HIGH
## 2. CLAUDE.md / memory
- 2.1.169 `--safe-mode` disables customizations (CLAUDE.md not loaded) — MED
- 2.1.157 plugins in `.claude/skills` auto-load; "closest .claude/ wins" precedence — HIGH
- 2.1.178 nested `.claude/skills` directory loading — MED
## 3. Hooks
- 2.1.178 new matcher syntax `Tool(param:value)` — HIGH
- 2.1.172 OTEL `model` attribute in hook output context — MED
- 2.1.169 new event `post-session` (self-hosted runners) — HIGH
- 2.1.163 `SubagentStop` supports `additionalContext` return — HIGH
- 2.1.157 hook "args: string[]" exec form (no shell wrap) — HIGH
- 2.1.152 new event `MessageDisplay` — HIGH
- 2.1.145 hook input fields `background_tasks`, `session_crons` — HIGH
- 2.1.141 hook output field `terminalSequence` — HIGH
- 2.1.139 hook condition "if" syntax — HIGH
## 4. MCP (.mcp.json)
- 2.1.153 `skipLfs` for Git-based sources — MED
- 2.1.149 managed `allowAllClaudeAiMcps` (auto-approve claude.ai connectors) — HIGH
- 2.1.141 env `CLAUDE_CODE_PLUGIN_PREFER_HTTPS` (transport) — MED
- (format stdio/SSE/HTTP, scopes local/project/user, OAuth: stable in window)
## 5. Plugins & marketplace
- 2.1.178 "closest .claude/" precedence; nested `.claude/skills` — HIGH
- 2.1.157 plugins auto-load from `.claude/skills/`; `claude plugin init` — HIGH/MED
- 2.1.154 plugin `defaultEnabled: false` manifest field — HIGH
- 2.1.149 managed `pluginSuggestionMarketplaces` — MED
- 2.1.143 plugin dependency enforcement (disable/enable) — MED
- 2.1.142 root-level `SKILL.md` plugin surfacing — MED
- 2.1.128 `--plugin-dir ./x.zip` + `--plugin-url` (below window but note) — HIGH
## 6. Skills
- 2.1.178 nested `.claude/skills` directory loading — HIGH
- 2.1.163 `\$` escape for literal `$` in descriptions — MED
- 2.1.152 frontmatter `Skill` type indicator + `disallowed-tools`; `/reload-skills` — HIGH
- 2.1.142 root-level `SKILL.md` — MED
- (below window, Fase 4) 2.1.105 `skillListingBudgetFraction`, `maxSkillDescriptionChars`
## 7. Subagents / agents
- 2.1.178 "closest .claude/" precedence for agent conflicts — HIGH
- 2.1.172 nested sub-agents (5-level depth limit) — HIGH
- 2.1.157 `subagent_type` case/separator-insensitive; `/agents` discovery — MED
- 2.1.154 settings `agent` key; plugin `defaultEnabled` applies to agents — HIGH
- 2.1.143 agent spawning in auto mode — MED
## 8. Slash commands
- 2.1.147 `/simplify``/code-review` rename — LOW
- 2.1.152 `/code-review --fix` — LOW
- 2.1.139 `/goal`, `/scroll-speed` new commands — LOW
- (frontmatter argument-hint/allowed-tools/model + namespacing: stable)
## 9. Permissions
- 2.1.178 `Tool(param:value)` parameter-level matching — HIGH
- 2.1.166 glob pattern support in deny rule tool-name position — HIGH
- 2.1.154 auto mode hard-deny rules — HIGH
- 2.1.162 WebFetch precedence fixes — MED
- 2.1.160 single-file grep satisfies read-before-edit — MED
- (allow/deny/ask arrays: stable)
## 10. Model lineup
- 2.1.154 **Opus 4.8** new default; `/effort xhigh` new effort level — HIGH
- 2.1.158 auto mode on Bedrock/Vertex/Foundry for Opus 4.7/4.8 — HIGH
- 2.1.170 **Fable 5** (Mythos-class) introduced — HIGH
- 2.1.173 Fable 5 name normalization (`[1m]` suffix) — HIGH
- 2.1.142 fast mode default Opus 4.6 → 4.7 — HIGH
- settings model keys: `model`, `availableModels`, `enforceAvailableModels`, `fallbackModel`, `effortLevel` (now incl. `xhigh`)
## 11. Output styles / statusline / background / managed / sandbox
- 2.1.181 `sandbox.allowAppleEvents` — HIGH
- 2.1.176 `footerLinksRegexes`; `language` key — MED
- 2.1.169 `--safe-mode`; env `CLAUDE_CLIENT_PRESENCE_FILE` — HIGH
- 2.1.166 env `MAX_THINKING_TOKENS=0` — MED
- 2.1.153 statusline env `COLUMNS`, `LINES` — MED
- 2.1.143 `worktree.bgIsolation` — HIGH
- 2.1.141 hook `terminalSequence`; background agent permission-mode preservation — HIGH/MED
- 2.1.140 settings symlink hot-reload — MED
- 2.1.133 managed `parentSettingsBehavior`, `policyHelper` — HIGH
- 2.1.128 env `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS` — MED
## Sources
- Official changelog: https://raw.githubusercontent.com/anthropics/claude-code/main/CHANGELOG.md
- What's new: https://code.claude.com/docs/en/whats-new · Settings: https://code.claude.com/docs/en/settings · Plugins: https://code.claude.com/docs/en/plugins

View file

@ -0,0 +1,101 @@
# Plan — Fase 4 Items 2 + 3 (skill-listing token-styring)
**Skrevet:** 2026-06-18 · **Scope:** kun `config-audit/` · **Status:****LEVERT 2026-06-18 (alle 4 sykluser).** Designvalg bekreftet: A=fold inn i SKL, B=**Syklus 2 NÅ LEVERT** — `CA-SKL-002` aggregat-budsjett implementert med «faktum-først, 200k-anker» (operatør-bekreftet 2026-06-18): low severity, sum av beskrivelser (hver tellet opp til 1 536-cap) > 2 %×200k = 4000 tok → funn m/ CALIBRATION_NOTE om 5×-skalering på 1M-kontekst. C=`skillOverrides` i KNOWN_KEYS (IKKE TYPE_CHECKS — verdien er objekt; avvik fra §C-skissen, verify-først). Suite **875/875**. Detaljer i `STATE.md`.
**Syklus 2-notat (avvik fra §B-skissen, verify-først):** (1) Aggregatet teller hver beskrivelse **capped på 1 536** (= det som faktisk lastes i listingen; unngår dobbelttelling av halen `CA-SKL-001` allerede flagger). (2) Finding-IDs er **sekvensiell teller** (`output.mjs:33`), ikke stabil semantisk ID — tester matcher på **tittel**, aldri NNN; aggregatet emitteres ETTER cap-loopen så vanlig-tilfellet leser 001=cap, 002=aggregat. (3) Eksponerte en latent **HOME-lekkasje** i `posture-grade-stability.test.mjs` (kjørte CLI uten hermetisk HOME → ekte ~/.claude lekket inn → Token Efficiency falt A→B); fikset med `hermeticEnv()` (samme mønster som 8 andre CLI-tester). (4) Egen tilpasset humanizer-`static`-entry for aggregat-tittelen (ikke generisk `_default`).
**Forrige:** Fase 4 Item 1 (4.7→modell-nøytral) levert som `8376dab`. Baseline **856/856 grønn**.
## Mål
Lukke matrise-radene 167 + 169 (skill-listing token-styring): gjøre config-audit i stand til å anbefale/varsle om skill-listing-budsjett før CC selv trunkerer det. Item 2(a) (false-positive) er **allerede gjort** (Batch 1).
---
## Verifiserte fakta (changelog-cache, CC 2.1.181 — re-verifisert 2026-06-18 FØR skriving)
| Fakta | Versjon | Changelog-linje |
|---|---|---|
| `disableBundledSkills` + `CLAUDE_CODE_DISABLE_BUNDLED_SKILLS` — skjuler bundled skills/workflows/built-in slash commands fra modellen | **2.1.169** | L176 |
| Skill-listing-truncation-varsel viser nå *antall* berørte beskrivelser | **2.1.178** | L64 |
| Truncation-varsel flyttet fra startup til `/doctor` | 2.1.144 | L666 |
| `skillOverrides`: `off` / `user-invocable-only` / `name-only` (kollapser beskrivelse) | 2.1.129 | L995 |
| Per-beskrivelse listing-cap hevet 250 → **1 536 tegn** + startup-varsel ved truncation | 2.1.105 | L1502 |
| Skill-tegnbudsjett skalerer med kontekstvindu = **2 % av kontekst** | 2.1.32 | L2860 |
**Matrise-korreksjon (verify-først):** rad 170s hypotetiske settings-nøkler `skillListingBudgetFraction` / `maxSkillDescriptionChars` **finnes ikke** som nøkler. Budsjettet (2 % kontekst) og cap-en (1 536 tegn) er interne CC-mekanismer. Ekte config-levere er `skillOverrides` (2.1.129) + `disableBundledSkills` (2.1.169).
---
## Item 2 — `disableBundledSkills` som token-lever (matrise-rad 166 + 167)
### 2(a) — false-positive-fiks → ✅ ALLEREDE GJORT (Batch 1)
`disableBundledSkills` ligger allerede i `KNOWN_KEYS` (`scanners/settings-validator.mjs:25`) + `TYPE_CHECKS` boolean (:53), og dekkes av testen `NEW_KEYS` (`tests/scanners/settings-validator.test.mjs:133,176-180`). **Ingen kode.** Matrise-rad 166 markert DONE.
### 2(b) — anbefaling (matrise-rad 167)
**Designfunn:** GAP-checks (`scanners/feature-gap-scanner.mjs:91` `GAP_CHECKS`) er binære `{id,tier,title,recommendation,check}` der finding fyres når `check()` = false. En ren binær «`disableBundledSkills` ikke satt»-check ville fyre for nær alle brukere → **støy + dårlig råd** (lever-en fjerner ALLE bundled skills; bare verdt det ved faktisk listing-press).
**Anbefalt (DESIGNVALG A — confirm):** IKKE en frittstående GAP-check. Folde `disableBundledSkills`-anbefalingen inn som en **remediation-lever i Item 3s SKL-scanner** (fyres bare ved kvantitativt budsjett-press). Substansen i Item 2(b) leveres dermed via Item 3. Matrise-rad 167 omklassifiseres «folded into SKL (Item 3)».
**Alternativ:** frittstående tier-4 GAP-check gated på antall aktive skills (krever skill-enumerering → samme kobling som Item 3; mer kode, mer støyrisiko).
---
## Item 3 — SKL-scanner: skill-listing-budsjett (matrise-rad 169)
**Ny scanner** `scanners/skill-listing-scanner.mjs`, prefiks **`SKL`** (ledig; bekreftet mot alle eksisterende prefikser). Finding-IDs `CA-SKL-NNN`.
### Hva den gjør
1. **Primær (høy konfidens):** flagg hver aktiv skill-beskrivelse > **1 536 tegn** — verifisert hard cap → CC trunkerer den. Severity medium. Dette er rock-solid (ikke estimat).
2. **Aggregat (estimat, lavere konfidens):** sum av skill-beskrivelse-tegn på tvers av ALLE aktive skills (user + plugin + bundled). Varsle når summen nærmer seg/overskrider budsjett. Budsjett = 2 % av kontekst → **krever kontekstvindu-antakelse** (se DESIGNVALG B). Merk med `CALIBRATION_NOTE`-mønster (jf. TOK) og ±-usikkerhet — IKKE lat som vi vet brukerens kontekstvindu.
3. **Remediation (alle funn):** trim beskrivelse < 1 536; `skillOverrides: name-only` (kollaps); `disableBundledSkills: true` (fjern bundled). ← her bor Item 2(b).
### Gjenbruk (ikke gjenoppfinn)
- **Aktiv-skill-enumerering:** `enumerateSkills(pluginList)` i `scanners/lib/active-config-reader.mjs:465``{name, source:'user'|'plugin', pluginName, path, bytes, estimatedTokens}`. Dekker user + plugin skills (bredere enn project-local).
- **Beskrivelse-uttrekk:** `parseFrontmatter(content)?.frontmatter?.description` (`scanners/lib/yaml-parser.mjs`), nøyaktig som TOK pattern F (`token-hotspots.mjs:381-411`). Les fil via `readTextFile` (`lib/file-discovery.mjs`).
- **`estimateTokens()`** (`active-config-reader.mjs:35`) — markdown = bytes/4. Brukes hvis budsjett uttrykkes i tokens.
### Avgrensning mot eksisterende TOK pattern F
TOK pattern F flagger project-local `skill-md` med beskrivelse > **500 tegn** (strukturell «bloat»-heuristikk, `SKILL_DESCRIPTION_THRESHOLD`, `token-hotspots.mjs:53`). SKL er en **annen linse**: ALLE aktive skills (ikke bare project-local) + den **verifiserte 1 536-cap-en** (hard truncation) + aggregat-budsjett. Dokumentér de to linsene eksplisitt i begge scanner-headere så det ikke leses som duplikat.
---
## ⚠️ Åpne designvalg — BEKREFT med operatør FØR kode (per syklus)
- **A — Item 2(b)-plassering:** anbefalt = folde inn i SKL-remediation (ikke frittstående GAP-check). [Alternativ: tier-4 GAP-check.]
- **B — kontekstvindu-antakelse for aggregat-budsjettet:** anbefalt = konservativ default **200k tokens** (2 % ⇒ budsjett), eksplisitt merket estimat m/ `CALIBRATION_NOTE`. Eventuelt: dropp aggregat-sjekken i v1 og lever KUN den verifiserte 1 536-cap-en (høyest konfidens, null gjetting) — minst risiko. [Anbefaler å starte med per-beskrivelse-cap; aggregat som stretch.]
- **C (valgfri relatert quick-win):** `skillOverrides` mangler i `KNOWN_KEYS` → egen unknown-key false-positive (matrise-rad 170, L/S). Hvis SKL anbefaler `skillOverrides`, bør nøkkelen samtidig legges i `KNOWN_KEYS` + `TYPE_CHECKS` (`'string'`) for konsistens. Egen mini-syklus.
---
## Utførelsesrekkefølge (hver = ratifisert failing-test-først, Iron Law)
**Syklus 1 — Item 3 SKL-scanner (per-beskrivelse 1 536-cap):**
1. Failing test: `tests/scanners/skill-listing-scanner.test.mjs` + fixture med en skill-beskrivelse > 1 536 tegn → forvent `CA-SKL-001`. (RED: scanner finnes ikke.)
2. Skriv `scanners/skill-listing-scanner.mjs` (gjenbruk enumerateSkills + parseFrontmatter).
3. Registrer: import + `SCANNERS`-entry i `scanners/scan-orchestrator.mjs:19-65`; humanizer-entry i `lib/humanizer-data.mjs` (`SKL.static` + `_default`) + `SCANNER_TO_CATEGORY` i `lib/humanizer.mjs:30-43` (kategori `'Wasted tokens'`).
4. README-badge `scanners-12``13` (validert av `self-audit.mjs countScannerShape()`); CLAUDE.md scanner-tabell + `docs/scanner-internals.md`.
5. Snapshot-reseed (HERMETISK): `SEED_SNAPSHOT=1` (v5.0.0/) + `UPDATE_SNAPSHOT=1` (default-output/) — scan-orchestrator + posture endres hvis fixturen `marketplace-medium` trigges. Verifiser `git diff tests/snapshots/` + kontaminerings-grep.
**Syklus 2 — Item 3 aggregat-budsjett (HVIS designvalg B = inkluder):** failing test for sum-over-budsjett → `CA-SKL-002`; CALIBRATION_NOTE; kontekstvindu-konstant.
**Syklus 3 — Item 2(b) remediation:** verifiser at SKL-funn lister `disableBundledSkills` + `skillOverrides` + trim som levere (test på recommendation-tekst).
**(Valgfri) Syklus 4 — designvalg C:** `skillOverrides` i `KNOWN_KEYS`+`TYPE_CHECKS` + utvid `NEW_KEYS`-testen.
---
## Verifisering (testbare kriterier)
- `node --test 'tests/**/*.test.mjs'` grønn (baseline **856** + nye SKL-tester).
- Failing-test-først bevist per syklus (RED → GREEN logget).
- Fixture med beskrivelse > 1 536 tegn → nøyaktig `CA-SKL`-funn; fixture med korte beskrivelser → 0 SKL-funn.
- `node scanners/self-audit.mjs --check-readme` ren (badge 13 matcher).
- `git diff tests/snapshots/` reseedet bevisst, kontaminerings-grep ren (sadhguru/vegnorm/pluginCount).
- Ingen ny `Opus 4.x`-æra-stempling i ny scanner (gjenbruk modell-nøytral framing fra Item 1).
## Scope-gjerde
Kun `config-audit/`. Item 2(a) ferdig. Items 2(b)/3 = ny kode → ratifisert design (A+B) + failing-test-først per syklus. IKKE auto-eskalér; bekreft A/B FØR første scanner-linje.
## Risiko
- **Aggregat-budsjett er et estimat** (ukjent kontekstvindu) — håndteres med CALIBRATION_NOTE + å lede med den verifiserte 1 536-cap-en. Ved tvil: lever KUN per-beskrivelse-cap i v1.
- **Snapshot-drift:** ny orkestrert scanner endrer scan-orchestrator/posture-snapshots → MÅ reseedes hermetisk (`hermeticEnv()`), ALDRI rå `node scanners/...`.
- **Matrise er agent-generert** (2 kjente versjonsfeil) — fakta over er re-verifisert mot changelog; re-verifiser nye påstander før skriving.
## Pekere
- Kartlegging (denne sesjonen): scanner-registrering, GAP-mekanisme, enumerateSkills, snapshot-mønster — alt ankret i koden over.
- Gap-matrise: `docs/cc-2.1.x-gap-matrix.md` rad 166 (✅ DONE), 167 (Item 2b), 169 (Item 3), 170 (skillOverrides quick-win).
- Mønster-katalog: `knowledge/prompt-cache-patterns.md` (rad 167/169 retargetet hit).

251
docs/cc-2.1.x-gap-matrix.md Normal file
View file

@ -0,0 +1,251 @@
# Gap-matrise — Claude Code 2.1.114→181 vs config-audit-dekning
**Generert:** 2026-06-18 · **Metode:** workflow `config-audit-cc-gap` (12 surface-verifikatorer + syntese) · **Scope:** kun `config-audit/`
> **Dekningsstatus:** 12/12 surfaces verifisert (130 + 32 = 162 rader). Re-run `config-audit-cc-gap-rerun`
> (task `wyk5istx7`) fullførte subagents/agents, token management, slash commands. Se «Surfaces 1012» nederst.
> Re-run fanget også versjonsfeil i draft-korpuset — se «Draft-korreksjoner».
> **Verifiseringsplikt:** Rader merket UNVERIFIED, og **særlig MCP `trust`-feltet**, må verifiseres mot offisiell
> `.mcp.json`-dokumentasjon FØR endring — agenten kan ta feil i begge retninger. Manuelt forhåndsverifisert av
> hovedkontekst: `xhigh` (settings-validator.mjs:66), manglende hook-events, 4.7-hardkoding, feature-evolution v2.1.111.
## v5.4.0 reconciliation (2026-06-19 — supersedes the v5.3.0 block + rows below)
The six M-effort candidates deferred at v5.3.0 were re-verified against HEAD (`fe686b6`) on both
axes: **code-state** (still open?) and **CC premise** (real? — primary-source where risky).
See `docs/v5.4.0-release-plan.md` for the full per-candidate evidence.
**Operator GO 2026-06-19: "Option A"** → ship #1, #5, #4. No new scanner (badge stays 13).
| # | Candidate (row) | Code-state @HEAD | CC premise | Verdict |
|---|---|---|---|---|
| 1 | PLH shadow-folder (164) | OPEN (parses only name/desc/version) | CONFIRMED ~2.1.140 | **ship-5.4**`CA-PLH-015` |
| 5 | PLH `skills:`-array dirs (165) | OPEN (no `parsed.skills` read) | CONFIRMED ~2.1.145 `plugin validate` | **ship-5.4**`CA-PLH-016` |
| 4 | autoMode.hard_deny structure (179) | OPEN (key known, no nested val) | CONFIRMED (primary source) | **ship-5.4**`CA-SET-NNN` |
| 2 | acceptEdits-writes shell/build (176) | OPEN (0 matches) | CONFIRMED ~2.1.160, file-list unpinned | **defer** (needs primary-source field list) |
| 6 | nested-.claude closest-wins (139/166) | OPEN (no scanner) | CONFIRMED ~2.1.178 | **defer** (NEW scanner, badge bump → own release) |
| 3 | Read-deny hides Glob/Grep (175) | PARTIAL (`permission-rules.mjs:158-159` hint only) | **REFUTED** by primary source | **wontfix** |
**Row 175 correction (Verifiseringsplikt).** The row was framed **backwards**. The CC permissions
doc states: *"Claude makes a best-effort attempt to apply `Read` rules to all built-in tools that
read files like Grep and Glob"* — so a Read deny **already** covers Glob/Grep. A "Read-deny is
bypassable → false security" finding would be a false positive (same failure mode as the invented
MCP `trust` field). The only real Read-deny bypass is a Bash subprocess (python/node script that
opens files itself), which is documented behavior, not a config mistake. The Windows-path half (CC
normalizes `C:\…``/c/…` before matching) is real but narrow/low-value → rolling maintenance.
The L-priority `update-knowledge` rows (env vars, model nuance, hook-output fields, plugin/skill
doc, nested-.claude doc) remain rolling knowledge maintenance — fold in opportunistically.
## v5.3.0 reconciliation (2026-06-19 — supersedes stale rows below)
This matrix was the **v5.2.0** plan. Verified against HEAD (`9b828fa`) during the v5.3.0 Session A
audit. **Read this block first** — individual rows below predate the v5.2.0 + 8-commit work.
**CLOSED in v5.2.0** (entire HIGH-priority false-positive cluster — verified in current code):
- `settings-validator.mjs` — all 11 keys + `xhigh` are in `KNOWN_KEYS`/`VALID_EFFORT_LEVELS`
(incl. `footerLinksRegexes`, `agent`, `parentSettingsBehavior`, `sandbox`, `enforceAvailableModels`,
`fallbackModel`, `disableBundledSkills`, `pluginSuggestionMarketplaces`, `requiredMin/MaxVersion`,
`allowAllClaudeAiMcps`, `wheelScrollAccelerationEnabled`). All `settings.json`/`model lineup`/
`permissions` now-wrong settings rows → CLOSED.
- `hook-validator.mjs` — 28 events incl. `MessageDisplay` + `post-session`; knowledge says 28. CLOSED.
- `mcp-config-validator.mjs` — POSIX/auto-injected env allowlisted; `trust` removed. CLOSED.
- `claude-md-linter.mjs` — HIGH@500 → MEDIUM token-cost reframe, context-window aware. CLOSED.
- DIS/CNF — parameter-aware identity (Agent(model:…)/WebFetch(domain:…) no longer collapsed). CLOSED.
**CLOSED by the 8 unreleased commits (v5.3.0 candidates):**
- disableBundledSkills recommendation (row 167) → `dfe9049`.
- permissions `'*'` deny-all + allow non-MCP glob (row 138) → `03949c6`.
- CLAUDE.md char budget on top of the reframe → `b0bf8c5` (new `CA-CML`).
- DIS forbidden-param check (rows 137/139, beyond matrix) → `d678765`.
- PLH namespace collision (search-first; **NOT** the row-129 shadow-folder check) → `c6c5f17`.
- SKL aggregate skill-listing budget (row 169) shipped in v5.2.0 (`CA-SKL-002`); `0a631e3` extracted
the shared `context-window.mjs`.
**STILL OPEN → all DEFER to v5.4 (none are bugs/false-positives; all M-effort enhancements):**
row 129 PLH shadow-folder · row 141 acceptEdits-prompts-on-config-writes · row 140 Read-deny-hides-
Glob/Grep + Windows path · row 144 autoMode.hard_deny structure · row 130 PLH `skills:`-array ·
rows 104/131 nested `.claude` closest-wins detection. Plus the L-priority `update-knowledge` rows
(env vars, model nuance, hook-output fields, plugin/skill doc) → rolling knowledge maintenance.
**v5.3.0 ship-list = EMPTY** (operator GO 2026-06-19, "Release-only"). 3 knowledge-backing entries
(`disableBundledSkills`, 40k char-budget, forbidden-param — absent from corpus, grep-verified)
fold into the v5.3.0 docs step. See `docs/v5.3.0-release-plan.md` § "v5.3.0 scope decision".
## Executive summary
config-audits kunnskapsbase og scanner-kjent-sett er frosset ved ~v2.1.111; CC har shippet til v2.1.181 (Opus 4.8-æra). Av 130 verifiserte rader er det mest akutte en **klynge aktive false positives** — gyldig, dokumentert konfig flagges som feil i dag:
- **`settings-validator.mjs`** avviser `xhigh` og flagger `sandbox`, `fallbackModel`, `enforceAvailableModels`, `disableBundledSkills`, `pluginSuggestionMarketplaces` + 5 til som «unknown».
- **`mcp-config-validator.mjs`** flagger `${CLAUDE_PROJECT_DIR}` og POSIX `${var%pattern}` (CC fikset dette i 2.1.142). Dessuten: `trust`-feltet ser ut til å være **oppfunnet/uverifisert** — gir en medium på hver normal `.mcp.json`.
- **`disabled-in-schema-scanner.mjs`** kollapser param-kvalifiserte permission-regler (`Agent(model:opus)` deny + `Agent(model:sonnet)` allow) → feil «dead config».
- **`claude-md-linter.mjs`** gir HIGH ved 500 linjer med «significantly reduce adherence» — motsagt av CC (2.1.169 skalerer terskel etter kontekstvindu) og av pluginens egen knowledge.
- **`hook-validator.mjs`** mangler `MessageDisplay` (2.1.152) + `post-session` (2.1.169) → flagger gyldige hooks som ukjente.
Resten er stort sett L-prioritets knowledge-oppdateringer (env vars, modell-lineup inkl. Fable 5, doc-nyanser) uten scanner-impact. De fleste fiksene er S-effort kjent-sett/terskel-edits; en håndfull (DIS param-bevissthet, nye PLH/permissions-checks) er M.
## Anbefalt release
**v5.2.0** — endringssettet domineres av additive kjent-sett-utvidelser, terskel/ordlyd-fikser og knowledge-oppdateringer. `--json` og `--raw` forblir byte-stabile; false-positive-fiksene *fjerner* funn, endrer ikke envelope-skjemaet. Kun MCP `trust`-demoteringen endrer eksisterende output materielt (ikke-breaking kvalitetsfiks). v6.0.0 reserveres for fremtidig humanizer/envelope-skjemaendring.
## Voyage-eskalering: **NEI**
Netto-nytt scanner-arbeid (DIS param-bevissthet, PLH shadow-folder-check, permissions deny-all/glob, disableBundledSkills-anbefaling) er en håndfull uavhengige M-poster med klare specs og eksisterende fixture-mønstre — ikke en stor/usikker greenfield-build. Vanlig TDD per scanner holder. Eskalér kun hvis permissions-param-matchingen viser seg å trenge et delt rule-parsing-refactor på tvers av conflict-detector + DIS + settings-validator.
## False positives (dagens bugs — verdict = now-wrong)
| Tittel | Fil | Fiks (kort) |
|---|---|---|
| `effortLevel: xhigh` avvist | `settings-validator.mjs` | Legg `xhigh` i `VALID_EFFORT_LEVELS`. IKKE legg til `ultracode` (slash-alias, ikke settings-verdi) |
| `sandbox`-blokk «unknown» | `settings-validator.mjs` | Legg `sandbox` i `KNOWN_KEYS` (top-level er triggeren) |
| `fallbackModel` «unknown» | `settings-validator.mjs` | Legg i `KNOWN_KEYS`; verdi er string ELLER array(≤3) — ikke single-type TYPE_CHECK |
| `enforceAvailableModels` «unknown» | `settings-validator.mjs` | Legg i `KNOWN_KEYS` (managed) |
| `disableBundledSkills` «unknown» | `settings-validator.mjs` | `KNOWN_KEYS` + boolean TYPE_CHECK |
| `pluginSuggestionMarketplaces` «unknown» | `settings-validator.mjs` | `KNOWN_KEYS` (managed, array) |
| `requiredMin/MaximumVersion` «unknown» | `settings-validator.mjs` | `KNOWN_KEYS` (managed) |
| `allowAllClaudeAiMcps` «unknown» | `settings-validator.mjs` | `KNOWN_KEYS` (managed, boolean) |
| `footerLinksRegexes` «unknown» | `settings-validator.mjs` | `KNOWN_KEYS` (array) |
| `wheelScrollAccelerationEnabled` «unknown» | `settings-validator.mjs` | `KNOWN_KEYS` + boolean |
| `parentSettingsBehavior` «unknown» | `settings-validator.mjs` | `KNOWN_KEYS` (enum first-wins\|merge) |
| MCP `${CLAUDE_PROJECT_DIR}` «unreferenced» | `mcp-config-validator.mjs` | Allowlist auto-injiserte runtime-env-vars |
| MCP POSIX `${var%pattern}` «missing env» | `mcp-config-validator.mjs` | Begrens env-sjekk til bart `${IDENTIFIER}`; skip POSIX-operatorer |
| DIS kollapser `Agent(model:…)` deny/allow | `disabled-in-schema-scanner.mjs` | Gjør identitet param-bevisst (`bareTool()` :29-33) |
| DIS kollapser `WebFetch(domain:…)` | `disabled-in-schema-scanner.mjs` | Samme rotårsak — behandle domain:/param: som distinkt |
| `claude-md-linter` HIGH@500 + «reduce adherence» | `claude-md-linter.mjs` | Skaler terskel etter kontekstvindu / fjern absolutt HIGH; reframe til cache-prefix (CPS) |
| hook `MessageDisplay` «unknown event» | `hook-validator.mjs` | Legg i `VALID_EVENTS` + knowledge |
| hook `post-session` «unknown» | `hook-validator.mjs` | Legg i `VALID_EVENTS` (kebab, ≠ SessionEnd) |
| stale «26 total» hook-count | `hook-validator.mjs:137` + `hook-events-reference.md:3` | 26 → 28 i takt med de to nye |
| ✅ MCP `trust` oppfunnet (VERIFISERT 2026-06-18: ikke i offisiell `.mcp.json`-schema) | `mcp-config-validator.mjs` | DONE — fjernet CA-MCP-001 + «Invalid trust level»; `trust` droppet fra `VALID_SERVER_FIELDS` (stray `trust` → unknown field); humanizer + 5 knowledge-filer + fixturer scrubbet. Godkjenning er dialog/settings-basert (`enableAllProjectMcpServers`/`enabledMcpjsonServers`/`disabledMcpjsonServers`), aldri et JSON-felt. |
| `feature-evolution.md`/`capabilities.md` frosset v2.1.111 | `knowledge/*` | Opus 4.8 default + Fable 5; `/simplify``/code-review` |
## Quick wins (S-effort, høy verdi)
1. **`settings-validator.mjs`:** +11 gyldige nøkler i `KNOWN_KEYS` + `xhigh` i `VALID_EFFORT_LEVELS` (fix-false-positive).
2. **`hook-validator.mjs`:** +`MessageDisplay` +`post-session` til `VALID_EVENTS`; 26→28.
3. **`mcp-config-validator.mjs`:** begrens env-sjekk til bart `${IDENTIFIER}`; allowlist `CLAUDE_PROJECT_DIR`.
4. **knowledge-refresh:** `feature-evolution.md`/`capabilities.md` → Opus 4.8 default, Fable 5, `/effort xhigh`, `/simplify``/code-review`, lean system prompt.
5. **`hook-events-reference.md`:** dokumenter `additionalContext`, `reloadSkills`/`sessionTitle`, `args[]`-exec, `terminalSequence`, `CLAUDE_CODE_SESSION_ID`.
6. **skills:** +`disallowed-tools` i frontmatter-kjent-sett.
7. **plugins:** +`defaultEnabled`, `.claude/skills` auto-load, `claude plugin init`, root-SKILL.md, `allowAllClaudeAiMcps` i knowledge.
8. **env:** +window-env-vars i `configuration-best-practices.md`.
9. **rett draft-feil** i `docs/cc-2.1.x-changelog-delta.md` (skipLfs/HTTPS mis-filed under MCP; `Tool(param:value)` er permissions, ikke hooks).
## Større poster (M-effort)
| Tittel | Surface | Begrunnelse |
|---|---|---|
| DIS (+conflict-detector) param/domain-bevisst regel-identitet | permissions | `bareTool()`-kollaps gir ekte false positives nå som CC har param-matching (2.1.178) + domain-regler (2.1.172). conflict-detector.mjs:156 deler blindsonen |
| Skaler CLAUDE.md-lengdeterskel etter kontekstvindu | CLAUDE.md/memory | CC sluttet å flagge lange filer på 1M-modeller (2.1.169); krever modell-kontekst-deteksjon + tekst-rework + snapshot-tester |
| Ny PLH: advarsel når plugin.json-nøkkel skygger default-mappe | plugins | Speiler CCs egen 2.1.139/140-advarsel; ekte config-fallgruve |
| Nye permissions-checks: `*`-deny = deny-all; ikke-MCP-glob i allow / ukjent tool i deny | permissions | CC avviser/varsler ved startup (2.1.166) |
| feature-gap/posture: `disableBundledSkills` som token-effektivitets-lever | skills/tokens | Skjuler bundled-skill-descriptions fra system-prompt — ekte token-saver tokens-scanneren kan anbefale |
| PLH-paritet med `claude plugin validate`: valider `skills:`-array | plugins | Lav prioritet; PLH parser ikke `skills:`-arrays i dag |
## Full matrise (9/12 surfaces)
| surface | cc change | versions | verdict | action | target file | priority | effort |
|---|---|---|---|---|---|---|---|
| settings.json | effortLevel gains 'xhigh' tier | 2.1.154 | now-wrong | fix-false-positive | scanners/settings-validator.mjs | H | S |
| settings.json | sandbox.* block (+ allowAppleEvents) valid top-level key | ~2.1.77-83; 2.1.181 | now-wrong | fix-false-positive | scanners/settings-validator.mjs | H | S |
| settings.json | enforceAvailableModels managed setting | 2.1.175 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | H | S |
| settings.json | fallbackModel (string or up-to-3 array) | 2.1.166 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | H | S |
| settings.json | disableBundledSkills setting | 2.1.169 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | H | S |
| settings.json | requiredMinimumVersion / requiredMaximumVersion | 2.1.163 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | M | S |
| settings.json | allowAllClaudeAiMcps managed setting | 2.1.149 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | M | S |
| settings.json | footerLinksRegexes setting | 2.1.181 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | M | S |
| settings.json | wheelScrollAccelerationEnabled setting | 2.1.174 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | L | S |
| settings.json | pluginSuggestionMarketplaces managed setting | 2.1.152 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | M | S |
| settings.json | worktree.bgIsolation nested value | 2.1.154 | partial | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| settings.json | env CLAUDE_CLIENT_PRESENCE_FILE | 2.1.181 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| settings.json | env OTEL_RESOURCE_ATTRIBUTES labels | 2.1.161 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| settings.json | env CLAUDE_CODE_PLUGIN_PREFER_HTTPS | 2.1.141 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| settings.json | env ANTHROPIC_WORKSPACE_ID | 2.1.141 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| settings.json | env MAX_THINKING_TOKENS=0 disables thinking | 2.1.166 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| settings.json | parentSettingsBehavior / policyHelper (UNVERIFIED) | 2.1.133 | missing | new-feature-gap-check | scanners/settings-validator.mjs | L | S |
| settings.json | skillOverrides | 2.1.129 | DONE | added to KNOWN_KEYS (NOT TYPE_CHECKS — value is per-skill object off/user-invocable-only/name-only) | scanners/settings-validator.mjs | L | S |
| settings.json | ~~skillListingBudgetFraction / maxSkillDescriptionChars~~ (keys DO NOT EXIST — verified) | 2.1.105 + 2.1.32 | DONE | budget=2% ctx + cap=1536 are internal CC mechanisms, not settings keys; both enforced by the SKL scanner — `CA-SKL-001` (1,536-char cap) + `CA-SKL-002` (aggregate 2%-of-context budget, 200k-anchored estimate) | scanners/skill-listing-scanner.mjs | L | M |
| CLAUDE.md / memory | 'too long' threshold now scales with context window | 2.1.169 | now-wrong | fix-false-positive | scanners/claude-md-linter.mjs | H | M |
| CLAUDE.md / memory | --safe-mode / CLAUDE_CODE_SAFE_MODE disables customizations | 2.1.169 | missing | update-knowledge | knowledge/configuration-best-practices.md | L | S |
| CLAUDE.md / memory | plugins in .claude/skills auto-load (cascade/token accounting) | 2.1.157 | partial | new-feature-gap-check | knowledge/configuration-best-practices.md | M | M |
| CLAUDE.md / memory | nested .claude closest-wins precedence + <dir>:<name> | 2.1.178, 2.1.161 | partial | new-feature-gap-check | knowledge/configuration-best-practices.md | M | M |
| CLAUDE.md / memory | CLAUDE_MEMORY_STORES team-memory env var | 2.1.172 | missing | update-knowledge | knowledge/configuration-best-practices.md | L | S |
| hooks | MessageDisplay hook event | 2.1.152 | now-wrong | fix-false-positive | scanners/hook-validator.mjs | H | S |
| hooks | post-session lifecycle hook | 2.1.169 | now-wrong | fix-false-positive | scanners/hook-validator.mjs | M | S |
| hooks | stale '26 total' hook-event count (scanner+knowledge) | derived | now-wrong | update-knowledge | knowledge/hook-events-reference.md | M | S |
| hooks | Stop/SubagentStop return additionalContext | 2.1.163 | partial | update-knowledge | knowledge/hook-events-reference.md | M | S |
| hooks | terminalSequence hook output field | 2.1.141 | partial | update-knowledge | knowledge/hook-events-reference.md | L | S |
| hooks | Stop/SubagentStop input background_tasks/session_crons | 2.1.145 | missing | update-knowledge | knowledge/hook-events-reference.md | L | S |
| hooks | SessionStart reloadSkills / sessionTitle output | 2.1.152 | partial | update-knowledge | knowledge/hook-events-reference.md | L | S |
| hooks | hook args: string[] exec form | 2.1.139 | partial | update-knowledge | knowledge/hook-events-reference.md | M | S |
| hooks | CLAUDE_CODE_SESSION_ID hook env var | 2.1.163 | missing | update-knowledge | knowledge/hook-events-reference.md | L | S |
| mcp | ${CLAUDE_PROJECT_DIR} auto-injected, flagged unreferenced | 2.1.139 | now-wrong | fix-false-positive | scanners/mcp-config-validator.mjs | H | S |
| mcp | POSIX ${var%pattern} expansions flagged as missing env (CC fixed) | 2.1.142 | now-wrong | fix-false-positive | scanners/mcp-config-validator.mjs | H | S |
| mcp | ${VAR} resolves from shell env, not just env block | 2.1.161 | now-wrong | fix-false-positive | scanners/mcp-config-validator.mjs | M | S |
| mcp | 'trust' field invented/enforced; unverified in docs/schema | n/a | now-wrong | fix-false-positive | scanners/mcp-config-validator.mjs | M | M |
| mcp | allowAllClaudeAiMcps managed key (knowledge) | 2.1.174/2.1.149 | missing | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| mcp | managed allowed/deniedMcpServers ${VAR} predicate support | 2.1.162, 2.1.169 | partial | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| mcp | agent-frontmatter mcpServers + --strict-mcp-config/--mcp-config flags | 2.1.154, 2.1.169 | partial | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| plugins & marketplace | plugin.json defaultEnabled: false | 2.1.154 | missing | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| plugins & marketplace | plugins auto-load from .claude/skills (no marketplace) | 2.1.157, 2.1.161 | partial | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| plugins & marketplace | claude plugin init scaffolding | 2.1.157, 2.1.161 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| plugins & marketplace | pluginSuggestionMarketplaces managed key flagged unknown | 2.1.152 | now-wrong | fix-false-positive | scanners/settings-validator.mjs | H | S |
| plugins & marketplace | disableBundledSkills setting flagged unknown | 2.1.169 | now-wrong | fix-false-positive | scanners/settings-validator.mjs | H | S |
| plugins & marketplace | plugin dependency enforcement (manifest deps) | 2.1.143-145 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| plugins & marketplace | root-level SKILL.md (no skills/) surfaced as skill | 2.1.142-143 | partial | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| plugins & marketplace | warn when plugin.json key shadows default component folder | 2.1.139-140 | missing | new-feature-gap-check | scanners/plugin-health-scanner.mjs | M | M |
| plugins & marketplace | validate skills: array entries are dirs within plugin | 2.1.145 | missing | new-feature-gap-check | scanners/plugin-health-scanner.mjs | L | M |
| plugins & marketplace | nested .claude closest-wins + <dir>:<name> vs flat conflict detection | 2.1.178 | partial | update-knowledge | knowledge/claude-code-capabilities.md | M | M |
| skills | disallowed-tools skill/command frontmatter | 2.1.152 | partial | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| skills | disableBundledSkills token-efficiency lever | 2.1.169 | missing | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| skills | nested .claude/skills load + <dir>:<name> disambiguation | 2.1.178, 2.1.181 | partial | update-knowledge | knowledge/claude-code-capabilities.md | M | M |
| skills | \$ escape syntax in command bodies | 2.1.178 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| skills | skillListingBudgetFraction / maxSkillDescriptionChars (below floor) | 2.1.105 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | M |
| permissions | Tool(param:value) rules — DIS collapses to bare tool | 2.1.178 | now-wrong | fix-false-positive | scanners/disabled-in-schema-scanner.mjs | H | M |
| permissions | glob in deny tool-name; '*' deny-all; allow non-MCP glob rejected | 2.1.166 | missing | new-feature-gap-check | scanners/conflict-detector.mjs | M | M |
| permissions | WebFetch(domain:*) / mid-pattern wildcards — DIS false positive | 2.1.172 | now-wrong | fix-false-positive | scanners/disabled-in-schema-scanner.mjs | H | M |
| permissions | Read deny hides Glob/Grep; Windows backslash/case path rules | 2.1.162 | missing | new-feature-gap-check | scanners/settings-validator.mjs | L | M |
| permissions | acceptEdits prompts on shell-startup/build-tool config writes | 2.1.160 | missing | new-feature-gap-check | scanners/settings-validator.mjs | M | M |
| permissions | sandbox.allowAppleEvents — sandbox block flagged unknown | 2.1.181 | now-wrong | fix-false-positive | scanners/settings-validator.mjs | H | S |
| permissions | parentSettingsBehavior admin-tier key flagged unknown | 2.1.133 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | M | S |
| permissions | autoMode.hard_deny sub-key (no structure validation) | 2.1.136 | partial | extend-scanner-known-set | scanners/settings-validator.mjs | M | M |
| model lineup | effortLevel missing 'xhigh' (Opus 4.8) | 2.1.154 | now-wrong | fix-false-positive | scanners/settings-validator.mjs | H | S |
| model lineup | Claude Fable 5 introduced (Mythos-class) | 2.1.170 | partial | update-knowledge | knowledge/configuration-best-practices.md | M | M |
| model lineup | auto mode on Bedrock/Vertex/Foundry for Opus 4.7/4.8 | 2.1.158 | partial | new-feature-gap-check | knowledge/configuration-best-practices.md | L | M |
| model lineup | enforceAvailableModels flagged unknown | 2.1.175 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | H | S |
| model lineup | fallbackModel flagged unknown | 2.1.166 | now-wrong | extend-scanner-known-set | scanners/settings-validator.mjs | H | S |
| feature evolution / staleness | Opus 4.8 new default (timeline frozen v2.1.111) | 2.1.154 | missing | update-knowledge | knowledge/feature-evolution.md | H | S |
| feature evolution / staleness | /effort xhigh new top tier | 2.1.154 | missing | update-knowledge | knowledge/feature-evolution.md | M | S |
| feature evolution / staleness | Claude Fable 5 new model family | 2.1.170 | missing | update-knowledge | knowledge/feature-evolution.md | H | S |
| feature evolution / staleness | lean system prompt now default for 4.8+ | 2.1.154 | partial | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| feature evolution / staleness | /simplify renamed to /code-review (stale bundled-skill list) | 2.1.147, 2.1.152 | now-wrong | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| feature evolution / staleness | window settings keys absent from capability register | 2.1.143-181 | missing | update-knowledge | knowledge/claude-code-capabilities.md | M | M |
| feature evolution / staleness | window hook events absent from timeline | 2.1.139-178 | missing | update-knowledge | knowledge/feature-evolution.md | M | M |
_(L-prioritets «covered/none» runtime/UI-rader utelatt for lesbarhet; full liste i workflow-output `tasks/wgxug2xfs.output`.)_
## Surfaces 1012 (re-run `wyk5istx7`, 12/12 komplett)
Mest actionable nye funn:
| surface | cc change | versions | verdict | action | target file | priority | effort |
|---|---|---|---|---|---|---|---|
| token management | `disableBundledSkills` flagges «unknown» (token-lever) | 2.1.169 | ✅ DONE (Batch 1) — i `KNOWN_KEYS`+`TYPE_CHECKS` | fix-false-positive | scanners/settings-validator.mjs | H | S |
| token management | `disableBundledSkills` kan ANBEFALES (skjuler bundled-skill-desc fra system-prompt) | 2.1.169 | missing | new-feature-gap-check | knowledge/prompt-cache-patterns.md | M | S |
| token management | `prompt-cache-patterns.md` (rendøpt fra `opus-4.7-patterns.md`); modell-nøytral mekanisme-tekst + ett Opus 4.8-anker | 2.1.154, 2.1.170 | ✅ DONE 2026-06-18 | update-knowledge | knowledge/prompt-cache-patterns.md | H | M |
| token management | CC truncerer skill-listing ved budsjett-overflow; config-audit kan varsle før | 2.1.178 | partial | new-scanner | knowledge/prompt-cache-patterns.md | M | M |
| token management | `skillListingBudgetFraction`/`maxSkillDescriptionChars` (UVERIFISERT, under gulv) | ~2.1.105 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| slash commands | **`/config key=value`** — in-session-anvendelse av settings-anbefalinger (kan refereres i fix/feature-gap-tekst) | 2.1.181 | missing | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| slash commands | `/simplify``/code-review` (stale bundled-skill-liste) | 2.1.147 | now-wrong | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| slash commands | `/goal`, `/reload-skills` ikke i bundled-command-inventar | 2.1.139, 2.1.152 | missing | update-knowledge | knowledge/claude-code-capabilities.md | L | S |
| subagents/agents | settings `agent`-felt (kjør main-thread som named subagent) ikke i kjent-sett | 2.1.157 | missing | extend-scanner-known-set | knowledge/claude-code-capabilities.md | M | S |
| subagents/agents | nested subagents (5-nivå), closest-`.claude`-precedence | 2.1.172, 2.1.178 | missing | update-knowledge | knowledge/claude-code-capabilities.md | M | S |
| subagents/agents | `availableModels`/`enforceAvailableModels` begrenser agent-`model:` | 2.1.172 | partial | update-knowledge | knowledge/claude-code-capabilities.md | M | M |
| subagents/agents | `defaultEnabled: false` i plugin.json (gater bundled agents/skills/commands) | 2.1.154 | missing | extend-scanner-known-set | knowledge/claude-code-capabilities.md | M | S |
_(Resten av de 32 radene er `none`/`covered` CLI/UI-runtime uten config-surface — full liste i `tasks/wyk5istx7.output`.)_
## Draft-korreksjoner (re-run fanget versjonsfeil i `cc-2.1.x-changelog-delta.md`)
Draftet (og dermed noen rader i første kjøring) hadde feil versjon på fire punkter — matrisen over er korrekt kilde:
- settings `agent`-felt: **2.1.157** (ikke 2.1.154)
- agent-spawn i auto-mode: **2.1.178** (ikke 2.1.143)
- `subagent_type` case/separator-insensitiv: **2.1.140** (ikke 2.1.157)
- `SubagentStop` `additionalContext`: **2.1.165** (ikke 2.1.163)
Disse korrigeres når knowledge-filene oppdateres (Batch 3). `docs/cc-2.1.x-changelog-delta.md` er nå superseded av denne matrisen.

View file

@ -0,0 +1,45 @@
# Plan — Claude Code 2.1.114→181 dekningsgjennomgang
**Ratifisert:** 2026-06-18 · **Scope:** kun `config-audit/` · **Metode:** A (orkestrert gap-matrise + verifikasjon per surface)
## Mål
Gå systematisk gjennom hver konfigurasjons-relevante Claude Code-endring i vinduet **v2.1.114 → v2.1.181** (~67 releaser), avgjøre om config-audit må endre hva den skanner/anbefaler, og produsere en prioritert tiltaksplan. Deretter (egen fase) token-optimalisering.
## Bakgrunn / hvorfor nå
Pluginens kunnskapsbase ble sist verifisert ved **v2.1.114** (`knowledge/claude-code-capabilities.md:4`). Installert CC er **2.1.181**. Verifisert staleness ved oppstart:
- `settings-validator.mjs:66` `VALID_EFFORT_LEVELS` mangler `xhigh` (CC 2.1.154) → **false positive** på gyldig konfig.
- Hook-events `post-session`, `MessageDisplay` udokumentert i `knowledge/hook-events-reference.md`.
- «Opus 4.7» hardkodet i `token-hotspots.mjs`, `opus-4.7-patterns.md`, `configuration-best-practices.md` — 4.8 er default, Fable 5 finnes.
- `feature-evolution.md` (staleness-tracker) stopper ved v2.1.111.
- Nyere settings-nøkler ukjente for validatoren (`fallbackModel`, `enforceAvailableModels`, `disableBundledSkills`, `skillOverrides`, `sandbox.allowAppleEvents` …).
## Versjonsgulv
**2.1.114** (pluginens faktiske sist-verifiserte punkt), ikke 2.1.130 — ellers mistes 2.1.115129.
## Metode (ratifisert: A)
Hovedkontekst orkestrerer; én verifikasjons-agent per config-surface, mekanisert som dynamisk workflow `config-audit-cc-gap`. Eskalering til Voyage (`/trekbrief→/trekplan`) **kun hvis** Fase 2 avdekker netto-nye scannere store nok til å forsvare det.
## Faser
- **Fase 0 — Lås changelog-korpus.** Re-verifiser mot offisiell `CHANGELOG.md` for hele vinduet, per surface, med kildelenker. *Leveranse:* `docs/cc-2.1.x-changelog-delta.md` (draft → verifisert i Fase 2).
- **Fase 1 — Dekningsbaseline.** Map hver scanner + knowledge-fil til surface + kjent-sett (settings-nøkler, hook-events, 25 feature-checks, mcp-felter, token-mønstre).
- **Fase 2 — Gap-matrise (kjernen).** Kryss delta × dekning. Hver rad: `surface → CC-endring → dekket/delvis/mangler/nå-feil → tiltakstype → mål-fil → prioritet H/M/L`. *Leveranse:* `docs/cc-2.1.x-gap-matrix.md`.
- **Fase 3 — Prioritert tiltaksplan + release-skisse.** Quick wins (xhigh, kjent-nøkler, hook-events, 4.7→4.8) vs større (sandbox/managed-settings, `Tool(param:value)`-permissions, plugin `defaultEnabled`). Avgjør v5.2 vs 6.0. Ratifiseres FØR kode.
- **Fase 4 (etter, egen tråd) — Token-optimalisering.** Refresh `opus-4.7-patterns` → 4.8/Fable 5; nye token-styrings-settings som TOK/feature-gap-checks.
## Tiltakstyper (klassifisering i matrisen)
`update-knowledge` · `extend-scanner-known-set` · `new-feature-gap-check` · `new-scanner` · `fix-false-positive` · `none`.
## Verifisering (testbare kriterier)
- Hver HIGH-impact-endring i vinduet har **én rad** i gap-matrisen med verdikt + kildelenke (#HIGH = #rader).
- Regresjonstest: fixture med `effortLevel: "xhigh"` flagges **ikke** (beviser false-positive fikset).
- `node --test 'tests/**/*.test.mjs'` grønn (baseline 792) etter hver scanner-endring.
- `node scanners/self-audit.mjs` ren på pluginen selv.
- `feature-evolution.md` siste rad = v2.1.181.
- Ny changelog-scout finner **0** netto-nye uoppdagede surfaces (completeness-sjekk).
## Scope-gjerde
Kun `config-audit/`. Fase 03 = analyse/planlegging. INGEN produksjonskode-endring uten ratifisert tiltaksplan + failing test først (Iron Law). Fase 4 startes ikke før 03 er ratifisert.
## Risiko
- Agent-draft kan inneholde feil/items under gulvet → Fase 2 re-verifiserer hver rad mot offisiell kilde.
- Web-fetch utilgjengelig i headless workflow-agenter → graceful fallback til draft, rader merkes `web_verified:false`.

View file

@ -0,0 +1,400 @@
# Brief — Delete-and-Rebuild (config subtraction)
**Status:** BRIEF, not a plan. No code, no chunk breakdown, no version number committed.
Written 2026-07-29 so a session starting Thursday evening has a durable starting point.
STATE.md is gitignored in this repo, so this file — not STATE — is the record.
---
## 1. Trigger and its provenance
The operator relayed a third-party YouTube summary (Hyper Automation Labs) of a talk
Boris Cherny reportedly gave at Y Combinator Startup School, one day after Opus 5
shipped. Claims attributed to him in that summary:
- Anthropic deleted ~80 % of Claude Code's own system prompt when Opus 5 landed.
- Advice to users: every six months, delete your CLAUDE.md, your skills, your hooks —
see what the model does.
- Rebuild method: delete everything, use it, add one line back only when the model
stumbles on the same thing repeatedly.
- The model measured *slightly more intelligent* with the built-in prompts stripped.
**Two levels of confidence here, and they must not be collapsed** (updated 2026-07-29
after the operator corrected the first draft):
- **Attribution — confirmed.** The operator watched the recording and identifies Boris
Cherny on stage. This is not a channel's claim about who spoke; it is direct
observation by the operator. The talk happened and it is him.
- **The verbatim figures — still summary-level.** "80 % of the system prompt", "the
model measured slightly more intelligent without the prompts", the Bun numbers: these
reach us through the channel's editing, not through a primary transcript. They are
plausible and consistent with §2, and they are not quoted as fact anywhere below.
So if this becomes a `BP-*` entry in the best-practices register, the source field reads
*"Boris Cherny, YC Startup School talk (attribution confirmed by operator); figures via
third-party summary, not primary-verified"* — not "unverified", and not a bare citation
either. Getting a primary transcript for the figures is a nice-to-have, never a
prerequisite: the feature is argued from this repo's own logic (§3), and this repo has
one scar from treating a plausible quote as load-bearing fact («Fable low ≈ Opus high»,
fabricated, rejected 2026-07-14).
**What the operator has affirmed as in scope:** delete CLAUDE.md, then rebuild by adding
back what the model needs help with — *starting with what it must have*. That last clause
is not a detail; it is the design constraint in §6.0.
## 2. Verified ground truth (checked against the local CLI, 2026-07-29)
These *are* facts, and they are what make an empirical variant buildable:
| Fact | How verified |
|---|---|
| `claude --bare` exists — "Minimal mode: skip hooks, LSP, plugin sync, attribution, auto-memory, background prefetches, keychain reads, and CLAUDE.md auto-discovery. Sets `CLAUDE_CODE_SIMPLE=1`." | `claude --help` |
| `claude --system-prompt <prompt>` — replaces the session system prompt | `claude --help` |
| `--append-system-prompt`, `--add-dir`, `--setting-sources <user,project,local>`, `--plugin-dir`, `--settings`, `--agents` — all present and all usable to compose an isolated config for a controlled run | `claude --help` |
So the `CLAUDE_CODE_SIMPLE=1` env var the video calls "undocumented" is, in this CLI
version, a documented flag (`--bare`). That matters: an A/B ablation harness would not
need an undocumented hook.
## 3. Why this fits this repo (the argument that does *not* depend on §1)
The plugin has three pillars — Health, Opportunities, Action. Every existing command
answers a question on the **addition** axis:
- `feature-gap` — what could you add?
- `optimize` — what would fit a better mechanism?
- `posture` / `tokens` / `manifest` — how good/expensive is what you have?
- `fix` / `implement` — apply changes.
**Nothing answers the subtraction question: what is no longer earning its rent?**
`feature-gap` has no inverse. That is a real hole, independent of who said what on
a stage.
Two further reasons this belongs *here* specifically:
1. **Nothing in the repo measures instruction age.** Verified by grep over `scanners/`:
`stale` appears only for knowledge-register entry age (`lib/knowledge-refresh.mjs`)
and for stale plugin-cache versions (`scan-orchestrator.mjs`, M-BUG-11). No scanner
touches `git blame`, `mtime`, or the vintage of a CLAUDE.md block. An instruction
written for Sonnet 3.5 and an instruction written last week are indistinguishable to
every current scanner.
2. **Deleting bravely is only sane if you can undo it.** That is already pillar three:
`lib/backup.mjs`, `rollback-engine.mjs`, `auto-backup-config.mjs` (PreToolUse), and
`drift`'s `saveBaseline` / `loadBaseline` / `diffEnvelopes`. The safety net exists;
the feature that would use it does not. This is arguably the strongest framing:
*config-audit is already the infrastructure that makes "delete it and see" a
measurement rather than a gamble.*
## 4. Reusable machinery (do not rebuild these)
| Need | Already exists |
|---|---|
| Snapshot config before deleting | `scanners/lib/baseline.mjs` (`saveBaseline`), used by `drift` |
| Diff before/after | `diffEnvelopes` in the same module |
| Backup + restore individual files w/ sha256 manifest | `scanners/lib/backup.mjs`, `scanners/rollback-engine.mjs` |
| Cost of each source, always-loaded subtotal | `scanners/manifest.mjs`, `token-hotspots.mjs` |
| Human-approved-writes command pattern | `knowledge-refresh`, `campaign` (both non-byte-stable by design) |
## 5. Candidate shapes (sketches — pick on Thursday, do not pre-commit)
All three inherit the floor constraint in §6.0: whatever the shape, load-bearing local
facts are never deletion candidates, and a rebuild restores them first. A shape that
cannot express that distinction is disqualified regardless of how cheap it is.
**A. Deterministic vintage scanner (`CA-VIN-*`).** Per instruction block in CLAUDE.md /
rules / skills / hooks: age from git history, plus a compensatory-phrasing signal
(blocks that exist to correct model behaviour — "ALWAYS", "never forget", "read the
whole file first", "don't guess"). Output: ranked deletion candidates with age + token
cost + why it looks compensatory. Cheapest, most testable, fits the existing scanner
architecture, and composes with `tokens`/`manifest` for the payoff figure.
> **🔴 The age half of this shape is DEAD — measured 2026-07-31 (økt #40), see §8's
> updated checklist.** Per-line `git blame` over four real instruction files gives
> single-date shares of 88 % / 55 % / 100 % / 58 %: instruction blocks trace back to bulk
> commits, so per-block age carries almost no information. Worse, `blame` reports
> *last touch*, not vintage — a reformatting commit (e.g. this repo's own `96e32df`,
> "trim project CLAUDE.md to invariants") makes old instructions look young, which biases
> the signal rather than merely thinning it. That also kills the `mtime` fallback.
> **§8 pre-registered this exact outcome and its consequence: shape A's primary signal
> collapses to the compensatory-phrasing heuristic alone, and §7.2 is answered — the
> classification needs the agent layer.** Do not resurrect age as a "weak secondary
> signal"; that is the drift the pre-registration existed to prevent. What survives of
> shape A is the *deterministic phrasing + token-cost pre-filter*, feeding a
> precision-gated judge (the `optimize` architecture).
**B. Protocol command (`/config-audit rebuild`).** The delete → live with it → earned
re-add ledger loop, spanning sessions: archive current config, record what was archived,
and maintain a ledger where a line only returns when the operator records that the model
actually stumbled on it. Highest fidelity to the source idea; needs cross-session state
(the `sessions/` machinery already exists) and is inherently not byte-stable.
**C. Measured ablation harness.** Use `--bare` + `--system-prompt`/`--add-dir` to run the
same prompt with and without a block and compare. Closest to a real verifier, most
expensive in tokens, weakest determinism. Probably a later `--experimental` sub-mode of
A or B rather than its own command.
Likely landing: **A first** (deterministic, testable, immediately useful), with B as the
workflow that consumes A's output. C stays a documented idea until A+B exist.
> **DECIDED 2026-07-31 (#40), superseding the sketch above — see §7 q2/q3.** Shape A does
> not ship as a standalone `CA-VIN-*` scanner: its age signal is dead, and what remains of
> it (a phrasing + token-cost pre-filter) is precisely the front half of `optimize`'s
> existing hybrid motor. **The landing is `/config-audit optimize --subtract`** — a fourth
> `lensCheck` class emitting `CA-OPT-*` findings, with a deterministic **floor-exclusion**
> step ahead of the judge so a load-bearing block is never a candidate. Shape B (the
> earned-re-add ledger) stays out of v1 precisely because it is what would demand
> cross-session state and therefore a command of its own. Shape C unchanged: documented,
> not built.
## 6. Hard constraints (carry these into any implementation)
### 6.0 The floor: compensatory vs load-bearing instructions
**This is the invariant all three shapes in §5 must respect, and it is what makes the
feature safe to ship at all.** Not every line in a CLAUDE.md is the same kind of thing:
| Class | What it is | Test | Disposition |
|---|---|---|---|
| **Compensatory** | An instruction correcting model *behaviour* — "read the whole file first", "don't guess", "ALWAYS verify", "think before you code" | A smarter model would do this unprompted | **Deletion candidate.** Returns only when earned. |
| **Load-bearing** | A *local fact* the model cannot derive at any capability level — "never GitHub, only Forgejo at git.fromaitochitta.com", "system bash is 3.2, no `declare -A`", "test with `node --test 'tests/**/*.test.mjs'`", "`~/.claude` is not git-tracked" | No amount of intelligence produces this from the codebase alone | **Floor. Never a deletion candidate.** |
Model capability erodes the first class and does nothing to the second. That is the whole
mechanism behind the source idea — and it means "delete your CLAUDE.md" is only correct
for one of the two classes. A tool that treats them alike would delete the operator's
Forgejo constraint because Opus 5 "is smart enough now", which is a category error: the
model isn't failing at intelligence there, it simply cannot know.
**Consequence for the rebuild ordering.** A rebuild is not one undifferentiated
add-back-on-stumble loop. It is three tiers:
1. **Floor — goes back immediately, no trial period.** Load-bearing facts. The config is
never in a state where these are absent; a "delete everything" that drops them is a
broken experiment, not a brave one.
2. **Earned — out, returns only on repeated observed stumbling.** Compensatory
instructions that turn out to still be needed by *this* model.
3. **Dead — out and never missed.** The payoff, measurable in tokens via `manifest`.
A third class sits deliberately outside the axis: **policy prohibitions** ("never commit
secrets"). These may well be things the model would honour unprompted, but their cost of
being wrong is asymmetric and they are cheap. They stay in the floor by decision, not by
classification. Do not let a "the model knows this now" argument reach them.
The hard part is tier 1 vs tier 2, and that classification — not the deletion mechanics —
is the real engineering problem in this brief. Precision is asymmetric: a missed dead line
costs a few tokens per turn; a deleted load-bearing line costs a wrong remote, a broken
bash script, or a lost afternoon.
### 6.1 Existing repo constraints
- **`~/.claude` IS git-tracked as of 2026-07-26 — but archive by `mv` into `_archive/`
anyway, never `rm`.** Premise corrected 2026-07-31 (økt #40); the rule it was used to
justify is unchanged. Ground truth: `/Users/ktg/.claude` is a git repo, 7 commits,
initial commit `13cd708` 2026-07-26, remote `backup
/Volumes/DharmaBackup/claude-config.git` (external volume — not Forgejo, not GitHub),
47 files tracked (`hooks/` 19, `commands/` 14, `scripts/` 9, `CLAUDE.md`,
`settings.json`, `learnings/`, `docs/`). The rule survives on different grounds: only
those 47 files are recoverable, the history is days old, and the remote lives on a
volume that may not be mounted. Two consequences for this feature — (a) `~/.claude`
age evidence is bounded by the repo's own birth date and is therefore worthless, and
(b) any untracked file there (`memory/`, `coord/`, session state) is still
unrecoverable after `rm`.
- **`settings.json` is pathguard-protected** — no Edit/Write. Use `jq` + temp file +
atomic `mv`, with operator OK ([[settings-json-pathguard-write]]).
- **New scanner ⇒ the 7-step byte-stability checklist** ([[adding-scanner-byte-stability]]):
scanner + orchestrator + scoring area-map + strip-added-scanner + humanizer count +
SC-5 + humanizer wiring (`SCANNER_TO_CATEGORY` entry and `TRANSLATIONS.static` per RAW
title — M-16/M-17). Frozen `tests/snapshots/v5.0.0/` must stay untouched.
- **TDD is inviolable** — red tests before implementation, including for .md contracts.
- **`feat:` commits require non-trivial diffs in both README.md and CLAUDE.md**
([[docs-gate-feat-requires-readme-claudemd]]).
## 7. Open questions for Thursday
> **STATUS 2026-07-31 (#40): q1, q2, q3 are CLOSED below — decided on measured evidence
> (dead age signal) plus the hand-built fasit (`docs/subtraction-fasit.local.md`), under
> the operator's standing delegation of technical/design calls. q4 was already closed
> 2026-07-29. Do not re-litigate without new evidence.**
>
> **The landing: `/config-audit optimize --subtract` — a fourth lens class in the existing
> hybrid motor, not a new command and not a new scanner.** Full rationale in q3.
1. Scope of the deletion candidate set: user-level `~/.claude` only, project only, or both?
(Age evidence is much stronger for project-level, per §5A.)
**CLOSED — both, and user-level is mandatory in v1.** The argument that pointed at
project-level was git-age evidence, and that evidence no longer exists (§5A box), so
scope now follows *payoff and testability* instead. Two things force user-level in:
§8's blocking floor test is defined over the operator's global CLAUDE.md, and that file
is where the always-loaded cost actually sits (~4 300 tok every turn, every repo,
every session — vs a project CLAUDE.md that loads only in its own repo). Project-level
comes along because the same motor reads it and because two of §8's four floor anchors
turned out to live there.
*Implementation note, not a reopening:* `optimize` today takes a repo `target`. Reading
`~/.claude/CLAUDE.md` is a target-resolution change, and per §6.1 the `~/.claude` write
rules apply — but this mode proposes, it never writes.
2. **The core question, given §6.0:** can compensatory-vs-load-bearing be classified
deterministically with acceptable precision, or does it need the agent layer (like
`optimize`'s precision-gated `optimization-lens-agent`, which stays silent when
unsure)? Note the two signals are independent — a load-bearing fact can be old, and a
compensatory instruction can be new — so age alone can never carry this call. Prior:
a deterministic pre-filter feeding a precision-gated judge, which is exactly the
`optimize` architecture already in the repo.
**Design the §8 fasit around the ambiguous middle, not the poles.** The four named
must-survive items are clear-cut and any mechanism will get them right; the gate is
really decided by blocks like "Conventional Commits: `type(scope): beskrivelse`"
(local convention, or a nag the model would follow anyway?), "commit ofte med
beskrivende meldinger" (pure behaviour correction?), or the model-routing rubric
(a table of local policy that reads like advice). Include 35 such blocks
deliberately. A fasit built only from obvious cases will pass a tool that fails on
real config.
**CLOSED — it cannot be done deterministically. Deterministic pre-filter → precision-
gated judge, which is the `optimize` architecture verbatim.** Two independent lines of
evidence, both gathered before any code:
- *The age signal is gone* (§5A box). It was the only deterministic input that was
going to do real discriminating work; phrasing heuristics alone are what remain.
- *The fasit shows phrasing alone is not enough.* The blocks that decide the gate all
require reading **content against container** — a load-bearing fact wearing generic
phrasing, or vice versa:
| Fasit block | The trap |
|---|---|
| B-27 (defensive shell-scripting) | Reads as universal advice ("quote your variables"), but each line records a *specific local incident*, and "keep test code ASCII-clean" is inseparable from bash 3.2 (B-26, a §8 floor anchor). **A phrasing regex deletes this and fails the gate.** |
| B-32b ("never work in another repo") | Sits inside a bullet list of compensatory anti-patterns, formatted identically to its neighbours, but encodes the polyrepo structure and names `coord-send`. Container says delete, content says floor. |
| B-05 vs B-06 | Adjacent preference bullets, near-identical form, opposite calls. |
| B-30a / B-30b / B-30c | Three consecutive bullets under one heading, three different calls (Conventional Commits is floor, "lesbarhet > cleverness" is not). |
No regex separates these; a judge reading them in context can.
- *Two failure modes the judge must be prompted against, both found in the fasit:*
**staleness is not a deletion signal** (B-29 pins Opus 4.8 while the session runs
Opus 5 — that is a `drift`/dead-reference finding about a **floor** block), and
**tier-2 is not tier-3** (B-22, premiss-verifisering, is compensatory by class yet
earned itself again during this very session).
**Design consequence — the floor is NOT the judge's call.** Load-bearing exclusion runs
as a deterministic pre-filter step *before* the judge sees anything, so a floor block is
never a candidate at all. §8's gate is blocking and precision is asymmetric; making it
depend on a probabilistic judge would be the wrong guarantee. The judge then decides
only compensatory-tier questions on the remainder, and stays silent when unsure exactly
as `optimization-lens-agent` already does.
3. Does this ship as its own command, or as a `--subtract` mode of `optimize`?
Both are defensible; command count is already 21.
**CLOSED — a `--subtract` mode of `optimize`.** The discriminator is whether v1 needs
cross-session state: it does not. v1 ranks deletion candidates and reports the payoff;
the earned-re-add **ledger** (shape B) is what would require `sessions/` state, and it
is deliberately not in v1. Without the ledger there is nothing a separate command buys.
What the mode inherits by not being new (verified against the code, 2026-07-31):
| Requirement | `optimize` already has it |
|---|---|
| Deterministic pre-filter → opus judge | `optimization-lens-scanner.mjs` (`lens-prefilter`) → `optimization-lens-agent` |
| Precision gate, silent when unsure | Stated in the agent's own prompt |
| "Not a mistake" framing | Every `optimize` finding is a *Missed opportunity* — exactly right for a line that works but no longer earns its rent |
| Register-backed provenance | `knowledge/best-practices.json`, CONFIRMED-only |
| Non-byte-stable by design | Already documented as such in CLAUDE.md |
| Finding IDs | `CA-OPT-*`**no new `CA-VIN-*` scanner** |
And what it avoids: the 7-step byte-stability checklist (§6.1), a scanners badge bump
16 → 17, snapshot risk against frozen `tests/snapshots/v5.0.0/`, and a 22nd command.
**Proportionality is the clinching argument.** The fasit puts the honest ceiling at
~8501 400 always-loaded tokens on a ~4 300-token file — real (~20 %), but nowhere near
the source anecdote's 80 %, and 26 of 34 blocks are floor. On a well-maintained config
the subtraction axis is mostly a no-op. That payoff justifies a mode on an existing
motor; it does not justify a new scanner plus a new command. ("Starte ambisiøse tiltak
når en konfig-justering holder" is a named anti-pattern.)
Registry shape: a 4th `lensCheck` class alongside the three in the agent's table, with
its own `BP-SUB-001` register rule. `--subtract` is the flag because the subtraction
axis should not fire on a plain `/config-audit optimize` run — it asks a different
question and needs the operator to have opted into it.
4. ~~Ordering against the existing queue.~~ **Decided 2026-07-29:** the operator
prioritized this work ahead of pipeline step 4 (`rollback`), to start Thursday
2026-07-30. Do not re-litigate. The open part is only what follows it — the
prior order stands underneath: step 4 `rollback`, then the M-11→M-20 batch
release, then the v5.13 plan.
## 8. Verifisering (testable criteria)
**Before building anything:**
- [x] **DONE 2026-07-31 (#40).** `grep -rniE 'blame|mtime|birthtime' scanners/` returns
zero hits → §3.1 confirmed still true at build time.
- [x] **DONE 2026-07-31 (#41).** `node scanners/drift-cli.mjs . --save --name pre-subtraction`
succeeds and `lib/baseline.mjs` round-trips → §4 reuse is real, not assumed.
Written envelope is `{meta, scanners[16], aggregate, _baseline}` with
`_baseline.target_path` = the repo (15 findings, score 55). Run with no flags
beyond `--save --name`, since M-BUG-21 makes an unknown flag's value silently
become the scan target.
- [x] **RESOLVED 2026-07-31 (#41) by not needing it.** `BP-SUB-001` was written
`confirmed`, and asserts **nothing** from §1 — no "80 %", no ablation figure, no
Cherny attribution. Its claim is grounded entirely in the Anthropic steering blog
already cited by BP-MECH-001004, re-verified the same day: *"Every line loads into
every session for every engineer working in the repo, whether it's relevant to their
task or not. This consumes tokens and dilutes adherence"*, *"Build commands,
directory layout, monorepo structure, coding conventions, and team norms all fit
naturally here"*, and *"Keep CLAUDE.md under 200 lines"*. The anecdote motivated the
feature; it is not a premise of the shipped rule.
- [x] **DONE 2026-07-31 (#40) — THE AGE SIGNAL DOES NOT EXIST. Collapse condition met.**
Per-line `git blame` (not `git log`; per-line is the question) over four real
instruction files:
| File | Lines | Largest single-date share |
|---|---|---|
| `~/.claude/CLAUDE.md` | 250 | **88 %** (2026-07-26 = repo init; max age 5 days) |
| `config-audit/CLAUDE.md` | 102 | **55 %** (2026-04-08) |
| `config-audit/.claude/rules/ux-rules.md` | 32 | **100 %** (one commit) |
| `llm-security/CLAUDE.md` | 110 | **58 %** (2026-04-08) |
Blocks trace back to bulk commits exactly as the pre-registration feared, and
`blame` measures last-touch rather than vintage (a reformat resets it), so the
signal is *biased*, not merely sparse. **Consequence, as pre-registered: shape A's
primary signal is gone and §7.2 is answered — deterministic phrasing/cost
pre-filter → precision-gated judge.** See the boxed note in §5A.
- [x] **DONE 2026-07-31 (#40): the §8 floor-test fasit is built, before any classifier
exists.** `docs/subtraction-fasit.local.md` (LOCAL-ONLY — it quotes the operator's
global CLAUDE.md verbatim and this repo's only remote is the public `open/` mirror).
34 blocks over `~/.claude/CLAUDE.md`, each labelled `FLOOR` / `POLICY-FLOOR` /
`DELETABLE`, with 13 marked ⚠ AMBIGUOUS per §7.2. Headline numbers: 26 of 34 blocks
are floor; the deletable set is ~1 400 always-loaded tokens, realistically ~850
after tier-2 earn-backs, against a ~4 300-token file (**~20 %, not 80 %**).
**Correction it forces:** two of §8's four named floor anchors are not in that file
at all — the test command is project-scope (`config-audit/CLAUDE.md`), and
"`~/.claude` is not git-tracked" is now false (§6.1). The floor gate must run over
a *set* of files, not one.
**When `optimize --subtract` is built** (was: "if shape A is built" — retitled 2026-07-31
per §7 q3; the two struck items below were premised on a new standalone scanner, which is
no longer the shape):
- [x] **DONE (#41).** Red tests written first against a synthesized fixture and confirmed
failing (module not found) before either module existed.
- [x] **DONE (#41).** Suite green at **1382/0** (was 1365).
- [x] **DONE (#41).** `git diff --stat tests/snapshots/v5.0.0/` empty, **and** a plain
`optimize-lens-cli.mjs` run diffed byte-identical against its pre-change output
(`--subtract` adds keys only when the flag is present; `LENS_DETECTORS` still
exposes exactly 3 detectors, asserted in the suite).
- [ ] ~~`node scanners/self-audit.mjs --check-readme` passes (badge counts updated:
scanners 16 → 17).~~ **No badge bump — no new scanner.** `--check-readme` must still
pass, and README/CLAUDE.md still need the docs-gate diffs for a `feat:` commit
([[docs-gate-feat-requires-readme-claudemd]]).
- [ ] Every new finding renders with a non-`Other` `userImpactCategory` and a
non-`_default` action language → humanizer wiring correct (M-16/M-17).
- [x] **PASSED 2026-07-31 (#41) — the blocking floor test.** Fasit built first (#40), tool
run after, and the comparison **machine-checked**: a script asserts the intersection
of the candidate list with every FLOOR / POLICY-FLOOR line range in the fasit is
empty. Result: **zero load-bearing blocks proposed for deletion.** Both §8 anchors
present in the subject file (B-25 "Aldri GitHub. Kun Forgejo", B-26 "System bash er
3.2") are excluded, as are B-09, B-13, B-17, B-23, B-27, B-29, B-32b and B-38b.
The first run found **five** violations the synthesized fixture missed; each was
fixed with a structural rule (list-stem merge, ordered-list-as-contract, security
terms, `unresolved-entity`, declarative guard) and a fixture shape added so it
cannot regress.
- [x] **MET (#41), with the number stated honestly.** 11 of 18 deletable groups surfaced,
≈756 always-loaded tokens ≈ **18 %** of the ~4 300-token file — inside the
pre-registered ~850 / ~20 % band. The 7 misses are the conservative default working
as intended (unresolvable entity names, code spans, and declarative/infinitive
phrasing carrying no imperative). Precision over recall held throughout: every fix
in this session traded recall away, never the gate.

View file

@ -0,0 +1,104 @@
# Brief — Devil's Advocate: motbevis at CC 2.1.114→181-feilrettingen er komplett
**Skrevet:** 2026-06-18 · **For:** NESTE sesjon (etter `/clear`) · **Scope:** kun `config-audit/` · **Metode:** Dynamic Workflow (adversariell fan-out)
## Hvorfor denne sesjonen finnes
Forrige sesjon konkluderte: «den akutte feilrettingen er ferdig; gap-matrisen er stale, alle 5 false-positive-klynger er fikset.» Den konklusjonen hviler på **spot-sjekk av 5 klynger + git-historikk**, IKKE på en uttømmende rad-for-rad-verifisering av matrisens 130+ rader. Operatøren bestilte derfor en **devil's advocate**: en uavhengig, skeptisk part som aktivt prøver å **motbevise** «alt er lukket». Default-holdning: en fiks er IKKE bekreftet før kode + test beviser den korrekt OG komplett.
## Hva som skal angripes (3 påstander)
1. **P1 — Ingen aktive false positives igjen.** Hver matrise-rad med `verdict=now-wrong` / `action=fix-false-positive` er faktisk fikset i dagens kode.
2. **P2 — Fiksene er korrekte, ikke overfladiske.** Ikke bare «nøkkel lagt i KNOWN_KEYS», men riktig TYPE_CHECK/enum/struktur. (Presedens: `skillOverrides` skulle IKKE i TYPE_CHECKS fordi verdien er objekt — en naiv `'string'`-sjekk ville skapt NY false positive. Let etter den slags.)
3. **P3 — Denne ukens nye arbeid er regresjonsfritt.** `CA-SKL-002` (aggregat-budsjett) + HOME-leak-klasse-fiksen har ingen skjulte feil.
**Eksplisitt UTENFOR scope:** rader merket `missing` / `new-feature-gap-check` / `new-scanner` er *manglende kapabiliteter*, ikke bugs. Devil's advocate skal IKKE rapportere dem som funn — MEN skal flagge hvis en `missing`-rad i virkeligheten er en feilmerket `now-wrong` (bug i forkledning).
## Hvorfor Dynamic Workflow (ikke solo-kontekst)
Matrisen spenner ~12 config-surfaces × 130+ rader. Uttømmende verifisering mot scanner-kode + tester er en fan-out-oppgave: uavhengige skeptikere per surface gir bredde + uavhengighet som én kontekst ikke klarer uten å skumme. Mønster: **review → find → adversarially verify → synthesize**. Workflow-verktøyet er forhåndsautorisert av operatøren for DENNE sesjonen (denne forespørselen = opt-in); bekreft kort med operatør før du brenner tokens, kjør så.
## Surface-klynger å fane ut over (fixed-claims å motbevise)
Ankret i `docs/cc-2.1.x-gap-matrix.md` + git. Hver klynge = én skeptiker-agent:
| Klynge | Mål-fil | Påstått fikset i | Angreps-vinkel |
|---|---|---|---|
| settings-keys + `xhigh` | `scanners/settings-validator.mjs` | `7309935` | Er ALLE ~11 nøkler + `xhigh` til stede? Er TYPE_CHECKS riktige (array vs string vs objekt)? Finn en gyldig CC 2.1.181-config som fortsatt flagges. |
| hook-events | `scanners/hook-validator.mjs` | `98ddd77` | `MessageDisplay`+`post-session` i `VALID_EVENTS`? Count = 28? Andre nye events (2.1.114181) glemt? |
| MCP env-vars + `trust` | `scanners/mcp-config-validator.mjs` | `4b94da0`,`b3c572a` | `CLAUDE_PROJECT_DIR`-allowlist + POSIX `${var%}`-skip korrekte? Shell-`${VAR}` (2.1.161)? Er `trust` *helt* borte (kode+humanizer+knowledge+fixturer)? |
| permissions / DIS + CNF | `scanners/disabled-in-schema-scanner.mjs`, `conflict-detector.mjs`, `lib/permission-rules.mjs` | `bec3f45` | Er `Tool(param:value)` / `WebFetch(domain:*)` param-bevisst i BEGGE scannere? Konstruer deny/allow-par som FØR kollapset — fyrer det fortsatt feil? |
| claude-md-lengde | `scanners/claude-md-linter.mjs` | `feaa7ed` | Er HIGH@500 reframet til MEDIUM token-kost? Finnes det en gjenværende `SEVERITY.high`-sti for ren lengde? |
| knowledge-korpus | `knowledge/*` | `624f5ed` | `/simplify``/code-review`, Opus 4.8-default, Fable 5, siste `feature-evolution`-rad = v2.1.181? Gjenværende «4.7»/«v2.1.111»-staleness? |
| token / skill-listing | `scanners/skill-listing-scanner.mjs`, `token-hotspots.mjs` | `8376dab`,`7bb2547`,`66433fe` | (se P3) |
## Workflow-skisse (script-skjelett — neste sesjon authorer/kjører via `Workflow`)
Hybrid: hovedkontekst skummer matrisen inline → bygger klynge-listen → fan-out skeptikere. Pipeline der mulig (per klynge uavhengig), barrier kun før syntese.
```js
export const meta = {
name: 'devils-advocate-gap-verification',
description: 'Adversarially disprove that the CC 2.1.114-181 feilretting is complete',
phases: [
{ title: 'Refute', detail: 'one skeptic per scanner surface — try to find a still-open or wrongly-fixed row' },
{ title: 'Session', detail: 'adversarial review of CA-SKL-002 + the HOME-leak fix' },
{ title: 'Synthesize', detail: 'punch-list of genuinely-open items, or a verified attestation' },
],
}
const VERDICT = { type:'object', additionalProperties:false,
required:['surface','rows','overallVerdict'],
properties:{
surface:{type:'string'},
rows:{type:'array', items:{type:'object', additionalProperties:false,
required:['claim','actualStatus','verdict','evidence'],
properties:{
claim:{type:'string'}, // the matrix row / fixed-claim
actualStatus:{type:'string'}, // what the code actually does now
verdict:{enum:['confirmed-fixed','still-open','incorrectly-fixed','uncertain']},
evidence:{type:'string'}, // file:line + the proving/refuting test or config
counterexample:{type:'string'} // a config that still false-positives, if found
}}},
overallVerdict:{enum:['all-closed','has-open-items']}
}}
// CLUSTERS: hardcode from the table above (surface, targetFiles, matrixRows) — keeps it deterministic.
const CLUSTERS = [ /* settings, hooks, mcp, permissions, claude-md, knowledge, token */ ]
phase('Refute')
const refutations = await parallel(CLUSTERS.map(c => () =>
agent(
`DEVIL'S ADVOCATE. Surface: ${c.surface}. Target file(s): ${c.targetFiles}.\n` +
`For each matrix row claimed fixed (${c.matrixRows}), READ the scanner code AND its test, then TRY TO REFUTE ` +
`that the fix is correct and complete. Construct a valid CC 2.1.181 config that SHOULD pass and check whether ` +
`the scanner still flags it; check TYPE_CHECK/enum/struct correctness (skillOverrides precedent: object, not string). ` +
`Default each row to 'still-open'/'uncertain' unless code+test prove 'confirmed-fixed'. Cite file:line. Do NOT report ` +
`'missing'/feature-gap rows as findings (out of scope) unless one is a now-wrong bug mislabeled.`,
{ label:`refute:${c.surface}`, phase:'Refute', schema:VERDICT, agentType:'Explore' }
)))
phase('Session')
const sessionReview = await parallel([
() => agent(`Adversarially review CA-SKL-002 in scanners/skill-listing-scanner.mjs + its tests. Hunt: off-by-one at the ` +
`4000-tok boundary, the min(desc,1536) capping logic, ID-ordering assumptions, calibration-note honesty, humanizer entry.`,
{ label:'review:CA-SKL-002', phase:'Session', schema:VERDICT, agentType:'Explore' }),
() => agent(`Adversarially review the HOME-leak fix (commits 66433fe + 325182d). Did isolating HOME hide real coverage? ` +
`Are there OTHER CLI-spawning tests still reading real ~/.claude that the grep audit missed? Is fix-cli REALLY leak-free?`,
{ label:'review:home-leak', phase:'Session', schema:VERDICT, agentType:'Explore' }),
])
phase('Synthesize')
const all = [...refutations, ...sessionReview].filter(Boolean)
const open = all.flatMap(r => r.rows).filter(x => x.verdict==='still-open' || x.verdict==='incorrectly-fixed')
return { open, attestation: open.length===0 ? 'all-closed (verified)' : `${open.length} open items`, raw: all }
```
## Suksesskriterier (verifiserbare)
- Hver `now-wrong`/`fix-false-positive`-rad har en eksplisitt `verdict` med `file:line`-bevis (ikke «ser bra ut»).
- Devil's advocate har FAKTISK forsøkt minst én counterexample-config per klynge (ellers er den ikke adversariell).
- Sluttleveranse: enten **«all-closed (verified)»** med rad-telling, ELLER en punch-list av reelt åpne items → hver blir en failing-test-først-fiks (Iron Law), IKKE auto-fikset uten godkjenning.
- `node --test 'tests/**/*.test.mjs'` forblir grønn; ingen workflow-agent endrer kode (read-only `Explore`-agenter).
## Scope-gjerde
Read-only analyse. Funn → rapport + forslag. INGEN kode-endring i devil's advocate-sesjonen uten at operatøren godkjenner punch-listen og ber om fiks (da: failing-test-først per item). Ikke utvid til feature-gap-implementering.
## Pekere
- Matrise: `docs/cc-2.1.x-gap-matrix.md` (stale ift. kode — verifiser mot kode, ikke mot matrisen).
- Overordnet plan: `docs/cc-2.1.x-gap-review-plan.md`. Fase 4-plan: `docs/cc-2.1.x-fase4-items-2-3-plan.md`.
- Fix-commits å revidere: `7309935`,`98ddd77`,`4b94da0`,`bec3f45`,`feaa7ed`,`624f5ed`,`b3c572a`,`8376dab`,`7bb2547`,`66433fe`,`325182d`.

View file

@ -21,7 +21,7 @@ User-impact category (added to each finding as `userImpactCategory`, derived fro
| Label | Scanners |
|-------|----------|
| Configuration mistake | CML, SET, HKV, RUL, MCP, IMP, PLH |
| Configuration mistake | CML, SET, HKV, RUL, MCP, IMP, PLH, OST |
| Conflict | CNF, COL |
| Wasted tokens | TOK, CPS |
| Dead config | DIS |

View file

@ -12,16 +12,19 @@ Scanner CLI: `node scanners/scan-orchestrator.mjs <path> [--global] [--full-mach
|---------|--------|---------|
| `claude-md-linter.mjs` | CML | Structure, length, sections, @imports, duplicates, TODOs |
| `settings-validator.mjs` | SET | Schema, unknown/deprecated keys, type mismatches, permissions |
| `hook-validator.mjs` | HKV | Format, script existence, event validity, timeouts |
| `hook-validator.mjs` | HKV | Format, script existence, event validity, timeouts, verbose-stdout (low), unfiltered `additionalContext` injection (info advisory, v5.10 B5) |
| `rules-validator.mjs` | RUL | Glob matching, orphan rules, deprecated fields, unscoped rules |
| `mcp-config-validator.mjs` | MCP | Server types, trust levels, env vars, unknown fields |
| `mcp-config-validator.mjs` | MCP | Server types, env vars, unknown fields |
| `import-resolver.mjs` | IMP | Broken @imports, circular refs, deep chains, tilde paths |
| `conflict-detector.mjs` | CNF | Settings conflicts, permission contradictions, hook duplicates |
| `feature-gap-scanner.mjs` | GAP | 25 feature checks across 4 tiers — shown as opportunities, not grades |
| `token-hotspots.mjs` | TOK | Cache-breaking volatile content, redundant tool permissions, deep import chains, oversized cascade, bloated SKILL.md descriptions, MCP tool-schema budget (Opus 4.7 patterns) |
| `cache-prefix-scanner.mjs` | CPS | Volatile content in lines 31150 of CLAUDE.md cascade (beyond Pattern A's top-30 window) |
| `disabled-in-schema-scanner.mjs` | DIS | Tools listed in BOTH `permissions.deny` AND `permissions.allow` — deny wins, allow entries are dead config |
| `token-hotspots.mjs` | TOK | Cache-breaking volatile content, redundant tool permissions, deep import chains, oversized cascade, bloated SKILL.md descriptions, MCP tool-schema budget, MCP tool-schema deferral (CA-TOK-006), stale plugin-cache disk-cleanup (prompt-cache patterns) |
| `cache-prefix-scanner.mjs` | CPS | Volatile content in lines 31150 of CLAUDE.md cascade (beyond Pattern A's top-30 window); plus volatile content inside `@import`-ed files (v5.10 B6, one hop) |
| `disabled-in-schema-scanner.mjs` | DIS | Dead/ineffective permission entries (low). (1) Tools in BOTH `permissions.deny` AND `permissions.allow` — deny wins; dominance is param-aware and treats the `Tool(*)` deny-all glob as equivalent to a bare deny (covers a bare allow). (2) Unanchored allow wildcards (`*`, `B*`, `mcp__*`) that Claude Code silently skips — CC accepts allow globs only after a literal glob-free `mcp__<server>__` prefix. Predicates shared with CNF live in `lib/permission-rules.mjs` |
| `collision-scanner.mjs` | COL | Cross-plugin skill name collisions (low); user-vs-plugin overlaps (medium); `details.namespaces` payload |
| `skill-listing-scanner.mjs` | SKL | (1) `CA-SKL-001` (medium): active skill descriptions over the verified 1,536-char listing cap (CC 2.1.105) → silently truncated in the model's skill listing. (2) `CA-SKL-002` (low): sum of active descriptions (each counted up to the cap) over the listing budget (~2% of context, CC 2.1.32), anchored on a conservative 200k window with a calibration note that the budget scales 5× on 1M-context models — leads with the measured sum, an estimate not telemetry. HOME-scoped (all user + plugin skills). Remediation surfaces `disableBundledSkills` / `skillOverrides` / trim. Distinct lens from TOK pattern F (project-local 500-char bloat heuristic) |
| `output-style-scanner.mjs` | OST | (1) `CA-OST-001` (medium): a user/project custom output style not setting `keep-coding-instructions: true` (defaults false) → silently strips Claude Code's built-in software-engineering instructions when active (V10). (2) `CA-OST-002` (low): a **plugin** style with `force-for-plugin: true` auto-applies and overrides the user's `outputStyle` (V11; plugin-styles-only per docs). (3) `CA-OST-003` (medium): a settings `outputStyle` matching no built-in (Default/Explanatory/Learning/Proactive, case-insensitive) nor discovered custom style → dead config (CC falls back to default). Reads each style's frontmatter via `parseFrontmatter`; fixture-gated (silent with no output styles). New scanner family in v5.6 C (count 13→14) |
| `optimization-lens-scanner.mjs` | OPT | `CA-OPT-001` (low, *Missed opportunity*): a CLAUDE.md procedure (≥6 consecutive numbered steps) that belongs in a skill (mechanism-fit, `BP-MECH-003`). Reads the machine-readable best-practices register (`best-practices-register.mjs`) for recommendation + provenance. Conservative — negative corpus proves null false-positive; prose-judgment cases (lifecycle→hook, "never"→permission) deferred to the Chunk 2b opus analyzer. Scoring area `CLAUDE.md` (existing → byte-stable). New scanner family in v5.7 Fase 1 Chunk 2a (count 14→15) |
## Scanner Lib (`scanners/lib/`)
@ -41,7 +44,7 @@ Scanner CLI: `node scanners/scan-orchestrator.mjs <path> [--global] [--full-mach
| `active-config-reader.mjs` | Read-only inventory: readActiveConfig(), detectGitRoot(), walkClaudeMdCascade(), readClaudeJsonProjectSlice() (longest-prefix match), enumeratePlugins(), enumerateSkills(), readActiveHooks(), readActiveMcpServers() (with cache → package.json tool-count fallback), estimateTokens() (v5: `'mcp'` kind = 500 + toolCount × 200) |
| `tokenizer-api.mjs` | Anthropic `count_tokens` wrapper for `--accurate-tokens` (v5 N5); 5s AbortController timeout, exponential 429 backoff, key masking |
| `humanizer.mjs` | Plain-language output translator (v5.1.0): `humanizeFinding`, `humanizeFindings`, `humanizeEnvelope`, `computeRelevanceContext`. Pure functions; never mutate inputs. Adds `userImpactCategory`, `userActionLanguage`, `relevanceContext` fields and replaces title/description/recommendation when a translation exists. Bypassed by `--raw` and `--json` paths. |
| `humanizer-data.mjs` | TRANSLATIONS table for 13 scanner prefixes (CML/SET/HKV/RUL/MCP/IMP/CNF/COL/TOK/CPS/DIS/GAP/PLH). Three-step lookup: exact title → regex pattern → `_default` → fall through to original |
| `humanizer-data.mjs` | TRANSLATIONS table for 16 scanner prefixes (CML/SET/HKV/RUL/MCP/IMP/CNF/COL/TOK/CPS/DIS/GAP/PLH/SKL/OST/OPT). Three-step lookup: exact title → regex pattern → `_default` → fall through to original |
## Action Engines (`scanners/`)
@ -52,8 +55,8 @@ Scanner CLI: `node scanners/scan-orchestrator.mjs <path> [--global] [--full-mach
| `fix-cli.mjs` | CLI: `node fix-cli.mjs <path> [--apply] [--json] [--global]` |
| `drift-cli.mjs` | CLI: `node drift-cli.mjs <path> [--save] [--baseline name] [--json]` |
| `whats-active.mjs` | CLI: `node whats-active.mjs <path> [--json] [--verbose] [--suggest-disables]` — read-only active-config inventory |
| `token-hotspots-cli.mjs` | CLI: `node token-hotspots-cli.mjs <path> [--json] [--global] [--output-file path] [--accurate-tokens] [--with-telemetry-recipe]`Opus-4.7 token hotspots ranking with optional API calibration |
| `manifest.mjs` | CLI: `node manifest.mjs <path> [--json]` — ranked system-prompt token-source table (v5 N2) |
| `token-hotspots-cli.mjs` | CLI: `node token-hotspots-cli.mjs <path> [--json] [--global] [--output-file path] [--accurate-tokens] [--with-telemetry-recipe]`prompt-cache token hotspots ranking (each hotspot tagged with its load pattern, v5.6 B2) with optional API calibration |
| `manifest.mjs` | CLI: `node manifest.mjs <path> [--json]` — ranked component-level token-source table, each source tagged with its load pattern + an always-loaded subtotal (v5 N2; load-pattern accounting v5.6 B) |
## Standalone Scanner
@ -67,10 +70,660 @@ Scanner CLI: `node scanners/scan-orchestrator.mjs <path> [--global] [--full-mach
| File | Content |
|------|---------|
| `claude-code-capabilities.md` | Feature register: 18 config surfaces, Anthropic guidance, relevance table |
| `configuration-best-practices.md` | Per-layer best practices (v5: Opus 4.7 cache-stability guidance replaces Sonnet-era 200-line rule) |
| `configuration-best-practices.md` | Per-layer best practices (v5: cache-stability guidance replaces Sonnet-era 200-line rule) |
| `anti-patterns.md` | Common mistakes mapped to scanner IDs |
| `hook-events-reference.md` | All 26 hook events with details |
| `hook-events-reference.md` | All 28 hook events with details |
| `feature-evolution.md` | Feature timeline for staleness detection |
| `gap-closure-templates.md` | Config-specific templates for closing gaps |
| `opus-4.7-patterns.md` | Token-cost dynamics for Opus 4.7 era — patterns powering the TOK scanner |
| `prompt-cache-patterns.md` | Token-cost dynamics (prompt-cache patterns) — patterns powering the TOK scanner |
| `cache-telemetry-recipe.md` | Manual `jq` recipe for verifying prompt-cache hit rate from session transcripts (v5 M7) |
## Implementation notes (per scanner / build block)
Detailed design rationale, primary-source verification, and byte-stability lessons for each scanner family and v5.6/v5.7 build block. Moved out of `CLAUDE.md` (kept lean per the "invariants only" rule); each note records why a change is correct and which frozen baselines it touched. Read on demand when working on the named scanner/block.
### active-config-reader — load-pattern model + rule/agent/output-style enumeration (v5.6 Foundation)
`scanners/lib/active-config-reader.mjs` now enumerates the three source kinds it previously
missed — **rules** (`enumerateRules`), **agents** (`enumerateAgents`), and **output styles**
(`enumerateOutputStyles`) — alongside the existing CLAUDE.md/plugins/skills/hooks/MCP enumerators.
Each new item, plus a pure `deriveLoadPattern(kind, {scoped})` helper, carries a
`loadPattern ∈ {always, on-demand, external}`, `survivesCompaction ∈ {yes, no, n/a}`, and
`derivationConfidence ∈ {confirmed, inferred}` derived from the published Claude Code loading
model (the V-rows in `docs/v5.5-steering-model-plan.md`). `readActiveConfig` exposes `rules`/
`agents`/`outputStyles` arrays + `totals` counts/subtotals (folded into `grandTotal`). This is
**internal plumbing** for v5.6 B (manifest/tokens rendering) — no command output changes yet, so
`--json`/`--raw`/SC-5 stay byte-stable. Output-style discovery is done directly (mirroring
`enumerateSkills`), **not** via a new `file-discovery` type, to keep the discovery surface stable.
The frontmatter parser (`scanners/lib/yaml-parser.mjs`) now also reads **YAML block sequences**
(`paths:\n - a\n - b`), not just inline `paths: "a, b"`. This resolves a pre-existing RUL
false-positive (a block-sequence-scoped rule was misread as unscoped). An empty-valued key with
no following `- ` items still resolves to `null` (backwards-compatible); only a real `- ` item
list becomes an array.
### manifest — load-pattern accounting (v5.6 B)
`buildManifest` (`scanners/manifest.mjs`) now consumes the Foundation enumeration. Two changes:
1. **Component-level sources (plugin roll-up dropped).** The coarse `kind:'plugin'` aggregate is
gone. A plugin contributes via its skills/rules/agents/output-styles/hooks/MCP — each already
enumerated **once** by `readActiveConfig` — so the old roll-up double-counted them (the plugin
aggregate's `estimatedTokens` already summed its components). Source kinds are now
`claude-md`/`skill`/`rule`/`agent`/`output-style`/`mcp-server`/`hook`.
2. **Load-pattern triple on every record + a `summary`.** Each source carries
`loadPattern`/`survivesCompaction`/`derivationConfidence`. Rules/agents/output-styles
**propagate** the foundation-derived values (rules vary by `scoped`); CLAUDE.md maps `scope`
kind via `CLAUDE_MD_SCOPE_KIND` (all cascade files walk **up**, so all are always-loaded);
skills are tagged **on-demand** via `deriveLoadPattern('skill-body')` — the measured tokens are
the skill **body** (paid on invoke), not the tiny always-loaded name+desc listing (tracked by
`skill-listing-budget`/posture), so tagging the body always would inflate the headline. The new
`summary` buckets sources into `always`/`onDemand`/`external`/`unknown` `{tokens,count}`; the
**always-loaded subtotal** ("≈X tokens enter context every turn before you type") is the headline.
**Byte-stability.** manifest is an **environment-aware CLI** → SC-6/SC-7 verify it by
**mode-equivalence** (`--json == --raw`), not byte-equal against a frozen snapshot, and it is not in
SC-5 default-output. Adding fields in place therefore keeps all snapshots green with **no regen**
(verified). `total` changes (de-duped, component-level) — that is the intended correctness fix.
### token-hotspots — load-pattern column (v5.6 B2)
TOK now annotates every ranked hotspot with the same load-pattern triple (`hotspotLoadPattern`
maps each discovery `type`→a `deriveLoadPattern` kind; rules reuse `activeConfig.rules` for precise
`scoped` handling; `claude-md` maps by scope). Two new `deriveLoadPattern` kinds back this:
**`command`** (on-demand — body loads on `/invoke`) and **`harness-config`** (external — settings/
keybindings/`.mcp.json`/hooks.json/plugin.json configure the CLI, **not** the model context, so they
cost no per-turn context tokens). Note the honest split: the `.mcp.json` **file** is `external`,
while the MCP **server**'s tool schemas are a separate `always` hotspot.
**Byte-stability — the opposite of manifest.** token-hotspots **is** a byte-equal SC-6/SC-7 CLI,
**and** its hotspots ride inside the scan-orchestrator + posture payloads, so the change broke
**six** frozen-v5.0.0 comparisons across five test files (json/raw-backcompat + the three Step 5/6/7
humanizer tests). Resolved by **preserving the frozen v5.0.0 baselines**: a shared
`tests/helpers/strip-hotspot-load-pattern.mjs` strips the additive triple before each byte-equal
compare (proves the original schema is byte-identical), and the **SC-5 default-output** snapshots
(scan-orchestrator + token-hotspots) were **regenerated** (`UPDATE_SNAPSHOT=1`) since their job is to
track current output — diff reviewed as additive-only. **Lesson for any future hotspot/scanner-output
field:** grep every frozen-v5.0.0 comparator (it is 5 files, not 2) before assuming the blast radius.
### token-hotspots — MCP tool-schema deferral (v5.10 B4, CA-TOK-006)
By default Claude Code **defers** MCP tool schemas: only tool *names* enter the always-loaded prefix
(~120 tokens total) and full schemas load on demand via tool search. Several signals force the FULL
schemas into the prefix every turn instead. CA-TOK-006 detects them from **config files only**, so the
finding is deterministic and hermetic-safe (mirrors Pattern G's project-local scoping):
| Signal | Source | Confidence |
|--------|--------|------------|
| `ENABLE_TOOL_SEARCH: "false"` | merged project+local settings.json `env` block | high |
| `"ToolSearch"` in `permissions.deny` | settings.json | high |
| configured `model` matches `/haiku/` | settings.json (Haiku lacks `tool_reference` support) | medium |
| per-server `alwaysLoad: true` | project `.mcp.json` (CC v2.1.121+) | high |
| `ENABLE_TOOL_SEARCH: "auto[:N]"` | settings `env` | threshold mode — **info, not a trigger** |
The engine (`lib/mcp-deferral.mjs`) splits a **pure** `assessMcpDeferral({settings, mcpServers})`
(fully unit-tested, no IO) from a thin IO wrapper `assessMcpDeferralForRepo(repoPath, {mcpServers})`
shared by TOK and GAP. Severity scales with the aggregate forced-upfront token cost
(`severityForForcedSchemas`: ≥5000→high, ≥1500→medium; **medium-confidence reasons cap at medium**).
**Honest scoping decision (Verifiseringsplikt).** The detector deliberately does **NOT** read
`process.env` shell vars. Tool search is also disabled on **Vertex AI**, with a custom
**`ANTHROPIC_BASE_URL`** (non-first-party host), or after a runtime **`/model`** switch to Haiku — but
those are launch/runtime state, not config files, so triggering on them would make the finding
machine-dependent (the marketplace-medium snapshot has MCP servers; an ambient `ANTHROPIC_BASE_URL`
would flap it). They are **disclosed** in every finding (`DEFERRAL_DISCLOSURE`), never triggered.
Tool-level `anthropic/alwaysLoad` (set server-side in the `tools/list` `_meta`) and claude.ai
connectors are likewise invisible to a static scan and disclosed. Mechanism verified 2026-06-23
against `code.claude.com/docs`: `context-window.md` (MCP deferred, ~120 tok),
`mcp.md#configure-tool-search` + `#exempt-a-server-from-deferral`, `costs.md`. The prefix-cache
connect/disconnect-invalidation claim from the raw research was **`[NOT CONFIRMED]`** in docs and is
NOT asserted by this finding.
**feature-gap companion.** `cliOverMcpLeverFinding` (GAP) fires **only** when CA-TOK-006's assessment
shows schemas forced upfront — recommends preferring CLI (`gh`/`aws`/`gcloud`) over MCP for common
operations (CLI adds zero context tokens until invoked). Deferred MCP is effectively free, so the
lever stays silent in the default case (opportunity, not noise — mirrors the bundledSkills lever).
`alwaysLoad` was added to CA-MCP's `VALID_SERVER_FIELDS` so it is never flagged as an unknown field.
### hook-validator — unfiltered additionalContext advisory (v5.10 B5)
A hook that emits `hookSpecificOutput.additionalContext` has that payload injected into Claude's
context **every time it fires** — plain stdout on exit 0 does NOT (it goes to the debug log only). A
hook that dumps large, un-grepped command output into `additionalContext` is therefore a recurring,
compaction-sensitive per-turn token cost. HKV flags it as an **`info` advisory** (weight 0 — never
severity-bearing, excluded from the self-audit `nonInfo` set), paired with a feature-gap lever.
The heuristic lives in `lib/hook-additional-context.mjs` as a **pure** `assessHookAdditionalContext({scriptContent})`
(unit-tested, no IO) plus a thin IO wrapper `assessHookContextForRepo(discovery)` (walk hooks → scripts
→ assess) used by GAP; HKV calls the pure function inline on scripts it already reads. The signal:
| Condition | Detected by | Effect |
|-----------|-------------|--------|
| references `additionalContext` | `/additionalContext/` | gate (else not applicable) |
| captures verbose-prone output | `cat`/`find`/`ls`/`git log\|diff\|status\|show`/`npm`/`pytest`/`jest`/`curl`/`execSync`/`readFileSync`… | `hasVerboseCapture` |
| applies any truncating filter | `grep`/`head`/`tail`/`sed`/`awk`/`jq`/`cut`/`wc`/`uniq`/`sort` or `.slice`/`.substring` | suppresses (assumed bounded) |
`flagged = buildsAdditionalContext && hasVerboseCapture && !hasFilter`. A filtered capture (e.g.
`cat … | grep ERROR`) or a cheap-only capture (`$(date)`) is not flagged.
**Why `info`, not a hard finding (Verifiseringsplikt).** This is deliberately **low precision** — a
static scan cannot run the hook or measure the real payload, and a filter we don't recognise would be
a false positive. So it ships as an advisory with the precision caveat in its own description, never a
graded/severity-bearing finding. Mechanism verified 2026-06-23 against `code.claude.com/docs`:
`context-window.md` — *"A PostToolUse hook … reports back via `hookSpecificOutput.additionalContext`.
That field enters Claude's context. Plain stdout on exit 0 does not."* + the tip to keep output concise
(it enters context without truncation).
**feature-gap companion.** `filterHookLeverFinding` (GAP) fires **only** when `assessHookContextForRepo`
returns ≥1 chatty hook — surfaces the documented **filter-before-Claude-reads** lever (`filter-test-output.sh`:
grep ERROR and return only matches instead of a 10,000-line log). No chatty hook → silent (opportunity,
not noise — same contract as the cliOverMcp / bundledSkills levers).
### cache-prefix-scanner — @import extension (v5.10 B6)
CPS originally scanned only the files discovery classifies as `claude-md`. But a CLAUDE.md can pull
arbitrary files into context with `@import` directives, and those targets are usually *not* `claude-md`
in discovery (e.g. `@shared/conventions.md`) — so their content was inlined into the cached prefix yet
never inspected. Neither TOK Pattern A (top-30 of cascade files) nor the in-file CPS scan reaches past
the importing file, so volatility in an imported file was invisible.
B6 closes the gap: for each `@import` whose **import site** sits within the cached-prefix window
(`imp.line ≤ CACHED_PREFIX_LINES`), CPS resolves the path (`resolveImportPath`, mirroring
import-resolver/token-hotspots semantics), reads the target, and runs `findVolatileLines` over its first
150 lines. A hit emits a distinct medium finding — *"Volatile content in @imported file breaks cached
prefix"* — keyed on the resolved file (so the fix points at the right place), with evidence naming the
importer (`imported by <file> (@<path> at line N)`).
**Scope boundaries (deliberate):**
- **One hop only.** Imports-of-imports are not followed — IMP owns deep-chain analysis. The verified win
is the direct import; transitive resolution adds cycle/depth complexity for marginal coverage.
- **No lines-130 skip for imported content.** That exclusion exists only to avoid duplicating TOK
Pattern A's territory in the *root* file; Pattern A never reads imported files, so all of the imported
prefix counts.
- **No double-reporting.** An import resolving to a file that is itself a discovered `claude-md` is
skipped (it gets its own in-file iteration); a `reportedImports` set dedupes a target imported by
several CLAUDE.md files.
**Byte-stability.** The in-file finding is emitted under exactly the same condition and with byte-identical
evidence/description as before (the `continue`-skip was refactored to an `if`-emit — behaviour-preserving).
New findings fire only when a discovered CLAUDE.md imports a volatile file, which no frozen v5.0.0 fixture
does — snapshots and SC-5 verified untouched by the full suite.
**Dropped from B6 (per plan `verdict`).** Confident behavioral cache-buster detection (opusplan /
model-switch is a *runtime* behaviour, not static config a scanner can reliably flag) and jq-transcript
automation. "No overstated behavioral finding ships" — so even the permitted opusplan *info*-advisory was
left out; the verified @import extension is the whole of B6.
### GAP scanner — authored-config scoping + direct cascade read (M-BUG-13)
The 25 presence checks ask "does the user's effective config have feature X?" and GAP **always**
runs `includeGlobal: true`. Two failure modes made the answer wrong on a real machine, both surfaced
by dogfooding `feature-gap`/`posture --global`:
1. **Demo/vendored config masks real gaps.** This plugin's own `examples/optimal-setup/` is a complete
config (sets `outputStyle`/`statusLine`/`worktree`/`model`/`keybindings.json`/`.lsp.json`), and its
copies vendored under `~/.claude/plugins/cache/.../config-audit/<ver>/examples/` are pulled into the
includeGlobal discovery. Because `anySettingsHas`/`files.some(...)` accept ANY discovered file, that
one demo file drove every tier-3 check to "present" → **GAP=0 on any target** (false negative).
Fix: `isAuthoredConfig` filters `ctx.files`/`parsedSettings` to the user's authored cascade —
excludes `~/.claude/plugins/` (absPath marker, mirrors CNF's M-BUG-2 exclusion) and any file whose
path **relative to the scan target** sits under `examples/` or `tests/fixtures/`. relPath (not
absPath) is deliberate: a fixture scanned AS the target keeps its own files, so the frozen v5.0.0
snapshots (scanned from `tests/fixtures/marketplace-medium`, which has no such nested trees) are
byte-stable.
2. **The real `~/.claude/settings.json` is invisible to the settings-key checks.** Discovery misses it
(its relPath carries no `.claude` segment when the walk root IS `~/.claude` — the gotcha) AND, when
vendored plugins flood the walk, the `maxFiles=2000` cap drops it. After (1) removed the demo
maskers, `statusLine`/`autoMode`/`permissions` (which the user HAS) would flip to false **positives**.
Fix: `readSettingsCascade` reads the four canonical cascade paths (user `settings.json`/`.local`,
project `settings.json`/`.local`) directly and merges them INTO `parsedSettings` — immune to the
cap and the gotcha. Merge (not replace) keeps non-canonical project settings and leaves the snapshot
(hermetic empty HOME → cascade adds nothing new) byte-stable.
Net: an empty target now surfaces ~18 humanized opportunities (was masked to ~0); config-audit's own
repo still shows 0 in output via its intentional `.config-audit-ignore` `CA-GAP-*` self-suppression
(a plugin repo legitimately lacks user-project features) — suppression is an envelope-layer concern,
orthogonal to this scanner fix. Scoped GAP-local; the includeGlobal discovery gotcha itself is left
to other consumers (see auto-memory `discovery-includeglobal-user-settings-gotcha`).
### CML scanner — context-window-scaled char budget
Beyond the line-count checks (200/500 lines, both MEDIUM), the CML scanner mirrors
Claude Code's own startup warning — *"Large CLAUDE.md will impact performance
(X chars > 40.0k)"* — as a `char`-based finding:
- **Char budget** — flags a CLAUDE.md over **~40.0k chars** (CC's startup-warning
figure at a 200k-context model). CC 2.1.169 scales that threshold with the model's
context window, so the finding anchors on the conservative 200k window (we cannot
observe the user's window; the anchor fires earliest) and discloses the relaxed
~200,000-char figure at 1M context. Severity MEDIUM (token cost, not an adherence
cliff). New `CA-CML` finding.
It keys on chars, not lines, so it is complementary to the line checks: a file can be
long by lines yet under budget (short lines), or short by lines yet over it (long lines).
The 200k/1M window constants live in the shared `scanners/lib/context-window.mjs`
(single source of truth, also re-exported by `skill-listing-budget.mjs`). The 40.0k
figure and context-window scaling are verified against the CC changelog (2.1.169) and
the live startup-warning text.
### DIS scanner — permission-rule hygiene
Beyond deny/allow overlap, the DIS scanner now also flags:
- **Ineffective allow wildcards** — unanchored tool-name globs in `permissions.allow`
(`*`, `B*`, `mcp__*`) that Claude Code silently skips (auto-approve nothing). Valid
only as a glob-free `mcp__<server>__*`. New `CA-DIS` finding, severity low.
- **`Tool(*)` deny-all glob** — treated as equivalent to a bare deny (`Bash(*)``Bash`),
so a bare allow killed by it is correctly reported as dead config.
- **Forbidden-param rules**`Tool(param:value)` whose key is the tool's own canonicalizing
field (`command` for Bash/PowerShell, `file_path` for Read/Edit/Write, `path` for
Grep/Glob, `notebook_path` for NotebookEdit, `url` for WebFetch). CC ignores these and
emits a startup warning. Severity follows intent: **deny/ask = false security (medium)**
the block never applies; **allow = dead config (low)**`param:value` matching is
deny/ask-only. Valid forms (`Bash(npm:*)`, `WebFetch(domain:host)`, `Agent(model:opus)`)
are never flagged. Predicate `forbiddenParamRule` in `permission-rules.mjs`.
These predicates live in `scanners/lib/permission-rules.mjs` (shared with the CNF
conflict-detector). Behavior verified against `code.claude.com/docs/en/permissions`.
### PLH scanner — plugin namespace collision
The standalone PLH scanner (cross-plugin checks in `scan()`) flags **plugin namespace
collisions**: two or more discovered plugins that declare the **same `name`** in
`plugin.json`. The search-first finding that shaped this check: Claude Code namespaces
every plugin component by the declared `name``/name:command`, `name:skill`, agent
`name` (verified against `code.claude.com/docs/en/plugins`, and observable in any session's
namespaced skill listing). A plugin component therefore can **never** shadow a user- or
project-level one; the only shadow that loses components is a same-`name` collision, where
the namespaces collapse into one and CC must pick a winner. Resolution between two installed
same-name plugins is **undocumented**, so the loser's commands/skills/agents go silently
unreachable — hence severity **MEDIUM** (dead config), `category: 'plugin-hygiene'`, with a
COL-shaped `details.namespaces` payload (`{ source: 'plugin:<dir>', name, path }`).
Two design notes: (1) the check keys on the declared `name` field, **not** `basename(dir)`
the folder name is irrelevant to the namespace; `scanSinglePlugin` now returns `declaredName`
for this. (2) Name-less plugins are excluded from the collision map (they are flagged by the
missing-field check and must never group on an `undefined` key).
The sibling cross-plugin **command-name** check was corrected to match the same model. Because
commands are namespaced (`/name:command`), a command name shared by two **differently-named**
plugins is ambiguity — not a hard conflict — so it now mirrors COL's plugin-vs-plugin skill
finding: severity **LOW**, `category: 'plugin-hygiene'`, COL-shaped `details.namespaces`, and a
group-first shape (one finding per command name listing every namespace, not pairwise). It keys
on the declared namespace and fires only when a name spans **2+ distinct** namespaces; when two
plugins share the same declared name, the namespace-collision finding above is the right (more
severe) signal, so the command check stays silent there to avoid a redundant `"dup, dup"` report.
The earlier HIGH `Cross-plugin command name conflict` finding (basename-keyed, "only one wins")
is gone, along with its now-inaccurate humanizer entry.
### PLH scanner — plugin-folder shadowing (`CA-PLH-015`)
Per-plugin check (in `scanSinglePlugin`, right after the required-field loop): a `plugin.json`
component-path key that **replaces** its default folder while that folder still exists on disk →
the folder is silently ignored (dead config). Severity **MEDIUM**, `category: 'plugin-hygiene'`,
`details: { field, ignoredDir, customPaths }`. Mirrors Claude Code's own warning in `/doctor`,
`claude plugin list`, and the `/plugin` detail view (v2.1.140+).
The field set is **primary-source-pinned** to the *replaces* category only —
`SHADOWING_PATH_FIELDS` = `commands`/`agents`/`outputStyles` (defaults `commands/`, `agents/`,
`output-styles/`). Deliberately excluded: **`skills`** (per
`code.claude.com/docs/.../path-behavior-rules` it *adds to* the default `skills/` scan — both
load, never a shadow), and **`hooks`/`mcpServers`/`lspServers`** (own merge rules, not a
folder-shadow). Experimental `themes`/`monitors` are omitted because the docs warn their manifest
schema may change between releases. The check also honors the doc's explicit-address exception: a
custom path that resolves *into* the default folder (`"commands": ["./commands/x.md"]`) is not
flagged, because Claude Code keeps scanning the folder in that case (`addressesDefaultDir`
predicate). The v5.4.0 plan originally listed `commands/agents/skills/hooks`; that set was
corrected here against the live docs (Verifiseringsplikt).
### PLH scanner — skills:-array validation (`CA-PLH-016`)
Per-plugin check (in `scanSinglePlugin`, after the shadow check): when `plugin.json` has a
`skills` field (string or array), each entry must resolve to an **existing directory inside the
plugin root**. The value is normalized `Array.isArray(v) ? v : [v]`, so a single string is one
entry — and a non-string top-level value (e.g. `42`) is naturally caught as a single non-string
entry (no separate top-level check needed). One finding per bad entry, severity **MEDIUM**,
`category: 'plugin-hygiene'`, `details: { field: 'skills', entry, problem }` where `problem` is
one of `non-string` / `escapes-root` / `not-found` / `not-a-directory`. Mirrors
`claude plugin validate` (~2.1.145).
Escape detection uses `skillsEntryEscapesRoot` (resolve + `startsWith(pluginDir + sep)`
containment — robust against a literal `..foo` dir name), backed by the docs' path-traversal rule
(*"Installed plugins cannot reference files outside their directory … such as `../shared-utils`"*).
`statOrNull` distinguishes missing from file-vs-dir. **Verifiseringsplikt note:** the v5.4.0 plan
claimed CC "suggests the parent directory when an entry points at a file"; that exact error text is
**not** in the primary docs, so it was dropped — the finding asserts only the four
primary-source-verified conditions. `skills` is deliberately *not* in `SHADOWING_PATH_FIELDS`
(it adds to the default scan, never shadows).
### PLH scanner — `scanDetailed`, `--output-file`, and the marketplace.json exemption (økt #46)
`scan()` returns the **frozen v5.0.0 envelope** (`scanner, status, files_scanned, duration_ms,
findings, counts`) and nothing else — `--raw`/`--json` print it verbatim and are snapshot-gated. Two
things the `/config-audit plugin-health` report requires therefore cannot live there: one row per
plugin (the `| Plugin | Grade | Commands | Agents |` table) and the cross-plugin/per-plugin split.
Both were computed inside `scan()` and discarded at the return: `pluginResults` never escaped, and
the only grade code — `formatPluginHealthReport` — had no caller anywhere in the repo.
`scanDetailed(targetPath)` is the seam. It returns `{ result, plugins, crossPluginFindings }`;
`scan()` is now `(await scanDetailed(p)).result`, so the byte-stable envelope is unchanged by
construction. `plugins[]` carries `name, declaredName, path, commandCount, agentCount, findingCount`
plus `score`/`grade` from the shared `pluginGrade(issueCount)` helper (which
`formatPluginHealthReport` now also calls, so the formula has exactly one home).
Cross-plugin findings are identified **positionally**, not by predicate: `crossPluginStart =
allFindings.length` is taken immediately before the namespace/command-name sections, and the tail is
sliced off at the end. A predicate would have to key on `category: 'plugin-hygiene'`, which the
per-plugin shadow and skills findings share. The marker (`crossPlugin: true`) is stamped only on the
**humanized copies** in the `--output-file` payload — never inside `finding()`, which would add a key
to the frozen envelope.
`--output-file` follows the `drift-cli` contract: humanized payload in default mode, stdout
untouched. This matters because default mode writes its report to **stderr**, and ux-rules rule 2
requires the command to run under `2>/dev/null` — before this, `commands/plugin-health.md` (and both
optional scanner calls in `commands/posture.md`) captured zero bytes. Argument parsing uses the same
`BOOL_FLAGS`/`VALUE_FLAGS` + unknown-flag-throws shape as `drift-cli`/`fix-cli` (M-BUG-21, third
arm); here the swallowed-flag failure mode was *worse than an error* — scanning the dropped flag's
value found no plugins, so the scanner answered `No plugins found` (info) with exit `0`.
**`marketplace.json` exemption:** `.claude-plugin/`'s known-file set is `plugin.json` **and**
`marketplace.json`. The catalog's location is documented and required (*"Create
`.claude-plugin/marketplace.json` in your repository root"*), and a marketplace entry with
`"source": "./"` makes the repo root its own plugin — so one `.claude-plugin/` legitimately holds
both. Verified against the primary docs before the change; the check was a false positive, latent in
this marketplace only because `catalog/` ships no `plugin.json` and is thus not scanned as a plugin.
### SET scanner — autoMode validation (`CA-SET`)
Per-file check in `settings-validator.mjs` (`autoMode` was in `KNOWN_KEYS` but had no nested
validation). Two sub-checks, both primary-source-verified against
`code.claude.com/docs/en/auto-mode-config`:
1. **Structure** (severity **MEDIUM**): `autoMode`, if present, must be an object whose only keys
are `environment`/`allow`/`soft_deny`/`hard_deny` (`AUTO_MODE_SUBKEYS`), each a **string
array** (the literal `"$defaults"` is a valid entry, so it passes the string check for free).
`problem``not-an-object` / `unknown-subkey` / `not-string-array` in `details`.
2. **Dead-config** (severity **LOW**): Claude Code does **not** read `autoMode` from *shared*
project settings — verbatim: *"The classifier does not read `autoMode` from shared project
settings in `.claude/settings.json`, so a checked-in repo cannot inject its own allow rules."*
The check keys on **`file.scope === 'project'`** (file-discovery's `classifyScope` returns
`'project'` for a committed `.claude/settings.json`; `'local'`/`'user'`/`'managed'` are read and
not flagged). `problem: 'shared-project-scope'`. This is why the plan's "test per-file scope
first" gate passed — `ConfigFile` already carries `scope`.
The two sub-checks are independent (a malformed autoMode in shared scope yields both). SET is in the
orchestrator, so SC-5 was re-checked after this change — byte-equal (the snapshot fixture has no
`autoMode`, so the block never fires there).
### OST scanner — output-style validation (`CA-OST`, v5.6 C, count 13→14)
New orchestrated scanner `output-style-scanner.mjs` — the first new scanner family since SKL
(v5.2.0). It reads the active config (`readActiveConfig`) and each output-style file's frontmatter
(via `parseFrontmatter`, keys hyphen→underscore-normalized, so it reads `keep_coding_instructions` /
`force_for_plugin`). Three findings, every claim pinned to a CONFIRMED row of
`docs/v5.5-steering-model-plan.md` (V9/V10/V11/V12), re-verified against
`code.claude.com/docs/en/output-styles` + `.../plugins-reference`:
- **`CA-OST-001`** (medium) — a **user/project** custom style not setting `keep-coding-instructions:
true`. The flag defaults to **false**, so the style silently **removes** Claude Code's built-in
software-engineering instructions when active (V10). Scoped to user/project (the styles the user
authors); a plugin author's choice is out of scope.
- **`CA-OST-002`** (low) — a **plugin** style with `force-for-plugin: true`, which auto-applies and
**overrides** the user's selected `outputStyle` (V11). **Verifiseringsplikt correction:** the v5.5+
plan's CA-OST-002 bullet said "in a project/user style," but `force-for-plugin` is
**plugin-styles-only** per the docs (its own cited V11 + `output-styles.md`), so the check keys on
`source === 'plugin'` — a user/project style with the flag is simply ignored, not an override.
- **`CA-OST-003`** (medium) — a settings `outputStyle` value resolving to **no** built-in
(`Default`/`Explanatory`/`Learning`/`Proactive`, matched case-insensitively) and **no** discovered
custom style → dead config (CC falls back to default; the configured behavior never applies).
**Byte-stability — a scanner addition, NOT a field addition.** Adding the 14th scanner grows
`envelope.scanners` by one entry and bumps `aggregate.scanners_ok` 12→13 on the deterministic
fixture **regardless of findings** — a field-strip helper cannot paper this over. The SKL precedent
(`7bb2547`) re-seeded the frozen v5.0.0 snapshots, but that predates B2's strip-preservation regime;
re-seeding now would **bake in** B2's hotspot triple + `claudeMdEstimatedTokens` drift (verified by
inspecting the seed diff). So, consistent with the B2 lesson ("preserve frozen via strip-helper;
regen ONLY SC-5"), C **preserves** the frozen v5.0.0 snapshots and **strips the OST entry at compare
time**: shared `tests/helpers/strip-added-scanner.mjs` (`stripAddedScanners` removes OST entries +
decrements `scanners_ok`; `stripAddedScannerStderr` drops the `[OST]` progress line) is wired into
json/raw-backcompat + the Step 5/6 humanizer wiring tests (cli-humanizer did **not** break — its
v5.0.0 compares don't grow a scanners array). Only **SC-5 default-output** (scan-orchestrator +
posture) is regenerated (additive OST entry only — diff reviewed). OST is fixture-gated: the
`marketplace-medium` fixture and the hermetic HOME have no output styles, so it emits nothing there.
Wiring: orchestrator import + `SCANNERS` entry; `humanizer.mjs` `SCANNER_TO_CATEGORY`
(`OST: 'Configuration mistake'`); `humanizer-data.mjs` OST family (title-coupled to the three exact
finding titles); `scoring.mjs` `SCANNER_AREA_MAP` (`OST: 'Settings'` — keeps the 10 quality areas,
byte-stable on zero-finding projects). Count badges: self-audit scanner count 13→14; humanizer-data
TRANSLATIONS families 14→15 (PLH is a translation family but not orchestrated).
### best-practices register — machine-readable knowledge layer (v5.7 Fase 1 Chunk 1)
`knowledge/best-practices.json`: provenance-stamped, schema-validated register (entry =
`id`/`claim`/`confidence`/`source` + optional `mechanism`/`lensCheck`/…). First runtime-consumed
file in `knowledge/` (the `*.md` stay human-only); source of truth for the v5.7 optimization lens
(`CA-OPT`); seeded from the v5.5 V-rows + the Anthropic "Steering Claude Code" blog. Only
**confirmed** entries are user-facing (Verifiseringsplikt). Loaded/validated by
`scanners/lib/best-practices-register.mjs` (`loadRegister`/`validateRegister`/`getEntry`; zero-dep
JSON, **not** YAML — `yaml-parser.mjs` can't do arrays-of-objects). Byte-stable until a scanner
consumes it (Chunk 2). Full design: `docs/v5.7-optimization-lens-plan.md`.
### OPT scanner — optimization lens / mechanism-fit (`CA-OPT`, v5.7 Fase 1 Chunk 2a, count 14→15)
First detector of the «optimal?» axis (vs «correct?»). `optimization-lens-scanner.mjs` reads the
best-practices register and flags config that works but fits a better mechanism. **`CA-OPT-001`**
(low, *Missed opportunity*): a CLAUDE.md procedure (≥6 consecutive numbered steps) that belongs in a
skill — recommendation/provenance from register `BP-MECH-003`. Conservative (negative corpus = null
false-positive); prose-judgment cases (lifecycle→hook, unscoped path→rule, «never»→permission) are
handled by the Chunk 2b opus analyzer (below). Wiring mirrors OST: orchestrator entry, humanizer
`OPT:'Missed opportunity'` + family, scoring `OPT:'CLAUDE.md'` (existing area → no new posture row →
byte-stable), strip-helper `OPT`, SC-5 regenerated (additive).
### Optimization lens Chunk 2b — opus analyzer (prose-judgment half, `/config-audit optimize`)
The hybrid motor's recall + precision halves for the three cases the deterministic OPT scanner skips.
**Pre-filter** (`scanners/lib/lens-prefilter.mjs`, pure + tested): cheap, recall-oriented line scan
of CLAUDE.md body for lifecycle phrasing (`BP-MECH-001`→hook), unscoped path-specific instructions
(`BP-MECH-002`→rule), and absolute «never» prohibitions (`BP-MECH-004`→permission); skips fenced
code, gates the path class on an instruction verb. Detector names = the register `lensCheck` fields.
**CLI** (`optimize-lens-cli.mjs`, `-cli` → not a scanner): runs discovery + OPT scanner + pre-filter,
attaches the **confirmed** register entry to each candidate (unverifiable → dropped, Verifiseringsplikt),
emits `{deterministic, candidates, register, counts}`. **Agent** (`optimization-lens-agent`, opus,
orange — the 7th agent, **precision gate**): reads the real CLAUDE.md, drops low-confidence candidates,
keeps only genuine opportunities, cites register id + source. **Command** `/config-audit optimize`
orchestrates pre-filter→agent→report. **Agent-driven → deliberately NOT byte-stable** (own command,
outside the snapshot suite); the pre-filter lib *is* unit-tested (13 tests). No new orchestrated
scanner → scanner count stays 15; agents 6→7, commands 18→19, suite 1055→1068.
### Subtraction lens (`optimize --subtract`, `BP-SUB-001`) — the inverse axis
**Why a mode, not a command or scanner.** Every other command asks an addition question; nothing
asked what is no longer earning its rent. A hand-built ground truth over a real 250-line global
CLAUDE.md put the honest payoff at ~8501400 always-loaded tokens of ~4300 (~20 %, not the source
anecdote's 80 %), with 26 of 34 blocks load-bearing. That proportion justifies a fourth `lensCheck`
on the existing hybrid motor — not a new scanner (no badge bump 16→17, no frozen-snapshot risk) and
not a 22nd command. Cross-session state is what would have forced a command, and the earned-re-add
*ledger* is deliberately out of v1.
**Polarity is flipped from `lens-prefilter`.** That module is recall-first because a false candidate
only costs the judge a moment. Here a false candidate is a proposal to *delete*, so
`subtraction-prefilter.mjs` is precision-first and carries a blocking deterministic guarantee.
**Floor-exclusion is a separate module on purpose** (`lib/floor-exclusion.mjs`). §6.0's asymmetry —
a missed dead line costs a few tokens per turn, a deleted load-bearing line costs a wrong remote —
means the guarantee must not rest on a probabilistic judge, so the floor runs *before* the agent
and is legible as its own unit. Markers: code span, URL/host, filename, rooted path, version pin,
policy invariant, and `unresolved-entity` (a mixed-case capitalized word mid-sentence). The last is
a deliberate conservative default — resolving "Forgejo" from an ordinary capitalized word needs a
dictionary, so the mechanism declines and keeps the block.
**Granularity: leaf block + two structural exceptions.** (1) A paragraph ending in `:` merges with
the list it introduces — a stem often carries no literal of its own, and deleting it without its
list is meaningless. (2) An *ordered* list is a contract: steps inherit floor from any sibling,
because deleting step 2 of a five-step protocol is not like dropping one platitude. Unordered lists
do **not** inherit — a load-bearing bullet and a disposable one routinely share a list, and
container-reasoning is exactly the error the ground truth was built to catch.
**Three lessons from the dogfood run**, all invisible to the synthesized fixture and worth keeping:
- **JS `\b` is ASCII-only.** `/\bunngå\b/` never matches — the trailing `å` is not a word character,
so there is no boundary after it. Every Norwegian keyword ending in æ/ø/å was silently dead. Use
the `LB`/`RB` lookaround constants, never `\b`, around that vocabulary.
- **A bare `word/word` is not a path.** `pros/cons` vetoed the single largest deletable block until
`PATH_RE` was tightened to rooted paths and globs; real filenames are `FILENAME_RE`'s job.
- **"Mid-sentence" must key on a preceding lowercase letter**, not on "anything that is not a full
stop". The loose version read `**Bold labels:**` and quoted openers (`"Som AI kan jeg ikke…"`) as
entities and cost 4 of 11 deletable groups.
**Measured against the ground truth:** zero load-bearing blocks proposed (the blocking §8 gate),
11/18 deletable groups surfaced, ≈756 tok ≈ 18 % of the file — inside the pre-registered band. The
misses are all the conservative default working as designed (entity names, code spans, and
declarative/infinitive phrasing that carries no imperative). No new scanner: scanners stay 16,
commands 21, agents 7; suite 1365→1382.
**Note for a future narrowing of the veto:** the two failure modes the agent prompt hardens against
— staleness-is-not-deletion and tier-2-is-not-tier-3 — are currently *also* covered by exclusion
(both example blocks carry code spans/version pins and never reach the judge). The prompt language
is the only protection if those markers are ever loosened.
**Test-isolation fix (this session):** `token-hotspots.test.mjs` `runScanner` now wraps `scan()` in
the shared `withHermeticHome` helper — the suite is green on BOTH a real and a clean `HOME` (the OPT
section's old «run with clean HOME» caveat is resolved). Snapshot/byte tests were already hermetic.
### knowledge-refresh — the "living" half of the register (v5.7 Fase 1 Chunk 3, commands 19→20)
Keeps `knowledge/best-practices.json` current so the optimization lens never reads stale rules.
Same hybrid split as Chunk 2b — a deterministic, byte-stable, unit-tested core + a web/judgment shell:
- **Deterministic core** (`scanners/lib/knowledge-refresh.mjs`, pure, 15 tests): `assessFreshness(register,
{referenceDate, staleAfterDays})` classifies each entry `fresh`/`stale` by the age of its
`source.verified` stamp. `referenceDate` is **injected** (not read from the clock) so the function is
fully deterministic; default threshold `STALE_AFTER_DAYS_DEFAULT = 90` (quarterly re-verify cadence).
An unparseable/missing `verified` → stale with `ageDays: null` (defensive; the schema-validated bundle
never hits this, but the command's hand-built candidates might). «Source changed» detection is a **web
responsibility** (command layer), **not** in this core.
- **CLI** (`scanners/knowledge-refresh-cli.mjs`, `-cli`**NOT** an orchestrated scanner → scanner count
stays 15, suite byte-stable; 8 tests): read-only — it NEVER writes the register and NEVER hits the
network. `--reference-date` (defaults to today; the **only** place the clock is read) makes it
deterministically testable against the bundled register. `--stale-after N`, `--dry-run` (implicit + only
mode, echoed as `requestedDryRun`). Exit **0** = all fresh, **1** = some stale (advisory), **3** = error.
- **Command** (`commands/knowledge-refresh.md`, opus): orchestrates CLI stale-report → re-verify each stale
entry by re-reading its `source.url` (WebFetch) → poll CC changelog + Anthropic blog for new/changed
practices (WebSearch) → present everything → **apply ONLY human-approved writes**, then re-run the
register schema test before declaring done. **No unverified claim is ever auto-written** (Verifiseringsplikt).
Web/judgment-driven → **deliberately NOT byte-stable** (own command, outside the snapshot suite), exactly
like `/config-audit optimize`. **No new agent** (web poll runs in the command's own context), **no new
orchestrated scanner**. suite 1068→1091.
### campaign-ledger — durable machine-wide campaign core (v5.7 Fase 2, Block 3a THIN)
`scanners/lib/campaign-ledger.mjs`: the durable ledger that sits ABOVE individual sessions for a
machine-wide audit campaign — repo list + per-repo lifecycle (`STATUSES` = pending→audited→planned
→implemented) + a machine-wide `rollUp` (counts by status + severity aggregated across repos). It
persists to a single JSON file **outside** the plugin dir (`~/.claude/config-audit/campaign-ledger
.json`, next to `sessions/`) so it survives uninstall/reinstall/upgrade. Same hybrid split as
knowledge-refresh: PURE transforms (`createLedger`/`addRepo`/`setRepoStatus`/`rollUp`) with `now`
**injected** (YYYY-MM-DD, never the clock) + soft `validateLedger` (returns `{valid,errors}`, never
throws) + a thin IO shell (`defaultLedgerPath`/`loadLedger`→null-on-ENOENT/`saveLedger`). Transforms
throw on programmer error (invalid status, unknown path); `schemaVersion` stamped from the start so a
Block 4 migration is cheap. **THIN**: ledger + roll-up + persistence only — NO execution, CLI, or
command surface (Blocks 3b/3c/4). **Internal plumbing, byte-stable until consumed**: no `export async
function scan` + lives in `lib/` → scanner count stays 15, no orchestrator wiring, SC-5 unchanged.
28 tests, suite 1091→1119.
### campaign-cli — read-only ledger reporter (v5.7 Fase 2, Block 3b)
`scanners/campaign-cli.mjs` (`-cli` → NOT an orchestrated scanner → scanner count stays 15, suite
byte-stable; 8 tests): the DETERMINISTIC, READ-ONLY half of the campaign motor, mirroring
`knowledge-refresh-cli`. It `loadLedger`s the durable ledger, `validateLedger`s it, and emits
`{status, initialized, ledgerPath, schemaVersion, createdDate, updatedDate, repos, rollUp}` as JSON.
It NEVER writes — a missing ledger is reported gracefully (`initialized:false`, all-zero roll-up),
**never created**; init + every status transition belong to the Block 3c command layer (human-approved
writes, Verifiseringsplikt). `--ledger-file` overrides the default path (deterministic testing);
`--output-file` mirrors the sibling. Exit codes: **0** = initialized & valid, **1** = not initialized
yet (advisory), **3** = error (parse/corrupt/invalid). suite 1119→1127.
### campaign-write-cli + `/config-audit campaign` — the WRITE half (v5.7 Fase 2, Block 3c, commands 20→21)
The human-approved mutation half of the campaign motor, completing the THIN campaign surface
(ledger + roll-up + status). Two pieces:
- **`scanners/campaign-write-cli.mjs`** (`-cli` → NOT an orchestrated scanner → scanner count stays
15, suite byte-stable; 11 tests): the sibling of `campaign-cli` that *mutates*. Subcommands
`init` / `add <path>...` / `set-status <path> <status>`, each a thin wrapper over the
invariant-enforcing lib transforms (`createLedger`/`addRepo`/`setRepoStatus`) + `saveLedger` — so
path-normalization/dedup, idempotent add, the status-lifecycle guard, and the `updatedDate` bump
are **never re-implemented by hand**. `init` refuses to clobber an existing (or corrupt) ledger
(**exit 1** advisory, file untouched); `add` **auto-inits** when no ledger exists and reports
`added` vs `skipped`; `set-status` accepts `--findings '<json>'` + `--session <id>`. Determinism
mirrors the lib + `knowledge-refresh-cli`: `--reference-date` is the **only** place the clock is
read (defaults to today), passed to the transforms as the injected `now`. Exit: **0** = write
performed, **1** = advisory no-op (init-clobber), **3** = error (unknown subcommand, bad args,
invalid status, untracked repo, no/corrupt ledger).
- **`commands/campaign.md`** (opus, `allowed-tools: Read/Write/Edit/Bash/Glob`**no Web**,
judgment-free): a thin orchestrator. It always **reports** first (read-only `campaign-cli`), then
for `init`/`add`/`set-status` it proposes the change and, **only on explicit human approval**,
invokes one write-CLI subcommand (Verifiseringsplikt — it never hand-edits the ledger JSON).
`add --discover <root>` finds git repos under a root and lets the user pick. When marking a repo
`audited` it attaches findings-by-severity from the repo's session (or user-provided counts) —
**never invented**.
**Not a new scanner, not byte-stable.** Both CLIs carry the `-cli` suffix (out of the
scan-orchestrator → scanner count stays **15**, snapshot suite untouched); the command's
orchestration is judgment-driven and deliberately outside the snapshot suite, exactly like
`/config-audit optimize` + `knowledge-refresh`. **No new agent** (web/judgment-free, runs in the
command's own context). suite 1127→1138.
### campaign backlog — cross-repo prioritized pick-list (v5.7 Fase 2, Block 4b)
The first half of Block 4 ("one cross-repo prioritized backlog the user picks from"). A pure
lib transform + a read-only CLI-payload field — **no schema change, no new scanner, byte-stable**.
- **`buildBacklog(ledger)`** (`scanners/lib/campaign-ledger.mjs`, pure, mirrors `rollUp`): the
single machine-wide prioritized work list. The actionable unit is a **repo** (the ledger tracks
per-repo severity *counts*, not individual findings — it tracks state, it does not re-run
audits), so each item is one repo: `{path, name, status, sessionId, findingsBySeverity
(normalized), totalFindings, weightedScore, rank}`. **Inclusion:** `status !== 'implemented'`
AND `totalFindings > 0` (implemented = done; pending / zero-finding repos have nothing known to
fix — they still surface in `rollUp.byStatus`). **Order:** DESC by `weightedScore` (exported
`SEVERITY_WEIGHTS = {critical:1000, high:100, medium:10, low:1}`), tie-broken lexicographically
by critical→high→medium→low count, then ascending `name` — fully deterministic, and the
tie-break keeps "criticals always win" even on a weighted-score collision (1 critical vs 10 high).
`rank` is 1-based after the sort.
- **`campaign-cli`** now emits `backlog: buildBacklog(ledger)` in both branches (uninitialized →
`[]`). Purely additive + read-only → fits the Block 3b read-only contract; the existing CLI
tests use targeted asserts (not full `deepEqual`), so the new field doesn't break them.
- **`commands/campaign.md`** renders the backlog as a "Prioritized backlog" pick-list and points
the user at the top item (still a pick-list, NOT an executor — execution is the later 4c block).
**Byte-stability.** `-cli`/lib/command only → scanner count stays **15**, snapshot/backcompat
suite untouched. suite 1138→1150 (lib +9, campaign-cli +3). **Deferred to 4c:** per-repo plan
export to each repo's `docs/` + reuse of backup/rollback for execution. **Deferred until the first
breaking schema change:** `migrateLedger` (4a) — backlog needs no schema bump, so building
migration now would be speculative (`schemaVersion` is already stamped for when it's needed).
### campaign plan-export + execution-by-reuse (v5.7 Fase 2, Block 4c — the rest of Block 4)
The second half of Block 4. **Asymmetric:** plan export is the new testable code; execution is
*pure reuse* (no new machinery), per the plan's "reuse existing backup/rollback".
- **Plan export** (`scanners/lib/campaign-export.mjs`, pure, 8 tests): `planExportPath(repoPath,
sessionId)` → `<repo>/docs/config-audit-plan-<sessionId>.md` (keyed on the timestamp-unique
sessionId, not the date, so same-day re-audits don't collide); `buildPlanExportDocument({...,now})`
→ provenance header (repo/session/how-to-execute-and-undo) + the verbatim session plan. `now`
injected → deterministic. **CLI** `scanners/campaign-export-cli.mjs` (`-cli`, read-only by
default, 10 tests): `--repo <path>` resolves the repo's linked session, reads its
`action-plan.md`, assembles the doc, emits `{exportable, problems, targetPath, document, ...}`.
Two gates → exit 1 advisory: `no-session-linked` (repo has no `sessionId`), `no-action-plan`
(linked session has no plan yet). Writes the file **only** under opt-in `--write` — the CLI does
the byte-faithful copy so a 200-line plan is never re-typed/mutated by the LLM. `--sessions-dir`
override for hermetic tests; exit 0/1/3 mirror the sibling CLIs.
- **Execution = reuse.** No campaign-side execution code. The exported `docs/` file is the repo's
durable record; `/config-audit implement` still reads the canonical plan from the session
(backup + apply + verify), `/config-audit rollback` undoes, then `set-status <path> implemented`
records it. The command (`commands/campaign.md`, new `export <path>` mode) previews → asks → on
approval invokes `--write` → routes the user to that existing machinery.
- **Byte-stable.** lib + `-cli` + command-doc only → scanner count stays **15**, agents **7**,
commands **21** (export is a *mode*, not a new command), snapshot/backcompat suite untouched.
suite 1150→1168 (lib +8, export-cli +10). **Block 4a (`migrateLedger`) still deferred** to the
first breaking schema change (export needs no schema bump).

View file

@ -0,0 +1,161 @@
# v5.13 Plan — Model Routing, Effort Awareness, Dead References
> **Target version is now v5.14.** The `5.13.0` slot was consumed by the pipeline-hardening batch
> (`optimize --subtract` is a feature, so that release could not be a patch). The filename is kept so
> existing references resolve; only the target moved.
Derived from an external video analysis ("The Model Isn't the Moat", 2026-07) cross-checked
against primary sources and against what config-audit already encodes. Every claim acted on
here was verified against Anthropic's own docs; video-only claims are explicitly rejected below.
## Source verification (done 2026-07-14)
| Claim from video | Verdict | Source |
|---|---|---|
| Orchestrator + cheaper worker models is a supported, recommended pattern | VERIFIED | code.claude.com/docs/en/sub-agents ("Control costs by routing tasks to faster, cheaper models like Haiku"), code.claude.com/docs/en/workflows |
| Reasoning effort is tunable per settings / session / launch / **per-agent frontmatter** / SDK; levels `low, medium, high, xhigh, max` | VERIFIED | code.claude.com/docs/en/model-config#adjust-effort-level, sub-agents doc |
| Leaked Fable 5 system prompt principles ("partial recognition ≠ current knowledge"; "a prompt implying a file is present doesn't mean one is"; answer-first-then-one-question; tool-call scaling 1 / 35 / 510) | VERIFIED near-verbatim, **provenance unconfirmed** (third-party leak repo, not Anthropic-confirmed) | github.com/asgeirtj/system_prompts_leaks `Anthropic/claude-fable-5.md` |
| "Fable 5 on low ≈ Opus 4.8 on high, slightly higher cost/quality" score-vs-cost chart | **CONTRADICTED** — no such chart/statement on Anthropic's pages; GPT-5.5 appears only in a testimonial | anthropic.com/news/claude-fable-5-mythos-5 |
## Already covered — no action
| Video idea | Existing coverage |
|---|---|
| "Process is the moat" (config/harness > raw model) | The plugin's entire thesis |
| Extract repeated procedure into a skill | BP-MECH-003 + CA-OPT-001 (`optimization-lens-scanner.mjs:121`) |
| CLAUDE.md size/ownership discipline | BP-SIZE-001 + CA-CML line/size checks (`claude-md-linter.mjs:109/:120/:140`) |
| Check that referenced imports exist | CA-IMP broken `@import` (`lib/import-resolver.mjs:88`) — but **only** `@import`, see Chunk 3 |
| Plugin's own agents are model-routed | Agents table already pins sonnet for mechanical, opus for judgment |
## Gaps → chunks
Verified gap summary (register-mapper sweep, 2026-07-14): no scanner audits per-agent
`model:`/`effort:` frontmatter; effort has validity-check only (`settings-validator.mjs:195`),
no recommendation; no dead-reference check for prose file mentions in CLAUDE.md; no
adversarial/failure-mode requirement in planner-agent; `fix-engine.mjs:26` effort list
omits `xhigh` (settings-validator has all five).
### Chunk 1 — Register entries: model routing + effort (dogfoods `knowledge-refresh`)
Add to `knowledge/best-practices.json` via the knowledge-refresh flow (human-approved write):
- **BP-MODEL-001** (`category: model-fit`): subagents doing mechanical/read-only work can pin a
cheaper model via `model:` frontmatter; orchestrator keeps the strong model. Source:
code.claude.com/docs/en/sub-agents → `confidence: confirmed`.
- **BP-MODEL-002** (`category: model-fit`): reasoning effort is tunable at five levels in five
places (settings `effortLevel`, `/effort`, `--effort`, per-agent `effort` frontmatter, SDK);
default `high`; higher effort is not universally better for simple tasks. Source:
code.claude.com/docs/en/model-config → `confidence: confirmed`.
Schema per `scanners/lib/best-practices-register.mjs:42-102` (id/claim/confidence/source.url/
source.verified required). This chunk doubles as the DEL B dogfood of `/config-audit
knowledge-refresh` (each chunk is also a plugin test).
### Chunk 2 — fix-engine effort hygiene (tiny, TDD)
`fix-engine.mjs:26` `VALID_EFFORT_LEVELS = ['low','medium','high','max']` — missing `xhigh`.
Consequence: nearest-match "fix" for a typo like `xhig` corrects to `high`, not `xhigh`.
Red test first: `findNearestEffortLevel('xhig') === 'xhigh'`. Align list with
`settings-validator.mjs:75`.
### Chunk 3 — CA-CML dead prose references (new deterministic check)
The strongest video-derived principle ("a prompt implying a file is present doesn't mean one
is") applied to CLAUDE.md quality: flag file paths mentioned in CLAUDE.md **prose** that do not
exist on disk. Today only `@import` targets are existence-checked; stale pointers like
`docs/foo.md` or `scripts/bar.sh` rot silently and burn always-loaded tokens on misdirection.
Conservative v1 to control false positives:
- Only backtick-quoted tokens that look like relative file paths (contain `/` or a known
extension), resolved against the CLAUDE.md's own directory.
- Skip URLs, globs (`*`), placeholders (`{...}`, `<...>`, `$VAR`, `${...}`), absolute and
`~/` paths (machine-specific), and paths under `.gitignore`d dirs if cheap to determine.
- Severity: low. New CA-CML-NNN (verify next free NNN at implementation — IDs are dynamic).
Byte-stability: follow [[adding-scanner-byte-stability]] steps for a new finding type in an
EXISTING scanner — frozen `tests/snapshots/v5.0.0/` must stay untouched; default-output
snapshots regenerate (`UPDATE_SNAPSHOT=1`) only if a fixture actually carries the new type;
humanizer step 7 (M-16/M-17 lessons): `TRANSLATIONS`-static entry for the new RAW title
(CML category mapping already exists).
### Chunk 4 — feature-gap + inventory: model/effort awareness
- New T3 opportunity check in `feature-gap-scanner.mjs`: authored agents
(`isAuthoredConfig`, M-BUG-13 lesson) where **no** agent sets `model:` or `effort:`
"all agents inherit the session model/effort — mechanical agents can be routed cheaper /
effort-calibrated" citing BP-MODEL-001/002. Fires only when authored agents exist
(M-BUG-15 lesson: no enhancement-check on empty collections). Opportunity framing, never
failure — deliberate max-model setups are a valid choice; finding is suppressable
(`.config-audit-ignore`).
- `whats-active` / `manifest`: surface `model`/`effort` per agent in the inventory tables.
- Humanizer wiring step 7 for the new GAP finding; verify via direct `scan()` output, not the
self-suppressed default output ([[agent-commands-need-scanner-scoping]]).
- feat commit → docs-gate: README + CLAUDE.md diffs required.
### Chunk 5 — planner-agent adversarial gate (do AFTER DEL B pipeline dogfood)
Add a required "Failure modes" section to `agents/planner-agent.md`'s action-plan contract:
before an action plan is emitted, list what could go wrong per change + rollback trigger.
Mirrors the video's scoping-vs-devil's-advocate distinction; currently absent (zero
adversarial requirements in agents/). **Sequencing constraint:** DEL B step 3.2 judges
planner-agent against a fasit — change the agent only after that dogfood pass, or the
fasit target moves mid-evaluation.
## Explicitly rejected (do not revisit without new evidence)
1. **"Fable low ≈ Opus high" cost/score framing** — contradicted by Anthropic's own pages.
Never encode in register, copy, or recommendations.
2. **Tool-call-count effort scaling (1 / 35 / 510) as a register entry** — source is an
unconfirmed third-party leak → would be `confidence: inferred`, never surfaced. Not worth
carrying.
3. **Cost/intelligence/"taste" model-routing table generator** — subjective scores don't fit
the deterministic, provenance-gated design. BP-MODEL-001 covers the actionable core.
4. **"Fable mode" skill** — a user-level skill, not configuration auditing. Out of plugin scope.
5. **CLAUDE.md prose contradiction detection** — real gap (CA-CNF only covers
settings/permissions/hooks) but not video-driven; keep this plan surgical.
## Known tension (named, not resolved here)
The operator's own global policy is Opus/max-effort for ALL subagents, never Haiku — the
opposite of Chunk 4's recommendation. Both are legitimate: the docs-backed routing advice
optimizes cost at equal quality for the general user; the operator deliberately buys maximum
quality. Chunk 4's copy must respect that (opportunity framing + suppressability), and on this
machine the finding will simply be suppressed or ignored. The plugin serves general users;
the operator's setup is not the target of the check.
## Verification
Global, after every chunk:
- `node --test 'tests/**/*.test.mjs'` → green (baseline 1359/0; count grows with new tests)
- `git status --porcelain tests/snapshots/v5.0.0/` → empty (frozen untouched)
- TDD: red test exists and fails BEFORE each production change
Per chunk:
- **C1:** `node --test tests/lib/best-practices-register.test.mjs` green;
`node scanners/knowledge-refresh-cli.mjs` classifies BP-MODEL-001/002 as fresh
- **C2:** `findNearestEffortLevel('xhig')``xhigh` (red first); `grep xhigh scanners/fix-engine.mjs` non-empty
- **C3:** fixture CLAUDE.md referencing `docs/missing.md` → finding fires; existing file /
URL / glob / placeholder / `~/` path → silent; humanized output has non-contradictory copy
- **C4:** authored-agent fixture without model/effort → opportunity fires; with either set →
silent; zero authored agents → silent; `userImpactCategory``Other` end-to-end via direct scan()
- **C5:** dogfood plan run produces a Failure-modes section; DEL B 3.2 fasit judged BEFORE the change
## Key assumptions (test before/at implementation)
1. **Per-agent `effort` frontmatter is official** — verified 2026-07-14 against
code.claude.com/docs/en/sub-agents + /model-config; re-fetch both pages at implementation
(docs move).
2. **New finding type in existing scanner leaves frozen snapshots untouched** — M-17 precedent
says yes when no v5.0.0 fixture carries the type; verify by running the suite and inspecting
which snapshots differ before committing.
3. **Next free CA-CML/CA-GAP finding numbers** — IDs are built dynamically; grep tests +
snapshots for the highest used NNN before assigning.
## Sequencing vs DEL B (one plan, no relitigation)
This plan does NOT preempt the active DEL B sequence. Recommended order:
1. DEL B step 3 pipeline dogfood (`analyze → plan → implement → rollback`) — unchanged, next.
2. Batch patch release M-11→M-17 — unchanged.
3. v5.13 chunks 1→5 (chunk 1 doubles as the `knowledge-refresh` dogfood already queued in
DEL B "Resten"; chunk 5 explicitly waits for step 3.2). Release as minor v5.13.0 via
`release-plugin.mjs` when all chunks land.

164
docs/v5.2.0-release-plan.md Normal file
View file

@ -0,0 +1,164 @@
# v5.2.0 Release — Execution Plan
**Status:** ✅ EXECUTED 2026-06-18. All edits applied, verified (875/875 tests, self-audit
configGrade A / pluginGrade A / readmeCheck.passed:true), committed and tagged `v5.2.0`. This
doc is retained as the release record. Operator chose "v5.2.0-release first, then later sessions
for net-new capabilities" (2026-06-18).
Bump: **5.1.0 → 5.2.0** (minor: one new scanner + fixes, backward-compatible). 17 commits
since `v5.1.0` tag.
---
## Pre-verified facts (do NOT re-derive)
- **Tests: 875/875 pass, 0 fail**`node --test 'tests/**/*.test.mjs'`` tests 875 / pass 875 / fail 0` (run 2026-06-18, 15.99s). 55 test files (was 52 at v5.1.0).
- **Scanner count = 13 orchestrated.** `countScannerShape` in `self-audit.mjs` returns 13:
CML, SET, HKV, RUL, MCP, IMP, CNF, TOK, CPS, DIS, COL, GAP, **SKL**. PLH excluded (standalone).
Badge `scanners-13` already correct. SKL exports `scan` + wired into scan-orchestrator (line 66).
- **No `package.json`** in this plugin. Canonical version lives in `.claude-plugin/plugin.json`.
- **self-audit README gate** (`node scanners/self-audit.mjs --json --check-readme`): currently
`passed: false`, single mismatch `{kind: tests, expected: 875, foundInReadme: 792}`. configGrade A (96), pluginGrade A (100). Fixing the tests badge → gate passes.
- **SKL history:** at `v5.1.0` the badge was `scanners-12` + prose "12" (consistent). The SKL feat
commit (`7bb2547`) bumped badge 12→13 and updated docs/scanner-internals (14 prefixes) +
humanizer-data (14 TRANSLATIONS), **but missed the 5 README prose locations** still saying "12".
Fixing those is part of shipping SKL. (Version-history row for 5.0.0 at README:602 says "→ 12
deterministic scanners" — that is HISTORICAL, leave it.)
- **Push window:** weekday 20:0023:00 CEST, weekends/holidays anytime. Re-check `date '+%u %H:%M'`
immediately before push. (At plan time: Thu 20:24 → in window.)
---
## Edit targets
### 1. Version pointers (bump)
| File:line | Find → Replace |
|-----------|----------------|
| `.claude-plugin/plugin.json` (the `"version"` field) | `"version": "5.1.0"``"version": "5.2.0"` |
| `README.md:9` badge | `version-5.1.0-blue``version-5.2.0-blue` |
| `README.md:15` badge **(required for gate)** | `tests-792+-brightgreen``tests-875+-brightgreen` |
### 2. Stale-by-SKL prose: "12" → "13" orchestrated scanners (5 exact-match edits)
Badge `scanners-13` is already correct — do NOT touch the badge. Only these prose fragments:
- `README.md:18``12 deterministic scanners across 10 quality areas``13 deterministic scanners across 10 quality areas`
- `README.md:95``(12 scanners + standalone PLH)``(13 scanners + standalone PLH)`
- `README.md:107``12 deterministic scanners verify correctness``13 deterministic scanners verify correctness`
- `README.md:336``12 Node.js scanners that perform structural analysis``13 Node.js scanners that perform structural analysis`
- `README.md:477``runs all 12 scanners + the standalone plugin-health scanner``runs all 13 scanners + the standalone plugin-health scanner`
### 3. README "What's New" section + TOC
- `README.md:24` TOC — `- [What's New in v5.1.0](#whats-new-in-v510)``- [What's New in v5.2.0](#whats-new-in-v520)`
- `README.md:48100` — replace the ENTIRE "## What's New in v5.1.0" section (heading line 48
through line 100, i.e. up to but NOT including `## What Is This?` at line 101) with the
**What's New block** below. (Verify boundaries with `grep -n '^## ' README.md` before editing —
next `## ` heading after What's New marks the end.)
### 4. README version-history table — insert new row ABOVE the `**5.1.0**` row (currently README:601)
Use the **version-history row** below.
### 5. CHANGELOG.md — insert the `[5.2.0]` block ABOVE `## [5.1.0]` (currently CHANGELOG:8)
Use the **CHANGELOG block** below.
### Leave unchanged (historical-accurate)
- CLAUDE.md humanizer "(v5.1.0)" references (lines 65, 67) — describe when the humanizer landed.
- `docs/humanizer.md` "(v5.1.0)" references.
- README:602 (5.0.0 history row).
### Open decision (operator)
- `scanners/lib/humanizer-data.mjs:2` header comment "config-audit v5.1.0" — trivial; bump to
v5.2.0 for consistency OR leave. Not required for correctness. Default: leave unless operator says bump.
---
## New content blocks
### CHANGELOG block (insert above `## [5.1.0]`)
```markdown
## [5.2.0] - 2026-06-18
### Summary
Claude Code 2.1.114→181 compatibility + skill-listing budget release. Adds a new
orchestrated scanner (SKL) for the model's skill-listing token budget, refreshes five
validators to recognize the settings/hook surface shipped across CC 2.1.114181, and
eliminates a batch of false positives surfaced by an adversarial gap-review. Scanner
internals only — no command, agent, or output-format changes.
### Added
- **`scanners/skill-listing-scanner.mjs` (SKL)** — new orchestrated deterministic scanner
(→ 13 orchestrated scanners). `CA-SKL-001` (medium): an active skill description over the
verified 1,536-char listing cap (CC 2.1.105) is silently truncated in the model's skill
listing. `CA-SKL-002` (low): the summed length of all active descriptions (each counted up
to the cap) over the listing budget (~2% of context, CC 2.1.32) — anchored on a conservative
200k window with a note that the budget scales 5× on 1M-context models; leads with the
measured sum, an estimate not telemetry. HOME-scoped (all user + plugin skills). Remediation
surfaces `disableBundledSkills` / `skillOverrides` / description trimming.
### Changed
- **settings-validator** — accepts CC 2.1.114181 settings keys and `xhigh` reasoning effort.
- **hook-validator** — recognizes `MessageDisplay` and post-session hook events (28 events).
- **claude-md-linter** — CLAUDE.md length reframed from a HIGH "adherence cliff" to a MEDIUM
token-cost finding (model-neutral, context-window aware).
- **tokens** — stale "Opus 4.7" framing refreshed to model-neutral with an Opus 4.8 anchor.
- **Knowledge corpus** — refreshed to the Opus 4.8 era (CC 2.1.114→181).
### Fixed
- **mcp-config-validator** — no longer flags auto-injected or POSIX-style env vars; removed an
invented `trust` field that does not exist in the MCP config schema.
- **permissions (DIS/CNF)** — dead-allow detection and conflict matching are now parameter-aware,
removing false positives on parameterized permission rules.
### Internal
- **Hermetic test isolation** — byte/snapshot tests and all 12 CLI-spawning test files now isolate
`HOME` via the `hermetic-home` helper, closing a leak class where the real `~/.claude` skills and
config bled into fixture-scoped runs (notably SKL, which is HOME-scoped regardless of
`includeGlobal`). Adversarially verified via a devil's-advocate gap-review pass.
### Test count
- 792 → 875 tests across 52 → 55 test files.
### Verification
- 875/875 tests pass (`node --test 'tests/**/*.test.mjs'`).
- `node scanners/self-audit.mjs --json --check-readme``configGrade: A` (96),
`pluginGrade: A` (100), `readmeCheck.passed: true`.
- README badge updated: `tests-792+``tests-875+`.
```
### What's New block (replaces README:48100)
```markdown
## What's New in v5.2.0
**Claude Code 2.1.114→181 compatibility + skill-listing budget.** A new orchestrated
scanner, **SKL**, checks the model's skill-listing token budget: `CA-SKL-001` flags any
active skill description over the 1,536-char listing cap (silently truncated by CC 2.1.105),
and `CA-SKL-002` flags when the summed descriptions exceed the listing budget (~2% of context).
Five validators were refreshed for the settings and hook surface that shipped across
CC 2.1.114181 (new settings keys, `xhigh` effort, `MessageDisplay` + post-session hook
events), and an adversarial gap-review eliminated a batch of false positives in the MCP and
permissions scanners. → **13 orchestrated deterministic scanners** (+ standalone plugin-health).
```
### Version-history row (insert above the `**5.1.0**` row at README:601)
```markdown
| **5.2.0** | 2026-06-18 | CC 2.1.114→181 compatibility + skill-listing budget. New orchestrated scanner **SKL** (`CA-SKL-001` 1,536-char listing cap, `CA-SKL-002` listing-budget sum) → 13 orchestrated scanners. Five validators refreshed for CC 2.1.114181 settings/hook surface (`xhigh` effort, `MessageDisplay` + post-session events, 28 hook events). False positives eliminated in MCP (auto-injected/POSIX env vars, invented `trust` field) and permissions (param-aware DIS/CNF). Hermetic HOME isolation across all CLI-spawning tests. 875 tests |
```
---
## Verification (after edits, before commit)
1. `node scanners/self-audit.mjs --json --check-readme``readmeCheck.passed: true`, `mismatches: []`, configGrade A, pluginGrade A.
2. `grep -nc "12 deterministic\|(12 scanners +\|12 Node.js scanners\|all 12 scanners" README.md``0` (no stale current-state prose; the 5.0.0 history row at :602 is allowed and uses "→ 12 deterministic scanners" which these patterns do not match).
3. `node --test 'tests/**/*.test.mjs'` → 875/875, no snapshot drift.
4. `git diff tests/snapshots/` → empty (contamination check; grep for sadhguru/vegnorm/linkedin/ms-ai if any snapshot changed).
5. `grep version .claude-plugin/plugin.json``5.2.0`; README badges show `version-5.2.0` + `tests-875+`.
## Commit / tag / push
1. `git add -A` (CHANGELOG.md, README.md, .claude-plugin/plugin.json [+ humanizer-data.mjs if operator approved]).
2. Commit: `release: v5.2.0 — CC 2.1.114→181 compat + skill-listing budget (SKL)` + standard Co-Authored-By/Session trailers.
3. `git tag v5.2.0`.
4. **Re-check push window** `date '+%u %H:%M'`. In window (weekday 20:0023:00 / weekend / holiday) → `git push origin main && git push origin v5.2.0`. Out of window → hold both, tell operator when window opens.
5. After: overwrite STATE.md (release done), and the net-new-capabilities track becomes the next open item.

240
docs/v5.3.0-release-plan.md Normal file
View file

@ -0,0 +1,240 @@
# v5.3.0 Release — Multi-Session Plan
_Status: PLANNED. No version bump yet. `plugin.json` = 5.2.0. All work below is gated on
explicit operator GO per session. Plans live next to STATE per the continuity system._
## Goal
Release **v5.3.0** (minor, backward-compatible) covering the work that has accumulated on
`main` since the v5.2.0 release (`1576909`, 2026-06-18), with a **fully updated README +
CHANGELOG**. The features are already implemented and on `main`; the bulk of this release is
documentation, changelog narrative, version sync, and a tag — plus a decision on whether any
remaining CC-gap items should ship in 5.3.0 first.
## Pre-verified facts (do NOT re-derive)
- Last release: **v5.2.0** = commit `1576909`, 2026-06-18. Prior: v5.1.0 (2026-05-01).
- Unreleased commits on `main` since v5.2.0 (8):
| Commit | Type | 5.3.0 changelog section |
|--------|------|-------------------------|
| `0a631e3` | refactor(skl): extract skill-listing budget to shared lib | Internal |
| `dfe9049` | feat(feature-gap): `disableBundledSkills` under listing pressure | Added |
| `03949c6` | feat(dis): ineffective allow wildcards; `Tool(*)` = deny-all | Added |
| `b0bf8c5` | feat(cml): context-window-scaled CLAUDE.md char budget | Added |
| `d678765` | feat(dis): forbidden-param permission rules | Added |
| `c6c5f17` | feat(plh): plugin namespace collision (same declared name) | Added |
| `0874188` | fix(plh): cross-plugin command overlap HIGH→LOW | **Changed** |
| `96743ec` | fix(readme): add SKL row to scanner table `[skip-docs]` | Fixed (docs) |
- **Scanner count unchanged: 13.** The 7 features extended EXISTING scanners (CML/DIS/PLH/GAP) —
no new scanner file. So the `scanners-13` badge, the table (now 13 rows after `96743ec`), and
`countScannerShape` need **no change** for 5.3.0. Do not bump the scanner badge.
- **One behavior change to call out** (`0874188`): the cross-plugin command-name finding went
**HIGH → LOW** and was reframed from "conflict" to "ambiguity" (namespacing keeps both
commands reachable). Scoring impact: any config with cross-plugin command-name overlap scores
slightly higher now. Not a `--json`-shape break; no test asserted the old HIGH (verified
2026-06-19). Document under **Changed**, not Breaking.
- Suite at HEAD: **936/936**. self-audit: configGrade A 97, pluginGrade A 100, readmeCheck.passed.
- Edit-target files (from the v5.2.0 release pattern, `docs/v5.2.0-release-plan.md`):
`.claude-plugin/plugin.json`, `README.md` (badges + What's New + version-history row + TOC),
`CHANGELOG.md` (`## [5.3.0]` block above `## [5.2.0]`), and `docs/cc-2.1.x-gap-matrix.md`.
---
## Session A — Release-readiness audit & scope decision (read/plan; 1 doc output)
**No code changes.** Decide exact 5.3.0 scope and surface the "build more first?" question.
1. **Reconcile `docs/cc-2.1.x-gap-matrix.md` against current code.** STATE flags it as stale.
For each matrix row, mark which of the 7 feature commits closed it; list every still-open gap.
2. **Classify each still-open gap**`ship-in-5.3.0` / `defer-to-5.4` / `wontfix`, with a
one-line rationale and effort tag. **Operator GO required** on the ship-list before any build.
3. **Draft the CHANGELOG narrative** — map all 8 unreleased commits to Added/Changed/Fixed/Internal
bullets (table above is the starting point). Draft the "What's New in v5.3.0" prose.
4. **Confirm version = 5.3.0** (minor) and write the Changed-note for the PLH severity downgrade.
5. **Knowledge review** — check whether any `knowledge/*.md` corpus file needs an update for the
7 features (e.g. permission-rule or plugin-namespacing corpus). List files to touch, if any.
**Output:** updated `cc-2.1.x-gap-matrix.md`; a `## v5.3.0 scope decision` block appended to THIS
file (ship-list + operator GO + draft changelog bullets + knowledge-touch list).
**Verifisering (testbar):**
- Every one of the 8 unreleased commits appears in exactly one draft changelog bullet.
- Every still-open matrix gap has a `ship/defer/wontfix` verdict recorded.
- Operator GO on the ship-list is recorded verbatim in the scope-decision block.
- `git log 1576909..HEAD` count matches the number of mapped commits (no commit dropped).
---
## Session B — (conditional) implement approved remaining features
**Runs only if Session A's ship-list is non-empty.** If empty → skip entirely; 5.3.0 is a
docs-and-release-only version.
- One feature per chunk, **/tdd: failing test FIRST** (Iron Law), then minimal implementation.
- Per feature commit: self-audit gates + docs-gate (README **and** CLAUDE ≥3 substantive lines, or
`[skip-docs]` for non-feature commits). Stage docs in a separate Bash call before commit.
- If a feature adds a **new scanner**, THEN (and only then) bump the `scanners-` badge, the table,
and `countScannerShape` — and note it for Session C's "What's New".
- Checkpoint STATE + this plan after each feature (chunk-before-compaction).
**Verifisering (per feature):** new tests RED→GREEN; full suite N/N; self-audit configGrade ≥ A,
pluginGrade ≥ A, readmeCheck.passed; contamination-grep + gitleaks clean on any snapshot/fixture touch.
---
## Session C — Release v5.3.0 (version sync + docs + tag + push)
1. **Bump** `.claude-plugin/plugin.json` 5.2.0 → 5.3.0. (Version lives ONLY here per STATE.)
2. **README:**
- `version-` badge → 5.3.0; `tests-` badge → current exact case count.
- Replace "What's New in v5.2.0" with **What's New in v5.3.0** (+ update the TOC anchor at the top).
- Insert a new **version-history table row** above the `**5.2.0**` row.
- Sweep for stale prose. **Scanner count stays 13** unless Session B added a scanner.
3. **CHANGELOG.md:** insert `## [5.3.0] - <release-date>` block above `## [5.2.0]`
(Added / Changed / Fixed / Internal / Test count / Verification — mirror the 5.2.0 block).
4. **Knowledge/docs:** apply any updates Session A flagged.
5. **Marketplace cross-repo (CONDITIONAL — verify first):** check whether the marketplace catalog
(`../.claude-plugin/marketplace.json` or the marketplace README) references config-audit's
version or feature list. If yes → update there too. Separate repo → **same push-window rules,
separate push**. (Not yet confirmed to exist; Session C verifies before assuming.)
6. **Gates (all must pass before commit):**
- `node --test 'tests/**/*.test.mjs'` → N/N green.
- `node scanners/self-audit.mjs --json --check-readme` → readmeCheck.passed, configGrade A,
pluginGrade A, scanner count consistent with badge+table.
- SC-5 snapshot byte-equal (or re-seeded + contamination-grep clean); gitleaks clean.
7. **Commit** `release: v5.3.0 — <one-line summary>`, **tag** `v5.3.0`, **push** (plugin repo;
marketplace repo if step 5 applies) inside the push window.
**Verifisering (testbar):**
- `git tag --list v5.3.0` returns `v5.3.0`.
- README `version-` badge value == `.claude-plugin/plugin.json` `version`.
- `self-audit --check-readme``readmeCheck.passed: true`, badge tests == suite case count.
- `git log -1` subject starts with `release: v5.3.0`.
- If marketplace updated: its config-audit entry shows 5.3.0.
---
## Key assumptions (test, don't trust)
- **5.3.0 is non-breaking.** Verify: `json-backcompat` + `raw-backcompat` + snapshot tests green
(they are at HEAD); the PLH severity change alters a severity VALUE, not the `--json` shape, and
no test/consumer asserts the old HIGH. If Session B adds anything that changes the JSON shape or
scoring math materially → re-classify as a Breaking note (like the 5.0.0 row).
- **docs-gate** blocks any commit lacking ≥3 substantive lines in BOTH README and CLAUDE — use
`[skip-docs]` for pure release/doc commits, or make both edits substantive. (Proven 2026-06-19.)
- **self-audit `--check-readme`** only checks badge NUMBERS vs filesystem, NOT prose-table
completeness — so manually re-verify the scanner table and "What's New" by eye (the SKL-row gap
`96743ec` slipped exactly because the gate is number-only).
## Out of scope (do not start without separate GO)
- New CC-gap features beyond Session A's approved ship-list.
- Any scanner refactor not required by an approved feature.
- Marketplace-wide changes beyond the config-audit catalog entry.
---
## v5.3.0 scope decision (Session A output, 2026-06-19)
### Operator GO (recorded verbatim)
> **"Release-only, ship-list tom"** — Session B SKIP; knowledge-backing (disableBundledSkills,
> 40k char-budget, forbidden-param) folded into Session C; open M-gaps → v5.4 backlog.
**Ship-list = EMPTY.** Session B is skipped entirely. v5.3.0 is a docs + knowledge + release
of the work already on `main`. No new feature build.
### Reconciliation result (matrix vs. current code)
The gap-matrix was the **v5.2.0** plan. Verified against HEAD: the entire HIGH-priority
false-positive cluster is **already CLOSED**`settings-validator` (all 11 keys + `xhigh`),
`hook-validator` (28 events incl. `MessageDisplay`/`post-session`), `mcp-config-validator`
(POSIX/auto-injected env + `trust` removed), `claude-md-linter` (HIGH→MEDIUM reframe), DIS/CNF
(param-aware). The 8 unreleased commits add incrementally on top. **No active false positive or
behavior bug remains unshipped.** Full row-by-row verdicts appended to `cc-2.1.x-gap-matrix.md`.
### Still-open gaps → verdicts (all defer; none are bugs)
| Gap (matrix row) | Effort | Verdict |
|---|---|---|
| PLH shadow-folder: plugin.json key shadows default component folder (129) | M | defer-5.4 |
| permissions: acceptEdits prompts on shell-startup/build-config writes (141) | M | defer-5.4 |
| permissions: Read-deny hides Glob/Grep; Windows backslash/case path (140) | M | defer-5.4 |
| settings: autoMode.hard_deny nested-structure validation (144) | M | defer-5.4 |
| PLH: validate `skills:` array entries are dirs within plugin (130) | M | defer-5.4 |
| CLAUDE.md: nested `.claude` closest-wins conflict detection (104, 131) | M | defer-5.4 |
| Knowledge L-rows: env vars, model-lineup nuance, hook-output fields, plugin/skill doc rows | S | defer (rolling knowledge maintenance) |
### Draft CHANGELOG narrative (all 9 commits `1576909..HEAD` accounted for)
**Added**
- **DIS forbidden-param rules** (`d678765`) — flags `Tool(param:value)` whose key is the tool's
own canonicalizing field (`command`/`file_path`/`path`/`notebook_path`/`url`); CC ignores these
and emits a startup warning. Severity by intent: **deny/ask = false security (medium)**,
**allow = dead config (low)**. Valid forms (`Bash(npm:*)`, `WebFetch(domain:host)`,
`Agent(model:opus)`) never flagged. Predicate `forbiddenParamRule` in `permission-rules.mjs`.
- **DIS ineffective allow wildcards + `Tool(*)` deny-all** (`03949c6`) — flags unanchored
tool-name globs in `permissions.allow` (`*`, `B*`, `mcp__*`) that CC silently skips (low);
treats `Tool(*)` as deny-all (`Bash(*)``Bash`) so a bare allow killed by it is reported
as dead config. Valid `mcp__<server>__*` never flagged.
- **CML context-window-scaled char budget** (`b0bf8c5`) — new `CA-CML` finding mirroring CC's
startup warning *"Large CLAUDE.md will impact performance (X chars > 40.0k)"*. Anchors on the
conservative 200k window; discloses the relaxed ~200,000-char figure at 1M context. Severity
**medium** (token cost). Char-keyed, complementary to the 200/500-line checks. Window constants
in shared `scanners/lib/context-window.mjs`.
- **PLH plugin namespace collision** (`c6c5f17`) — flags 2+ discovered plugins declaring the same
`name` in `plugin.json`. Namespaces collapse; CC picks an undocumented winner; the loser's
commands/skills/agents go silently unreachable. Severity **medium** (dead config),
`category: 'plugin-hygiene'`, COL-shaped `details.namespaces`. Keys on declared `name`, not
`basename(dir)`; name-less plugins excluded.
- **feature-gap `disableBundledSkills` lever** (`dfe9049`) — conditional recommendation to set
`disableBundledSkills` when the active skill listing is over budget; remediation companion to
SKL `CA-SKL-002`.
**Changed**
- **PLH cross-plugin command-name overlap: HIGH → LOW** (`0874188`) — reframed from "conflict" to
"ambiguity". Commands are namespaced (`/name:command`) so both stay reachable; only a same-`name`
plugin collision loses components (covered by the new namespace-collision finding). Now
group-first (one finding per command name listing every namespace), COL-shaped
`details.namespaces`. The old HIGH `Cross-plugin command name conflict` finding **and its
humanizer entry are removed**. Scoring impact: configs with cross-plugin command overlap score
slightly higher. **Not a `--json`-shape break** — no test/consumer asserted the old HIGH.
**Fixed**
- **README scanner table** (`96743ec`) — added the missing SKL row (12 → 13 rows); the prose table
lagged the badge/self-audit count (the `--check-readme` gate is number-only and didn't catch it).
**Internal**
- **Extract skill-listing budget to shared lib** (`0a631e3`) — `scanners/lib/context-window.mjs`
as single source of truth for the 200k/1M window constants, re-exported by
`skill-listing-budget.mjs` and consumed by the new CML char-budget finding.
**Excluded from changelog:** `9b828fa` `docs(plan): v5.3.0 multi-session release plan [skip-docs]`
— the plan meta-commit itself, not a product change. (9 commits total = 8 mapped + 1 excluded.)
### What's New in v5.3.0 (draft prose)
> **v5.3.0 — permission-rule & plugin-hygiene hardening.** Five additive scanner findings extend
> the permission, CLAUDE.md, and plugin surfaces: DIS now catches forbidden-parameter rules and
> ineffective allow-wildcards that Claude Code silently ignores; CML mirrors CC's own 40.0k-char
> "large CLAUDE.md" startup warning, context-window scaled; PLH flags two plugins that declare the
> same name (one silently shadows the other); and feature-gap recommends `disableBundledSkills`
> when the skill listing is over budget. The cross-plugin command-name finding is reframed from a
> HIGH conflict to a LOW ambiguity to match Claude Code's namespacing. Scanner count stays **13**
> (all extend existing scanners). `--json`/`--raw` remain byte-stable.
### Version & non-breaking confirmation
- **5.3.0** (minor). Additive findings + one severity downgrade (PLH HIGH→LOW). No envelope/JSON
shape change; `json-backcompat` + `raw-backcompat` + SC-5 snapshot green at HEAD. The PLH change
alters a severity VALUE, not the shape. **Document the downgrade under Changed, not Breaking.**
### Knowledge-touch list (apply in Session C step 4)
Three short backing entries for the v5.3.0 features (rows currently absent from the corpus —
verified by grep, 0 files each). Pure docs, no scanner/test impact:
1. `knowledge/prompt-cache-patterns.md``disableBundledSkills` as a token-efficiency lever
(hides bundled-skill descriptions from the system prompt). (matrix row 167)
2. `knowledge/claude-code-capabilities.md` *or* `prompt-cache-patterns.md` — the CC 40.0k-char
"Large CLAUDE.md will impact performance" startup warning + 2.1.169 context-window scaling
(backs the new CML finding).
3. `knowledge/claude-code-capabilities.md``Tool(param:value)` permission semantics: param
matching is deny/ask-only, and the canonicalizing-field rules CC ignores (backs DIS
forbidden-param). (matrix rows 137, 139)
Everything else in the matrix (env-var rows, model-lineup nuance, hook-output fields, plugin/skill
doc rows) → **defer** to rolling knowledge maintenance; not required for 5.3.0 correctness.

258
docs/v5.4.0-release-plan.md Normal file
View file

@ -0,0 +1,258 @@
# v5.4.0 Release — Multi-Session Plan
_Status: PLANNED. No version bump yet. `plugin.json` = 5.3.0. All work below is gated on
explicit operator GO per session. Plans live next to STATE per the continuity system.
Mirrors the structure of `docs/v5.3.0-release-plan.md`._
## Goal
Release **v5.4.0** (minor, backward-compatible): three **net-new additive scanner findings**
that extend the plugin-hygiene (PLH) and settings (SET) surfaces, mirroring real Claude Code
warnings/validation that config-audit does not yet detect. Unlike v5.3.0 (which documented work
already on `main`), **the features here do not exist yet** — Session B is the core build, not a
conditional. No bugs/false-positives are involved; all three are enhancements with confirmed CC
premises.
**Theme:** *plugin-hygiene & settings-validation hardening.*
## Pre-verified facts (do NOT re-derive)
- Last release: **v5.3.0** = commit `fe686b6`, tag `v5.3.0`, 2026-06-19. `git rev-list --count
v5.3.0..HEAD` = **0** → working tree is at the tag; v5.4.0 builds net-new code.
- Suite at HEAD: **936/936** (clean tree, unchanged since the v5.3.0 release). self-audit at
v5.3.0: configGrade A 97, pluginGrade A 100, readmeCheck.passed.
- **Scanner count = 13** (authoritative source `countScannerShape`, `scanners/self-audit.mjs:56`;
README badge `scanners-13`). All three approved features **extend existing scanners** (PLH, SET)
**no new scanner file**, so the badge, the scanner table, and `countScannerShape` need **no
change**. (An early grep of `export async function scan` counted 15 shapes — that set includes
lib/non-badge shapes like CNF/IMP/SAL and is NOT the badge count. Do not bump the badge.)
- PLH currently emits `CA-PLH-001`..`CA-PLH-014` (014 = namespace collision, 013 = command-name
ambiguity, both shipped v5.3.0). Next free: `CA-PLH-015`, `CA-PLH-016`.
- Edit-target files (v5.3.0 release pattern): `.claude-plugin/plugin.json`, `README.md`
(badges + What's New + version-history row + TOC), `CHANGELOG.md` (`## [5.4.0]` above `[5.3.0]`),
`docs/cc-2.1.x-gap-matrix.md`, the two target scanners, their tests, and 3 knowledge entries.
- **CC premises primary-source verified** (code.claude.com/docs, 2026-06-19) — see Session A
reconciliation. The matrix is NOT ground truth; row 175 was found to be framed backwards.
---
## Session A — Reconciliation & scope decision (read/plan; this session) ✅ DONE
**No code changes.** Re-reconcile the gap-matrix against HEAD, verify each candidate is still open
AND that its CC premise is real, classify ship/defer/wontfix, get operator GO on the ship-list,
draft the changelog. Output recorded below + a `## v5.4.0 reconciliation` block appended to
`docs/cc-2.1.x-gap-matrix.md`.
**Verifisering (testbar):**
- Every v5.4 candidate has a recorded code-state verdict (IMPLEMENTED/OPEN/PARTIAL) with file:line.
- Every candidate has a CC-premise verdict (CONFIRMED/REFUTED/UNVERIFIABLE) with a source.
- Ship-list has a recorded operator GO before any build. ✅
- The refuted premise (#3) is corrected in the gap-matrix. ✅
---
## Session B — Implement the 3 approved features (CORE build)
**One feature per chunk. Iron Law: failing test FIRST (/tdd), then minimal implementation.**
Re-confirm each feature's exact CC behavior against primary docs at the start of its chunk
(Verifiseringsplikt — premises were verified in Session A but the exact field lists / warning text
must be pinned before coding). Checkpoint STATE + this plan after each feature.
### Feature 1 — PLH shadow-folder (`CA-PLH-015`) ✅ DONE (commit `7abc5a1`)
- **Verifiseringsplikt correction (2026-06-19):** the field set was pinned against
`code.claude.com/docs/.../path-behavior-rules`. Only the **replaces** category truly shadows:
shipped set = `commands`/`agents`/`outputStyles`. **`skills` excluded** — it *adds to* the
default `skills/` scan (both load, never a shadow); **`hooks`/`mcpServers`/`lspServers`
excluded** — own merge rules, not a folder-shadow. Experimental `themes`/`monitors` omitted
(docs warn their schema may change). Explicit-address exception honored
(`"commands": ["./commands/x.md"]` not flagged). +5 tests (936→941), suite green, self-audit
A/A, count 13, gitleaks clean. Pushed.
- **What:** flag when a `plugin.json` component-path key (`commands`/`agents`/`outputStyles`)
points at a **custom path** while the **default folder of that name also exists** in the plugin,
so the default is silently ignored. Mirrors CC's warning in `/doctor` & `claude plugin list`
(~2.1.140). _(Original plan listed `commands/agents/skills/hooks` — corrected above.)_
- **Where:** `scanners/plugin-health-scanner.mjs` (per-plugin loop; the scanner already reads
`plugin.json` but only checks `name`/`description`/`version`). Add component-path-key parsing.
- **Severity:** MEDIUM — silently-shadowed components = dead config (a whole command/skill/agent
set goes unreachable). `category: 'plugin-hygiene'`.
- **Tests:** fixture plugin with both `"commands": "custom/"` AND a `commands/` folder → flagged;
fixture with only one of the two → not flagged; fixture with no component keys → not flagged.
### Feature 2 — PLH `skills:`-array validation (`CA-PLH-016`) ✅ DONE (commit `76d5eda`)
- **Verifiseringsplikt correction (2026-06-19):** the plan's claim that the CC error "suggests the
parent directory when an entry points at a file" is **not** in the primary docs — dropped. The
finding asserts only the four primary-source-verified conditions. Escape backed by docs'
path-traversal rule (*"Installed plugins cannot reference files outside their directory … such as
`../shared-utils`"*). string|array normalized, so a non-string top-level value (e.g. `42`) is
caught as one non-string entry. +3 tests (941→944), suite green, self-audit A/A, count 13.
- **What:** when `plugin.json` has a `skills:` field (string|array), validate each entry resolves to
an existing **directory inside the plugin root**. One finding per bad entry, `problem`
{`non-string`, `escapes-root`, `not-found`, `not-a-directory`}. Mirrors `claude plugin validate`.
- **Where:** `scanners/plugin-health-scanner.mjs` (read `parsed.skills` after the parse-success
guard).
- **Severity:** MEDIUM (broken/partly-broken plugin manifest). `category: 'plugin-hygiene'`.
- **Tests:** fixture `skills: ["valid-dir", "points-to-file.md", "missing", "../escape"]` → one
finding per bad entry, none for `valid-dir`; fixture with no `skills:` key → not flagged.
### Feature 3 — SET autoMode structure + dead-config (`CA-SET-NNN`) ✅ DONE (commit `3633571`)
- **Premise CONFIRMED against primary source** (`code.claude.com/docs/en/auto-mode-config`) — the
plan was correct this time (unlike #1/#2). 4 sub-keys verbatim; `"$defaults"` valid; exact scope
exclusion quote confirmed. **Per-file-scope gate PASSED:** `ConfigFile` already carries `scope`
(`classifyScope``'project'` for shared `.claude/settings.json`), so the dead-config sub-check
shipped (no fallback to structure-only needed). SET is in the orchestrator → **SC-5 re-checked,
byte-equal** (snapshot fixture has no autoMode). +5 tests (944→949), suite green, self-audit A/A,
count 13, gitleaks clean. Fixtures force-added (`.claude/` gitignored).
- **What (two sub-checks):**
1. **Structure validation**`autoMode`, if present, must be an object whose only keys are
`environment`, `allow`, `soft_deny`, `hard_deny`, each a **string array** (entries are prose
rules; `"$defaults"` is a valid literal). Flag unknown sub-keys and wrong value types.
2. **Dead-config**`autoMode` placed in **shared project settings** (`.claude/settings.json`)
is **ignored by CC** ("Not read from shared project settings" — official docs). Flag it as
dead config; valid scopes are user (`~/.claude/settings.json`), local
(`.claude/settings.local.json`), and managed.
- **Where:** `scanners/settings-validator.mjs` (`autoMode` is already in `KNOWN_KEYS:21` but has no
TYPE_CHECKS / nested validation). The dead-config sub-check needs the settings file's **scope**;
confirm the scanner already knows per-file scope (it validates user/project/local files).
- **Severity:** structure = LOW/MEDIUM (malformed config); dead-config = LOW (dead config).
- **Tests:** valid autoMode in user settings → clean; unknown sub-key / non-array value → flagged;
autoMode in shared `.claude/settings.json` → dead-config finding.
**Per-feature gates (all must pass before that feature's commit):**
- New tests RED→GREEN; full suite N/N green.
- `node scanners/self-audit.mjs --json --check-readme` → configGrade ≥ A, pluginGrade ≥ A,
readmeCheck.passed, scanner count still 13.
- Any snapshot/fixture touch → SC-5 re-seed + contamination-grep clean; gitleaks clean.
- **docs-gate** (`feat:` commits): README **and** CLAUDE.md each get ≥3 substantive lines
documenting the new finding — OR `[skip-docs]`. Prefer real docs (each finding deserves a
CLAUDE.md "Scanner" note + README mention). Stage docs in a separate Bash call before commit.
- Humanizer: each new finding needs a `userImpactCategory` mapping (dead config / conflict /
configuration mistake) — verify `--json`/`--raw` stay byte-stable for *existing* findings
(`json-backcompat` + `raw-backcompat` green); new findings are additive.
**Scanner count stays 13.** No badge/table/`countScannerShape` change (no new scanner file).
---
## Session C — Release v5.4.0 (version sync + docs + tag + push)
1. **Bump** `.claude-plugin/plugin.json` 5.3.0 → 5.4.0. (Version lives ONLY here.)
2. **README:** `version-` badge → 5.4.0; `tests-` badge → exact suite count after Session B
(will rise — new tests added). Replace "What's New in v5.3.0" → **v5.4.0** (+ TOC anchor).
Insert a new version-history row above `**5.3.0**`. **Scanner count stays 13.**
3. **CHANGELOG.md:** insert `## [5.4.0] - <date>` above `## [5.3.0]` (Added / Test count /
Verification — mirror the 5.3.0 block; this release is Added-only + Internal if any).
4. **Knowledge:** apply the 3 backing entries (list below).
5. **Marketplace cross-repo (`../catalog/`, separate repo):** bump the config-audit `ref`/version
in `marketplace.json` + README **after** the plugin tag is pushed. Same push-window rules,
separate push. **Cross-repo edit requires operator GO.**
6. **Gates:** suite N/N; self-audit configGrade A, pluginGrade A, readmeCheck.passed, count 13;
SC-5 byte-equal or re-seeded + contamination-grep clean; gitleaks clean (both repos).
7. **Commit** `release: v5.4.0 — <one-line>`, **tag** `v5.4.0`, **push** (in push window).
**Verifisering (testbar):** `git tag --list v5.4.0` returns it; README `version-` badge ==
`plugin.json` version; `self-audit --check-readme` passed + badge tests == suite count;
`git log -1` subject starts `release: v5.4.0`; marketplace entry shows 5.4.0.
---
## Key assumptions (test, don't trust)
- **5.4.0 is non-breaking.** Three additive findings; no envelope/JSON-shape change. Verify
`json-backcompat` + `raw-backcompat` + SC-5 snapshot stay green; new findings appear only in
configs that trigger them. If anything changes the JSON shape or existing-finding output → it is
a Breaking note, not Added.
- **settings-validator knows per-file scope.** Feature 3's dead-config sub-check needs to know a
settings file is the *shared project* one. **Test this assumption first in Feature 3's chunk**
if the scanner doesn't already carry scope per file, either thread it through or drop the
dead-config sub-check to structure-only (still ships).
- **docs-gate blocks `feat:` commits** lacking ≥3 substantive lines in BOTH README and CLAUDE —
proven 2026-06-19. (`release:`/`fix:`/`chore:`/`docs:` pass freely.)
- **self-audit `--check-readme` is number-only** (badge vs filesystem) — it does NOT catch prose
gaps. Manually eyeball the scanner table + "What's New" (the SKL-row gap slipped exactly here).
## Knowledge-touch list (Session C step 4)
Three short backing entries (mirrors v5.3.0's 3-entry pattern; grep-verify each is absent first):
1. `knowledge/claude-code-capabilities.md` — plugin component-path keys shadow default folders;
CC warns in `/doctor` & `claude plugin list` (~2.1.140/142). Backs `CA-PLH-015`.
2. `knowledge/claude-code-capabilities.md``claude plugin validate` requires `skills:` entries to
be directories (~2.1.145). Backs `CA-PLH-016`.
3. `knowledge/claude-code-capabilities.md` *or* `configuration-best-practices.md``autoMode`
schema (`environment`/`allow`/`soft_deny`/`hard_deny` string arrays, `"$defaults"` literal) and
the "not read from shared project settings" rule. Backs `CA-SET` autoMode finding.
Everything else in the matrix (env-var rows, model-lineup nuance, hook-output fields, nested
.claude doc) → **defer** to rolling knowledge maintenance.
## Out of scope (do not start without separate GO)
- **#2 acceptEdits-on-shell/build-config** (deferred — needs the exact CC special-cased file-category
list primary-source-confirmed before it's buildable). Candidate for a later release.
- **#6 nested-.claude closest-wins** (deferred — a NEW scanner: ancestor-chain walker + grouping +
badge bump 13→14. Heaviest item; natural headline for its own release, e.g. v5.5.0).
- **#3 Read-deny/Glob-Grep** — **wontfix** (premise refuted by primary source; see reconciliation).
- Any scanner refactor not required by the 3 approved features; marketplace-wide changes beyond the
config-audit catalog entry.
---
## v5.4.0 scope decision (Session A output, 2026-06-19)
### Operator GO (recorded verbatim)
> **"Option A"** — ship #1 PLH shadow-folder + #5 PLH skills:-array + #4 autoMode structure/
> dead-config. No new scanner (badge stays 13). #2 and #6 deferred, #3 wontfix.
**Ship-list (3 features):** `CA-PLH-015` shadow-folder · `CA-PLH-016` skills:-array · `CA-SET-NNN`
autoMode structure + dead-config. All extend existing scanners.
### Reconciliation result (6 candidates vs HEAD, code-state + CC-premise)
| # | Candidate (matrix row) | Code-state @HEAD | CC premise | Verdict |
|---|---|---|---|---|
| 1 | PLH shadow-folder (164) | OPEN — `plugin-health-scanner` parses only name/desc/version | CONFIRMED (~2.1.140; `/doctor`+`plugin list` warning) | **ship-5.4** |
| 5 | PLH `skills:`-array dirs (165) | OPEN — does not read `parsed.skills` | CONFIRMED (~2.1.145 `claude plugin validate`) | **ship-5.4** |
| 4 | autoMode.hard_deny structure (179) | OPEN — key in `KNOWN_KEYS:21`, no TYPE_CHECKS/nested val | CONFIRMED (primary source: 4 string-array sub-keys; not read from shared project settings) | **ship-5.4** |
| 2 | acceptEdits-writes shell/build (176) | OPEN — 0 matches in `scanners/` | CONFIRMED (~2.1.160) but exact file-category list unpinned | **defer** (needs primary-source field list) |
| 6 | nested-.claude closest-wins (139/166) | OPEN — no scanner; CML walks only the CLAUDE.md cascade | CONFIRMED (~2.1.178) | **defer** (NEW scanner, badge bump) |
| 3 | Read-deny hides Glob/Grep (175) | PARTIAL — hint in `permission-rules.mjs:158-159`, but `dominates()` treats Read/Glob/Grep as distinct | **REFUTED** — docs: "Claude makes a best-effort attempt to apply `Read` rules to … Grep and Glob"; Read-deny already covers them | **wontfix** |
**Matrix correction (Verifiseringsplikt):** row 175's framing was **backwards**. CC's permissions
doc states Read deny rules already apply (best-effort) to Glob/Grep, so a "Read-deny is bypassable"
finding would be a false positive — the same failure mode as the invented MCP `trust` field. The
only real Read-deny bypass is via Bash subprocesses (a python/node script that opens files), which
is already documented behavior, not a config mistake to flag. The Windows-path half (CC normalizes
`C:\``/c/`) is real but narrow/low-value and is folded into "defer rolling maintenance".
### Draft CHANGELOG narrative (finalize after Session B)
**Added**
- **PLH plugin-folder shadowing** (`CA-PLH-015`) — flags a `plugin.json` component-path key
(`commands`/`agents`/`skills`/`hooks`) that points at a custom path while the default folder of
that name also exists, so one silently shadows the other. Mirrors Claude Code's `/doctor` &
`claude plugin list` warning. Severity **medium** (dead config), `category: 'plugin-hygiene'`.
- **PLH `skills:`-array validation** (`CA-PLH-016`) — validates each entry in a `plugin.json`
`skills:` array resolves to a directory inside the plugin; flags non-string entries, missing
paths, file-not-directory, and path-escape. Mirrors `claude plugin validate`. Severity
**medium**, `category: 'plugin-hygiene'`.
- **SET autoMode structure + dead-config** (`CA-SET-NNN`) — validates `autoMode` contains only
the four known sub-keys (`environment`/`allow`/`soft_deny`/`hard_deny`), each a string array,
and flags `autoMode` placed in shared project settings (`.claude/settings.json`), which Claude
Code ignores. Structure = low/medium, dead-config = low.
**Internal / Test count / Verification** — fill in Session C (mirror the 5.3.0 block).
Scanner count stays **13** (all extend existing PLH/SET scanners). `--json`/`--raw` byte-stable.
### What's New in v5.4.0 (draft prose)
> **v5.4.0 — plugin-hygiene & settings-validation hardening.** Three additive findings extend the
> plugin and settings surfaces: PLH now flags a plugin.json component-path key that silently shadows
> the default folder of the same name (mirroring Claude Code's `/doctor` warning), and validates
> that `skills:`-array entries resolve to real directories (mirroring `claude plugin validate`); the
> settings validator now checks the structure of `autoMode` and flags it when placed in shared
> project settings, where Claude Code ignores it. Scanner count stays **13** (all extend existing
> scanners). `--json`/`--raw` remain byte-stable.
### Version & non-breaking confirmation
- **5.4.0** (minor). Additive findings only; no envelope/JSON-shape change. `json-backcompat` +
`raw-backcompat` + SC-5 snapshot green at HEAD; new findings are additive (appear only in configs
that trigger them).

View file

@ -0,0 +1,245 @@
# v5.5+ Plan — Steering-Model Coverage
_Status: PLAN (awaiting GO per feature). Created 2026-06-20. Plans live next to the work (continuity rule)._
## Why this plan exists
A "seven ways to steer Claude Code" framing (CLAUDE.md / rules / skills / sub-agents /
hooks / output styles / mechanism-fit) was used as a **check-against** reference — not a
spec. Checking it against the live docs (`code.claude.com/docs`, map updated 2026-06-19)
both corrected the video **and** surfaced gaps in config-audit's own coverage.
The unifying insight: every steering mechanism has a different **loading model** and
**compaction-survival** profile, and the docs now publish both explicitly. config-audit
already audits *structure* well; it does **not** yet model *when a thing is loaded* or
*whether it survives compaction* — which is exactly the "wrong mechanism = wasted tokens /
unloaded instruction" axis. This plan closes that.
The small correctness items the check surfaced (HKV hook-events, RUL globs wording) are
**already landed** on `main` (`b6a62d7`) and bound for the pending **v5.4.1** patch. This
document covers only the **new functionality** (v5.5+), greenlit by the operator: A, B, C,
D, E.
---
## Verification log (Verifiseringsplikt)
Every claim this plan builds on, with source and status. `CONFIRMED` = stated in primary
docs; `REFUTED` = docs contradict; `UNVERIFIED` = not found in primary docs (do not assert).
| # | Claim | Status | Source |
|---|-------|--------|--------|
| V1 | Project-root CLAUDE.md + unscoped rules are **re-injected from disk after compaction** | CONFIRMED | `context-window.md#what-survives-compaction` |
| V2 | Rules with `paths:` frontmatter are **lost after compaction** until a matching file is read again | CONFIRMED | `context-window.md#what-survives-compaction` |
| V3 | Nested (subdir) CLAUDE.md is **lost after compaction** until a file in that dir is read again | CONFIRMED | `context-window.md`, `memory.md` |
| V4 | Path-scoped rules trigger on **Read** of a matching file (not every tool use) | CONFIRMED | `memory.md#path-specific-rules` |
| V5 | The only documented rule-scoping frontmatter field is **`paths`** (not `globs`) | CONFIRMED | `memory.md#path-specific-rules` |
| V6 | `.claude/rules/` is an **official** feature; all `.md` auto-discovered recursively; unscoped = always-on | CONFIRMED | `memory.md#organize-rules-with-claude/rules/` |
| V7 | Skill **name+description load every turn**; body loads on invoke | CONFIRMED | `skills.md`, `features-overview.md` |
| V8 | Skill listing description cap = **1,536 chars** default of `maxSkillDescriptionChars` (configurable, v2.1.105+) | CONFIRMED | `skills.md`, `settings.md` |
| V9 | Output styles **still exist** (not deprecated); only the standalone `/output-style` command was removed (v2.1.91) → use `/config` | CONFIRMED | `output-styles.md`, changelog |
| V10 | Output styles **modify the system prompt** (add to end); a **custom** style without `keep-coding-instructions: true` **removes built-in software-engineering instructions** | CONFIRMED | `output-styles.md` |
| V11 | `force-for-plugin: true` auto-applies a plugin's style, **overriding the user's `outputStyle`** | CONFIRMED | `output-styles.md` |
| V12 | CLAUDE.md is "**a user message after the system prompt**", whereas output styles are a system-prompt mechanism | CONFIRMED | `output-styles.md` (comparison table) |
| V13 | Output styles are "the most expensive" steering mechanism | **UNVERIFIED** — docs say cost rises but prompt-cache mitigates; no ranking | `output-styles.md` |
| V14 | Sub-agent runs in isolated context; only a summary returns | CONFIRMED | `sub-agents.md` |
| V15 | For **plugin** subagents, `hooks` / `mcpServers` / `permissionMode` frontmatter are **silently ignored** | CONFIRMED | `sub-agents.md` |
| V16 | Agent `skills:` frontmatter **preloads full skill content at startup** | CONFIRMED | `sub-agents.md` |
| V17 | Required agent frontmatter = **only `name` + `description`** | CONFIRMED | `sub-agents.md` |
| V18 | Hooks: **~30 events** (not "five"); hook scripts run outside context, but injected `additionalContext` IS saved to transcript (subject to compaction) | CONFIRMED | `hooks.md` |
| V19 | An official **mechanism-fit comparison table** exists (output style vs CLAUDE.md vs `--append-system-prompt` vs agents vs skills) | CONFIRMED | `output-styles.md` |
| U1 | hook event `post-session` (kebab) | **RESOLVED → REFUTED** (2026-06-20). The 2.1.169 changelog `post-session` is a **self-hosted-runner** workspace-lifecycle hook, NOT a settings.json hook event; absent from `hooks.md` (all 30 events PascalCase). **Removed** from HKV `VALID_EVENTS` in v5.4.1. | `hooks.md`, changelog 2.1.169 |
| U2 | A plugin.json `outputStyles` path-override key (PLH `SHADOWING_PATH_FIELDS`) | **RESOLVED → CONFIRMED** (2026-06-20). `outputStyles` (camelCase) is a documented plugin.json key in the **replaces** category (default `output-styles/` ignored when set). PLH is correct — no change. See `[[plugin-json-path-behavior]]`. | `plugins-reference.md` |
---
## The five features
Design constraints throughout: deterministic where possible; additive findings that do
**not** fire on the SC-5 snapshot fixture stay byte-stable (the v5.3/v5.4 pattern); every
new finding gets a stable title (fix-engine + humanizer couple on title — see `b6a62d7`).
### Foundation (prerequisite for B/A/C/E): enumeration
`active-config-reader.mjs` today enumerates CLAUDE.md cascade, hooks, MCP, skills, plugins —
but **not** rules, agents, or output styles (confirmed by the coverage scan). B, and parts
of A/C/E, need these enumerated with a `loadPattern` tag. This is the first build step.
- Add enumeration for: `.claude/rules/*.md` (+ `~/.claude/rules/`), agents (`.claude/agents/`,
`~/.claude/agents/`, plugin agents), output styles (`.claude/output-styles/`,
`~/.claude/output-styles/`).
- Tag each source kind with `loadPattern ∈ { always, on-demand, external }` and
`survivesCompaction ∈ { yes, no, n/a }` derived from V1V3, V6, V7, V18.
- **Test:** fixture with root+subdir CLAUDE.md, scoped+unscoped rule, project agent, output
style → assert each is enumerated with the correct `loadPattern`/`survivesCompaction`.
### A — Durability / compaction-survival findings
**Problem.** A must-always-hold instruction placed where it does **not** survive compaction
(nested CLAUDE.md per V3, path-scoped rule per V2) silently disappears mid-session after a
`/compact`. The docs publish exactly which mechanisms survive (V1V3) — so this is
doc-grounded, not a guess.
**Shape (additive, no new scanner — count stays 13).**
- **RUL** new finding `CA-RUL-NNN` — a **large** path-scoped rule (e.g. > 50 lines, reuse the
existing unscoped-size heuristic) carries an informational note: its content is **lost after
compaction** until a matching file is re-read (V2/V4). Severity **low** (awareness, not a bug).
- **CML** new finding `CA-CML-NNN` — a **nested** (non-root) CLAUDE.md of meaningful size
notes it is **not re-injected after compaction** (V3). Severity **low**.
- Keep it strictly **structural** (size + location), never semantic ("is this instruction
critical?") — that stays deterministic.
**Key assumption + test.** Discovery distinguishes root vs nested CLAUDE.md and scoped vs
unscoped rules. → Test with a fixture asserting the finding fires for nested/scoped+large and
**not** for root/unscoped.
**Byte-stability.** Additive; snapshot fixture has no nested CLAUDE.md / large scoped rule →
SC-5 byte-stable (re-verify).
### B — Load-pattern accounting in `tokens` + `manifest`
**Problem.** `manifest.mjs` and TOK rank all sources uniformly by `estimated_tokens`. A skill
listing entry (paid **every turn**, V7) ranks identically to a skill body (paid **on invoke**).
The docs give a precise per-mechanism loading + survival model (V1V3, V7, V18) that we can
surface. This is the **core** of the "wrong mechanism = wasted tokens" thesis.
**Shape.**
- Add a `loadPattern` (and `survivesCompaction`) column to `manifest` output and to TOK
source records, derived from the foundation enumeration.
- Add an **always-loaded subtotal** to the manifest summary: "≈X tokens enter context every
turn before you type" (project-root CLAUDE.md + unscoped rules + skill/agent listing +
output style + MCP schemas), vs an on-demand subtotal.
- Optional follow-up: cost-benefit hint (a large `always` source dwarfing the on-demand pool).
**Key assumption + test.** Source-kind → `loadPattern` is a deterministic mapping. → Test:
manifest on a mixed fixture yields the correct always-subtotal; `--json` includes the new field.
**Byte-stability — RISK.** B changes the **manifest/tokens output format** → **not**
byte-stable. Requires snapshot regen (SC-5) and a decision on `--json` back-compat (add field
vs version the schema). This is the highest-format-risk item; sequence it deliberately.
### C — Output-style scanner (new `CA-OST`, count 13 → 14)
**Problem.** Output styles are live (V9) and the **most surprising** surface: a custom style
silently strips built-in software-engineering instructions (V10), and `force-for-plugin`
overrides the user's choice (V11). config-audit does not scan them at all today (GAP/SET/PLH
only touch adoption / a settings key / plugin folder-shadow).
**Shape (new scanner — output styles are a genuinely new file surface, so a new scanner is
warranted, unlike the additive v5.3/v5.4 work).**
- `CA-OST-001`**custom** output style missing `keep-coding-instructions: true` → flags that
built-in SWE instructions are removed (V10). Severity **medium** (silently changes coding
behavior). The headline footgun.
- `CA-OST-002``force-for-plugin: true` in a project/user style → it overrides the user's
`outputStyle` (V11). Severity **low**.
- `CA-OST-003` — settings `outputStyle` value resolving to a non-existent style → dead config.
Severity **low/medium**. (Could live in SET instead; decide during build.)
- Active output-style token cost → feed B's `always` accounting (V10: system-prompt, every turn).
- **Not** asserting V13 ("most expensive") anywhere — unverified.
**Key assumption + test.** Discovery can find `.claude/output-styles/*.md` and
`~/.claude/output-styles/*.md` (new discovery type). → Test: fixture with a custom style
missing the flag fires `CA-OST-001`; a built-in style reference does not.
**Byte-stability.** New scanner; fixture-gated → SC-5 byte-stable if it does not fire on the
snapshot project (re-verify). Self-audit scanner-count badge moves 13 → 14 (update README +
CLAUDE.md inventory + the "count stays 13" lore).
### D — Mechanism-fit detector (heuristic — lowest precision, sequence last)
**Problem.** The docs publish a mechanism-fit table (V19) and the architectural distinction
that CLAUDE.md is a per-turn user message while output styles are system-prompt (V12). Common
mismatch: an "every time / before each" instruction in CLAUDE.md that should be a **hook**; a
path-specific instruction in root CLAUDE.md that should be a **path-scoped rule** (V4/V6).
**Shape (heuristic; additive to CML or a small new check).**
- `CA-???-001` — imperative lifecycle phrasing in CLAUDE.md (`every time`, `before each`,
`always run`, `after you …`, `whenever you …`) describing a tool/lifecycle action → suggest
a hook. Framed as **Missed opportunity**, severity **low**, never "Fix this now".
- `CA-???-002` — path-specific phrasing in **root** CLAUDE.md (`in src/**`, `for *.ts files`)
→ suggest a path-scoped rule.
**Key assumption + test — THE RISK.** Phrase heuristics must have low false-positive rate. →
Test with a positive corpus (clear mismatches → flagged) **and** a negative corpus (normal
project conventions → silent). Must be suppressible (`.config-audit-ignore`). If precision is
poor in testing, **ship behind a flag or defer** — do not ship a noisy heuristic.
**Byte-stability.** Additive; risk it fires on the snapshot's CLAUDE.md → check carefully,
may need snapshot regen.
### E — Agent-listing cost + plugin-agent dead-config (additive to PLH; feeds B)
**Problem.** Agents' name+description load for delegation (every turn), like skills — but
unlike skills they're unmetered. And `hooks`/`mcpServers`/`permissionMode` in a **plugin**
agent's frontmatter are silently ignored (V15) — dead config, and a false sense of
isolation/security if `permissionMode` is among them.
**Shape (additive to PLH).**
- `CA-PLH-NNN` — plugin agent declaring `hooks` / `mcpServers` / `permissionMode` → dead config
(V15). Severity **low** generally; **medium** for `permissionMode` (false security).
- Agent description cost → feed B's `always` accounting. **No hard cap claimed** — research
found no documented agent-description limit (unlike skills' 1,536, V8). Treat as cost signal,
not a limit.
- `skills:` preload in agent frontmatter → informational note: full skill content injected at
startup (V16).
**Key assumption + test.** PLH can tell a **plugin** agent from a user/project agent (V15
only applies to plugin subagents). → Test: plugin-agent fixture with `permissionMode` fires
medium dead-config; identical user-level agent does **not**.
---
## Dependencies & phased rollout
```
Foundation (enumeration: rules, agents, output styles + loadPattern/survivesCompaction)
├── A (durability) additive RUL+CML doc-grounded, low risk
├── E (plugin-agent dead) additive PLH doc-grounded, low risk
├── C (output-style) new CA-OST (13→14) doc-grounded, new surface
└── B (load-pattern) manifest/tokens format doc-grounded, FORMAT risk
D (mechanism-fit) heuristic independent, precision risk
```
**Recommended phasing** (chunk-work rule — session-sized, checkpoint STATE between):
- **v5.5.0 "steering-model I"** — A + E only. **DONE on main** (E `f75ed56`, A `f3aadb5`,
2026-06-20). Additive to RUL/CML/PLH, doc-grounded, byte-stable, count stays 13. **Foundation
was dropped from v5.5.0**: on inspection A/E are additive to scanners that already read the
files and do NOT consume the `active-config-reader` enumeration — that serves B, so it moves to
v5.6. (Release-cut of v5.5.0 is a separate GO step.)
- **v5.6.0 "steering-model II"****Foundation** (active-config-reader enumeration +
`loadPattern`/`survivesCompaction`) + B (load-pattern accounting, manifest format change +
snapshot regen + `--json` back-compat decision) + C (new `CA-OST`, count → 14). Also fix the
frontmatter-parser block-sequence limitation (inline `paths:` only today) as part of Foundation.
- **v5.7.0 (optional)** — D, only if the negative-corpus test shows acceptable precision;
otherwise behind a flag or dropped.
(Operator may prefer a single larger v5.5 — but B's format change and D's precision risk argue
for separating them from the low-risk additive batch.)
---
## Acceptance criteria (whole programme)
- [ ] Every new finding's primary claim traces to a `CONFIRMED` row above (no `UNVERIFIED`
assertions in user-facing text — V13 and U1/U2 must be resolved or omitted).
- [ ] `node --test 'tests/**/*.test.mjs'` green; README test badge == suite count
(`self-audit --check-readme` `passed: true`).
- [ ] `self-audit` stays **A / A**, no critical/high on the plugin itself.
- [ ] Each new finding has: a positive fixture (fires) **and** a negative fixture (silent).
- [ ] fix-engine + humanizer-data entries added/updated for any new finding title (title
coupling — see `b6a62d7`).
- [ ] SC-5 snapshot: byte-stable for additive findings (A/E); regenerated + reviewed for B
(and D if it touches the snapshot project).
- [ ] Scanner-count lore updated everywhere if C lands (README badge, CLAUDE.md inventory,
`docs/scanner-internals.md`, the "count stays 13" notes).
- [x] U1 (`post-session`) and U2 (plugin `outputStyles` key) resolved (2026-06-20): U1 refuted
→ removed from HKV in v5.4.1; U2 confirmed → PLH unchanged.
## Open decisions for the operator
1. **Phasing**: three releases (recommended) vs one v5.5 with everything?
2. **C placement**: new `CA-OST` scanner (count → 14) vs folding output-style checks into
SET+a file check (keeps 13)? (Recommend new scanner — distinct surface.)
3. **D**: build now (heuristic, precision-gated) vs defer until A/B/C/E prove the theme?
4. **B `--json`**: add `loadPattern` field in place (mild back-compat risk) vs schema version bump?

View file

@ -0,0 +1,126 @@
# v5.7 — Optimization Lens + Living Knowledge Base (Plan)
> Outcome of the 2026-06-20 vision discussion. This is the **first concrete realization**
> of the operator's "F1-tuning" north star (see auto-memory `config-audit-vision`). The
> shift is from *"is the config **correct**?"* (today's health scanners) to *"is the config
> **optimal / best-practice-tuned**?"*.
>
> **Status: GO-ready design. Implementation is a SEPARATE GO, chunk by chunk.** The
> discussion session deliberately stopped before code (per STATE.md). This doc mirrors the
> verification-protocol format of `docs/v5.5-steering-model-plan.md`.
## The four building blocks (full vision)
The vision decomposes into four blocks. v5.7 builds **Fase 1** (blocks 1+2 lite). Blocks
3+4 are **Fase 2** (deferred, own GO).
| # | Block | Phase | One-liner |
|---|-------|-------|-----------|
| 1 | **Optimization lens** | **Fase 1** | New finding family: "you USE mechanism X, but Y fits this content better" (mechanism-fit). |
| 2 | **Living knowledge base** | **Fase 1** | Structured, provenance-stamped register the lens reads + semi-auto refresh. |
| 3 | Machine-wide campaign | Fase 2 | A durable ledger above sessions: audit N repos over many sessions, resumable, machine-wide roll-up. |
| 4 | Durable backlog + execution | Fase 2 | One cross-repo prioritized backlog (critical/high/med/low) the user picks from; per-repo plans exported to each repo's `docs/`. |
**Why Fase 1 first (operator-confirmed):** the value of a machine-wide campaign (Fase 2)
depends on the lens being good. Building the campaign shell on today's correctness-only
scanners would underdeliver the vision. So: establish + validate the lens on ONE repo,
then scale.
## Decisions locked (2026-06-20)
- **Sequence:** Fase 1 (lens + knowledge) first, then Fase 2 (campaign + backlog). _(operator)_
- **Knowledge format:** **structured register** (YAML/JSON) the lens reads directly;
markdown kept as a human-readable mirror. _(operator, over markdown-only)_
- **Lens engine:** **new finding family** with a **hybrid motor** — deterministic
pre-filter → opus analyzer that judges mechanism-fit and cites the register rule;
precision-gated. _(operator, over extend-feature-gap / pure-deterministic-scanner)_
## Fase 1 — two deliverables
### Leveranse A — Living knowledge register
Today `knowledge/*.md` (8 files) is prose with a `Source: … verified DATE` header, and the
v5.5 V-rows are an ad-hoc table. The foundation exists; what's missing is a
**machine-consumable** form with provenance per claim.
- **Format:** one register (e.g. `knowledge/best-practices.yaml`), one entry = one
best-practice rule, with fields:
`id / claim / mechanism / recommendation / source-url / verified-date / confidence
{confirmed|inferred|unverified} / lens-check (which detector consumes it)`.
- **Markdown stays** as the readable mirror; the v5.5 V-rows are **migrated into** the
register (formalizing the existing claim→source→CONFIRMED protocol).
- **"Living" =** `/config-audit knowledge-refresh` (semi-auto, **human-approved writes**):
polls CC changelog + Anthropic docs/blog, flags `stale` (older than N days / source
changed) and `candidate` (new practice found), presents for approval. Mirrors the
`architect` plugin's kb-update poll. **No unverified claim is ever auto-written**
(Verifiseringsplikt).
### Leveranse B — Optimization lens (new family, e.g. `CA-OPT`)
A new finding layer that reads CLAUDE.md / rules / skills / hooks and checks each against
the register's mechanism-fit rules. Content is already source-anchored from the Anthropic
"Steering Claude Code" blog (read 2026-06-20):
| Signal in config | Best-practice rule | Source |
|---|---|---|
| Lifecycle phrasing ("after every commit, do X") in CLAUDE.md | → hook (deterministic) | blog |
| Path-specific instruction, unscoped | → path-scoped rule (`paths:`) | blog |
| 30-line procedure in CLAUDE.md | → skill | blog |
| "Never do X" as an instruction | → permission/hook (an instruction is the wrong tool for absolute prohibitions) | blog |
| Custom output-style missing `keep-coding-instructions` | (already covered by CA-OST-001) | blog/docs |
This is **v5.7 D "mechanism-fit" promoted to a real family**, driven by the register
instead of hardcoded rules.
- **Hybrid motor:** cheap deterministic pre-filter (line counts, lifecycle keywords,
path-specificity) → opus analyzer agent (sibling of `feature-gap-agent`) that judges fit
and cites the register rule. **Precision-gated** (emit only on high confidence).
- **Overlap with `feature-gap`:** feature-gap = "you DON'T use feature X"; optimization
lens = "you USE mechanism X, but Y fits THIS content better." Decision: **separate
family** to keep those two intents clean.
- Findings fold into existing posture/report/plan flow → **Fase 2 campaign inherits them
for free**.
## Proposed chunking (one GO'd session each; per `chunk-work-before-compaction`)
1. **Chunk 1 — Register foundation.** Structured register format + schema validation +
migrate existing `knowledge/` + v5.5 V-rows into it + tests. (No output change → byte-stable.)
2. **Chunk 2 — The lens (`CA-OPT`).** Deterministic pre-filter + opus analyzer agent +
humanizer/scoring wiring + fixtures + byte-stability strip. (New family → additive.)
3. **Chunk 3 — `knowledge-refresh`.** Semi-auto poller: `--dry-run`, stale/candidate
flagging, human-approved writes. (The "living" part.)
Dependency order: Chunk 1 → Chunk 2 (lens reads register) → Chunk 3 (keeps register fresh).
## Verification (per plan-quality rule)
- **Register:** schema validates; every entry has source + verified-date;
`knowledge-refresh --dry-run` lists stale/candidate **without writing**.
- **Lens:** a fixture repo with KNOWN mechanism-fit problems (procedure-in-CLAUDE.md,
unscoped path-rule) → lens flags **exactly** those; **zero false positives** on a clean
fixture. Precision target stated explicitly (lens is precision-gated).
- **Byte-stability:** new `CA-OPT` family is additive → follow the **post-B2 "preserve
frozen + strip"** regime (do NOT re-seed); regen only SC-5 default-output. New scanner
family also bumps `scanners_ok` on the deterministic fixture → mirror the `strip-added-scanner`
precedent from CA-OST (v5.6 C).
## Fase 2 — deferred (own GO, after Fase 1 validated)
- **Block 3 — Machine-wide campaign:** a durable campaign ledger above sessions (repo
list + per-repo status pending/audited/planned/implemented + machine-wide roll-up by
severity), resumable across sessions. Start **thin** (ledger + roll-up), not full
orchestration.
- **Block 4 — Durable backlog + execution:** make persistence explicit + versioned/
migratable; one cross-repo prioritized backlog the user picks from; per-repo plans
optionally exported to each repo's own `docs/` ("planer følger arbeidsstedet"); reuse
existing backup/rollback for execution.
- **Persistence note:** sessions already live in `~/.claude/config-audit/sessions/`
(OUTSIDE the plugin dir → survive uninstall/reinstall/upgrade — verified). Fase 2 adds
an explicit version field + migration so upgrades don't break the ledger.
## Synergy / open threads
- **`/repo-init` synergy:** once the lens exists, it becomes the quality meter for what
`/repo-init` produces — evaluate repo-init output with the same lens.
- **Open:** final ID prefix for the family (`CA-OPT` proposed); register file format
(YAML vs JSON); `knowledge-refresh` poll cadence + which sources beyond changelog/blog.

View file

@ -2,8 +2,7 @@
"mcpServers": {
"filesystem": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-filesystem", "."],
"trust": "local"
"args": ["-y", "@modelcontextprotocol/server-filesystem", "."]
}
}
}

View file

@ -6,7 +6,11 @@ import { readdirSync, readFileSync, existsSync } from 'fs';
import { join, basename } from 'path';
import { homedir } from 'os';
const sessionsDir = join(homedir(), '.config-audit', 'sessions');
// Canonical location since v2.2.0. The pre-v2.2.0 path is kept as a fallback so
// sessions created before the move are still detected (see commands/cleanup.md).
const canonicalSessionsDir = join(homedir(), '.claude', 'config-audit', 'sessions');
const legacySessionsDir = join(homedir(), '.config-audit', 'sessions');
const sessionsDir = existsSync(canonicalSessionsDir) ? canonicalSessionsDir : legacySessionsDir;
if (!existsSync(sessionsDir)) {
process.exit(0);

View file

@ -6,7 +6,11 @@ import { readdirSync, readFileSync, statSync, existsSync } from 'fs';
import { join, basename, dirname } from 'path';
import { homedir } from 'os';
const sessionsDir = join(homedir(), '.config-audit', 'sessions');
// Canonical location since v2.2.0. The pre-v2.2.0 path is kept as a fallback so
// sessions created before the move are still detected (see commands/cleanup.md).
const canonicalSessionsDir = join(homedir(), '.claude', 'config-audit', 'sessions');
const legacySessionsDir = join(homedir(), '.config-audit', 'sessions');
const sessionsDir = existsSync(canonicalSessionsDir) ? canonicalSessionsDir : legacySessionsDir;
if (!existsSync(sessionsDir)) {
console.log('{}');

View file

@ -27,11 +27,10 @@
| 21 | Rules file glob doesn't match any project files | CA-RUL-002 | low | Fix the glob pattern. `src/**/*.ts` won't match `./src/file.ts` — test actual paths. |
| 22 | Deprecated frontmatter field in rules file | CA-RUL-003 | low | Remove/replace deprecated fields. Check official docs for current frontmatter schema. |
| 23 | `.claude/rules/` directory missing entirely | CA-RUL-004 | medium | Create directory and split CLAUDE.md by domain. Path-specific rules dramatically reduce context overhead. |
| 24 | MCP server with no trust level set | CA-MCP-001 | medium | Set `"trust": "workspace"` or `"trusted"` explicitly. Default is untrusted/sandboxed; may cause unexpected failures. |
| 25 | User MCP servers in project `.mcp.json` | CA-MCP-002 | low | Move personal MCP servers to `~/.claude.json`. Project `.mcp.json` is for servers the whole team needs. |
| 26 | No custom skills when team has repeated workflows | CA-GAP-001 | medium | Create skills for `/deploy`, `/review-pr`, `/fix-issue`. Repeated multi-step workflows are the target. |
| 27 | Custom agents without `description` field | CA-GAP-002 | medium | Add a description explaining when to delegate to this agent. Without it, Claude never auto-invokes it. |
| 28 | No hooks configured at all | CA-GAP-003 | high | Add at minimum a `Stop` hook for session summaries. Zero hooks is the most common high-value gap. |
| 24 | User MCP servers in project `.mcp.json` | CA-MCP-002 | low | Move personal MCP servers to `~/.claude.json`. Project `.mcp.json` is for servers the whole team needs. |
| 25 | No custom skills when team has repeated workflows | CA-GAP-001 | medium | Create skills for `/deploy`, `/review-pr`, `/fix-issue`. Repeated multi-step workflows are the target. |
| 26 | Custom agents without `description` field | CA-GAP-002 | medium | Add a description explaining when to delegate to this agent. Without it, Claude never auto-invokes it. |
| 27 | No hooks configured at all | CA-GAP-003 | high | Add at minimum a `Stop` hook for session summaries. Zero hooks is the most common high-value gap. |
---

View file

@ -0,0 +1,153 @@
{
"version": 1,
"note": "Machine-readable best-practices register. SOURCE OF TRUTH for the optimization lens (v5.7 CA-OPT). Human-readable mirror lives in knowledge/*.md. Every entry is provenance-stamped (source.url + source.verified) and carries a confidence; only CONFIRMED claims are consumed user-facing (Verifiseringsplikt). Curated manually + by /config-audit knowledge-refresh (human-approved). Seeded from docs/v5.5-steering-model-plan.md V-rows + the Anthropic 'Steering Claude Code' blog.",
"entries": [
{
"id": "BP-MECH-001",
"claim": "Lifecycle automation phrased as an instruction in CLAUDE.md (\"every time\", \"before each\", \"always run X after Y\") should be a hook — a behavior the model chooses to follow is not deterministic.",
"mechanism": "hook",
"appliesTo": "claude-md",
"recommendation": "Move the behavior to a PreToolUse/PostToolUse/Stop hook so it runs deterministically, outside the model's discretion.",
"confidence": "confirmed",
"severity": "low",
"category": "mechanism-fit",
"lensCheck": "claude-md-lifecycle-phrasing",
"source": { "url": "https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more", "title": "Steering Claude Code: skills, hooks, rules, subagents and more", "verified": "2026-06-20" }
},
{
"id": "BP-MECH-002",
"claim": "A file- or path-specific constraint placed in root CLAUDE.md or an unscoped rule should be a path-scoped rule (paths: frontmatter), so it loads only when a matching file is touched.",
"mechanism": "rule",
"appliesTo": "claude-md",
"recommendation": "Move it to .claude/rules/ with a paths: frontmatter; unscoped instructions cost tokens every turn whether relevant or not.",
"confidence": "confirmed",
"severity": "low",
"category": "mechanism-fit",
"lensCheck": "unscoped-path-specific-instruction",
"source": { "url": "https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more", "title": "Steering Claude Code: skills, hooks, rules, subagents and more", "verified": "2026-06-20" }
},
{
"id": "BP-MECH-003",
"claim": "A multi-step procedure (deploy/release checklist) in CLAUDE.md should be a skill — CLAUDE.md is for facts Claude should hold all the time; procedures belong in skills.",
"mechanism": "skill",
"appliesTo": "claude-md",
"recommendation": "Extract the procedure into .claude/skills/; its body then loads only on invoke instead of every turn.",
"confidence": "confirmed",
"severity": "low",
"category": "mechanism-fit",
"lensCheck": "procedure-in-claude-md",
"source": { "url": "https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more", "title": "Steering Claude Code: skills, hooks, rules, subagents and more", "verified": "2026-06-20" }
},
{
"id": "BP-MECH-004",
"claim": "An absolute prohibition phrased as a \"never do X\" instruction is the wrong tool; for something that absolutely must not happen, use permissions or a PreToolUse hook.",
"mechanism": "permission",
"appliesTo": "claude-md",
"recommendation": "Enforce hard prohibitions via permission deny rules or a PreToolUse hook (exit code 2 denies the call), not prose instructions.",
"confidence": "confirmed",
"severity": "low",
"category": "mechanism-fit",
"lensCheck": "never-instruction",
"source": { "url": "https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more", "title": "Steering Claude Code: skills, hooks, rules, subagents and more", "verified": "2026-06-20" }
},
{
"id": "BP-MECH-005",
"claim": "A custom output style without keep-coding-instructions: true removes Claude Code's built-in software-engineering instructions when active.",
"mechanism": "output-style",
"appliesTo": "output-style",
"recommendation": "Set keep-coding-instructions: true, or prefer a built-in style (Explanatory / Learning / Proactive) before writing a custom one.",
"confidence": "confirmed",
"severity": "medium",
"category": "mechanism-fit",
"lensCheck": "CA-OST-001",
"source": { "url": "https://code.claude.com/docs/en/output-styles", "title": "Output styles", "verified": "2026-06-20" }
},
{
"id": "BP-LOAD-001",
"claim": "Project-root CLAUDE.md and unscoped rules are re-injected from disk after compaction (they survive a /compact).",
"appliesTo": "claude-md",
"confidence": "confirmed",
"category": "loading-model",
"lensCheck": null,
"source": { "url": "https://code.claude.com/docs/en/context-window", "title": "Context window — what survives compaction", "verified": "2026-06-20" }
},
{
"id": "BP-LOAD-002",
"claim": "Path-scoped rules are lost after compaction until a matching file is read again, and they trigger on Read of a matching file (not on every tool use).",
"appliesTo": "rule",
"confidence": "confirmed",
"category": "loading-model",
"lensCheck": null,
"source": { "url": "https://code.claude.com/docs/en/memory", "title": "Memory — path-specific rules", "verified": "2026-06-20" }
},
{
"id": "BP-LOAD-003",
"claim": "A nested (non-root) CLAUDE.md is lost after compaction until a file in its directory is read again.",
"appliesTo": "claude-md",
"confidence": "confirmed",
"category": "loading-model",
"lensCheck": null,
"source": { "url": "https://code.claude.com/docs/en/context-window", "title": "Context window — what survives compaction", "verified": "2026-06-20" }
},
{
"id": "BP-LOAD-004",
"claim": "A skill's name + description load every turn; its body loads only on invoke.",
"appliesTo": "skill",
"confidence": "confirmed",
"category": "loading-model",
"lensCheck": null,
"source": { "url": "https://code.claude.com/docs/en/skills", "title": "Skills", "verified": "2026-06-20" }
},
{
"id": "BP-LOAD-005",
"claim": "Hook scripts run outside the model context, but any additionalContext they inject is saved to the transcript and is therefore subject to compaction.",
"appliesTo": "hook",
"confidence": "confirmed",
"category": "loading-model",
"lensCheck": null,
"source": { "url": "https://code.claude.com/docs/en/hooks", "title": "Hooks", "verified": "2026-06-20" }
},
{
"id": "BP-LOAD-006",
"claim": "A subagent runs in an isolated, fresh context window; only its final summary returns to the main session (parent instructions are not auto-injected).",
"appliesTo": "agent",
"confidence": "confirmed",
"category": "loading-model",
"lensCheck": null,
"source": { "url": "https://code.claude.com/docs/en/sub-agents", "title": "Subagents", "verified": "2026-06-20" }
},
{
"id": "BP-SIZE-001",
"claim": "Keep CLAUDE.md under 200 lines; give it an owner and review changes to it like code. Every line costs tokens whether relevant or not.",
"appliesTo": "claude-md",
"recommendation": "Trim CLAUDE.md to facts; move procedures to skills and path-specific rules to .claude/rules/.",
"confidence": "confirmed",
"severity": "medium",
"category": "size-budget",
"lensCheck": "CA-CML-001",
"source": { "url": "https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more", "title": "Steering Claude Code: skills, hooks, rules, subagents and more", "verified": "2026-06-20" }
},
{
"id": "BP-SIZE-002",
"claim": "The skill-listing description cap is 1,536 characters (maxSkillDescriptionChars, configurable, v2.1.105+); the name + description load every turn.",
"appliesTo": "skill",
"confidence": "confirmed",
"severity": "low",
"category": "size-budget",
"lensCheck": "CA-SKL-002",
"source": { "url": "https://code.claude.com/docs/en/skills", "title": "Skills", "verified": "2026-06-20" }
},
{
"id": "BP-SUB-001",
"claim": "Every line of CLAUDE.md loads into every session whether or not it is relevant, which consumes tokens and dilutes adherence. A line that states a local fact the model cannot derive (build commands, directory layout, conventions, team norms) earns that cost; a line that only restates general engineering behaviour pays it without being the kind of content CLAUDE.md is for, and is a candidate for removal.",
"mechanism": "deletion",
"appliesTo": "claude-md",
"recommendation": "Review the block for removal, then re-add it only if the model actually stumbles on it repeatedly. Local facts (remotes, versions, paths, conventions) and policy invariants are the floor and are never removal candidates.",
"confidence": "confirmed",
"severity": "low",
"category": "subtraction",
"lensCheck": "compensatory-instruction",
"source": { "url": "https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more", "title": "Steering Claude Code: skills, hooks, rules, subagents and more", "verified": "2026-07-31" }
}
]
}

View file

@ -109,6 +109,6 @@ when structural signal alone isn't enough.
## See also
- `knowledge/opus-4.7-patterns.md` — structural patterns the TOK scanner detects (CA-TOK-001..005)
- `knowledge/prompt-cache-patterns.md` — structural patterns the TOK scanner detects (CA-TOK-001..005)
- `knowledge/configuration-best-practices.md` — CLAUDE.md cache-stability guidance
- `/config-audit tokens --with-telemetry-recipe` — surfaces a pointer to this file in JSON output

View file

@ -2,6 +2,7 @@
> Source: Official Claude Code documentation (code.claude.com/docs), 75 pages, verified 2026-04-03.
> Delta layer: research/03-claude-code-changes-config-surfaces.md (verified 2026-04-19) — sandbox/managed-only/prompt-cache surfaces added between v2.1.14 and v2.1.114.
> 2026-06 model/effort lineup (below) re-verified against the official changelog on 2026-06-18 (window v2.1.114v2.1.181).
## 2026-04 deltas (research/03)
@ -18,6 +19,33 @@
| Env: `ENABLE_PROMPT_CACHING_1H`, `FORCE_PROMPT_CACHING_5M` | v2.1.108 | Explicit prompt-cache TTL control. |
| Env: `CLAUDE_CODE_DISABLE_1M_CONTEXT`, `CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING` | 2026-04 | Behavior opt-outs for new defaults. |
## 2026-06 model & effort lineup (changelog, verified 2026-06-18)
> Window v2.1.114v2.1.181. The Opus-4.7 framing elsewhere in the corpus predates these.
| Item | Added in | Notes |
|------|---------|-------|
| **Opus 4.8 — new default** | v2.1.154 | Defaults to high effort. `/effort xhigh` (`effortLevel: "xhigh"`) is the new top tier for the hardest tasks. |
| Lean system prompt now default | v2.1.154 | Default for all models except Haiku, Sonnet, and Opus 4.7 and earlier — lowers baseline system-prompt token cost. |
| **Claude Fable 5** (Mythos-class) | v2.1.170 | New top-capability model family; 1M context by default (a `[1m]` id suffix is normalised away). Auto mode falls back to the best available Opus for orgs without Opus 4.8. |
| `/code-review` bundled review skill | v2.1.147 | Renamed from the former cleanup skill; reports correctness bugs at a chosen effort (`/code-review high`), `--comment` posts inline PR comments. |
| `/config key=value` | v2.1.181 | Apply any setting from the prompt (e.g. `/config thinking=false`) — interactive, `-p`, and Remote Control. Lets `/config-audit` settings recommendations be applied without hand-editing settings.json. |
## v5.3.0 scanner-backing facts (verified 2026-06-19)
| Fact | Source | Backs |
|------|--------|-------|
| **"Large CLAUDE.md will impact performance (X chars > 40.0k)"** — Claude Code emits this startup warning when a CLAUDE.md exceeds ~40,000 chars. CC 2.1.169 scales the threshold with the model's context window: ~40.0k at a 200k-context model, relaxing to ~200,000 chars at 1M context. The figure is a **char** count, not a line count — complementary to the 200/500-line adherence guidance (a file can be long by lines yet under budget, or short by lines yet over it). | CC changelog 2.1.169 + live startup-warning text | CML `CA-CML` char-budget finding — anchors on the conservative 200k window (fires earliest, since the user's window is unobservable) and discloses the relaxed 1M figure. Window constants in `scanners/lib/context-window.mjs`. Severity medium (token cost, not an adherence cliff). |
| **`Tool(param:value)` permission matching is deny/ask-only.** The `param:value` form is honored only in `permissions.deny` and `permissions.ask`; in `permissions.allow` it matches nothing (dead config). Separately, a rule whose key is the tool's own *canonicalizing field*`command` (Bash/PowerShell), `file_path` (Read/Edit/Write), `path` (Grep/Glob), `notebook_path` (NotebookEdit), `url` (WebFetch) — is **ignored** by CC, which emits a startup warning. Valid parameter forms (`Bash(npm:*)`, `WebFetch(domain:host)`, `Agent(model:opus)`) are honored. | code.claude.com/docs/en/permissions | DIS forbidden-param finding (`forbiddenParamRule` in `permission-rules.mjs`): deny/ask = false security (medium, the block never applies), allow = dead config (low). |
## v5.4.0 scanner-backing facts (verified 2026-06-19)
| Fact | Source | Backs |
|------|--------|-------|
| **plugin.json component-path keys in the *replaces* set shadow their default folder.** `commands`, `agents`, and `outputStyles` each replace a default folder (`commands/`, `agents/`, `output-styles/`); setting one to a custom path while the same-named default folder still exists makes the folder silently ignored — Claude Code warns in `/doctor`, `claude plugin list`, and the `/plugin` detail view. **`skills` is different — it *adds to* the default `skills/` scan (both load, never a shadow).** `hooks`/`mcpServers`/`lspServers` have their own merge rules (not folder shadows). A custom path that resolves *into* the default folder (e.g. `"commands": ["./commands/x.md"]`) keeps the folder scanned (explicit-address exception). | code.claude.com/docs path-behavior-rules; CC `/doctor` + `claude plugin list` + `/plugin` detail (v2.1.140+) | PLH `CA-PLH-015` plugin-folder shadowing — `SHADOWING_PATH_FIELDS` = `commands`/`agents`/`outputStyles` only; `addressesDefaultDir` exception. Severity medium, `category: 'plugin-hygiene'`. |
| **`claude plugin validate` requires each `skills:` entry to be a directory inside the plugin.** A `plugin.json` `skills` field (string or array) must point at an existing directory within the plugin root; installed plugins **cannot reference files outside their directory** (no `../shared-utils` path traversal). | `claude plugin validate` (~2.1.145); code.claude.com/docs path-traversal rule | PLH `CA-PLH-016` `skills:`-array validation — normalizes string→`[string]`, flags `non-string` / `escapes-root` / `not-found` / `not-a-directory`; escape detection is plugin-root containment. Severity medium. |
| **`autoMode` is a four-sub-key object and is not read from shared project settings.** Its only valid sub-keys are `environment`, `allow`, `soft_deny`, `hard_deny`, each a **string array** (entries are prose rules; the literal `"$defaults"` is valid). The classifier **does not read `autoMode` from shared project settings** (`.claude/settings.json`), so a checked-in repo cannot inject its own allow rules; valid scopes are user (`~/.claude/settings.json`), local (`.claude/settings.local.json`), and managed. | code.claude.com/docs/en/auto-mode-config | SET `CA-SET` autoMode — structure (object + four string-array sub-keys; `not-an-object`/`unknown-subkey`/`not-string-array`) = medium; dead-config (`shared-project-scope`, keyed on `file.scope === 'project'`) = low. |
## Official Configuration Guidance (Anthropic)
These principles are backed by official docs and verified community reports. Use them to ground recommendations.
@ -215,7 +243,7 @@ paths: ["src/**/*.ts"]
"type": "stdio|http",
"command": "...", "args": [...],
"url": "...",
"env": {}, "timeout": 30000, "trust": "workspace|trusted|untrusted"
"env": {}, "timeout": 30000
}
}
}
@ -223,7 +251,7 @@ paths: ["src/**/*.ts"]
**Fully utilizing:** Team-shared MCP servers (GitHub, Jira, DBs); MCP resources via `@server:path`; MCP prompts as slash commands; `enableAllProjectMcpServers: true` for zero-friction team onboarding.
**Common gaps:** No `.mcp.json`; MCP only configured in `~/.claude.json` (not shared); trust levels not set; MCP resources not used.
**Common gaps:** No `.mcp.json`; MCP only configured in `~/.claude.json` (not shared); project servers not approved via `enableAllProjectMcpServers`/`enabledMcpjsonServers`; MCP resources not used.
---
@ -291,7 +319,7 @@ paths: ["src/**/*.ts"]
**Dynamic context:** `` !`command` `` executes shell command and inlines output.
**Bundled skills:** `/batch`, `/claude-api`, `/debug`, `/loop`, `/simplify`
**Bundled skills:** `/batch`, `/claude-api`, `/debug`, `/loop`, `/code-review` (renamed from the former cleanup skill, v2.1.147)
**Fully utilizing:** Custom deploy/review workflows; `disable-model-invocation: true` on side-effect skills; `context: fork` for isolated research; `!`git diff HEAD`` for dynamic context; `argument-hint` for UX.

View file

@ -6,7 +6,7 @@
## CLAUDE.md
1. **Optimise for prompt-cache stability.** Place stable content in the first 30 lines (cache-friendly prefix); volatile content (timestamps, dynamic counts, rolling activity logs) goes below that threshold or moves to an `@import`-ed file outside the cache prefix. On Opus 4.7 the dominant cost lever is cache reuse, not file length.[^200lines]
1. **Optimise for prompt-cache stability.** Place stable content in the first 30 lines (cache-friendly prefix); volatile content (timestamps, dynamic counts, rolling activity logs) goes below that threshold or moves to an `@import`-ed file outside the cache prefix. On Opus 4.8 the dominant cost lever is cache reuse, not file length.[^200lines]
2. **Use `@import` for specs/docs.** `@path/to/spec.md` inlines the file at session start. Max 5 hops, but keep chains ≤ 2 hops — every `@import` boundary fragments the prompt-cache prefix. Keeps the main file scannable.
3. **Use HTML comments for maintainer notes.** `<!-- Updated 2026-01-01: reason -->` is stripped before context injection — zero token cost.
4. **Put personal dev notes in `CLAUDE.local.md`**, not `CLAUDE.md`. Add `CLAUDE.local.md` to `.gitignore`. Team members' sandbox URLs should never appear in git.
@ -56,7 +56,7 @@
1. **Commit `.mcp.json` to git.** Team-shared MCP servers belong in `.mcp.json` at project root, not in individual `~/.claude.json` files. One commit, everyone gets the servers.
2. **Set `enableAllProjectMcpServers: true` in project settings.json** for zero-friction team onboarding. New team members don't have to manually approve each server.
3. **Set trust levels explicitly.** `"trust": "workspace"` for project-specific servers; `"trust": "trusted"` only for servers you fully control. Default is untrusted (sandboxed).
3. **Approve servers via settings, not a `trust` field.** `.mcp.json` has no per-server `trust` key — approval is recorded by the workspace-trust and MCP-approval dialogs. For non-interactive, granular control set `enabledMcpjsonServers: ["memory", "github"]` / `disabledMcpjsonServers: [...]` in settings.json instead of approving each server by hand.
4. **Use `@server:resource/path` for dynamic data.** `@github:repos/owner/repo/issues` pulls live data into context. More reliable than asking Claude to fetch and parse.
5. **Deny MCP tools you don't want Claude to invoke.** `{"permissions": {"deny": ["mcp__filesystem__write_file"]}}` — even with a server connected, specific tools can be blocked.
@ -94,4 +94,4 @@
---
[^200lines]: The "keep CLAUDE.md under 200 lines" threshold was a Sonnet-era adherence heuristic — Sonnet's attention quality dropped on longer files, so trimming raw line count was the optimisation lever. Opus 4.7 uses prompt-cache structure as the dominant cost driver: the first 30 lines must stay byte-stable across turns to keep the cache hit, and `@import` boundaries fragment the cached prefix. A 400-line CLAUDE.md with stable structure outperforms a 150-line file whose top contains a daily-rolling activity log. See `knowledge/opus-4.7-patterns.md` for detection IDs (CA-TOK-001..003).
[^200lines]: The "keep CLAUDE.md under 200 lines" threshold was a Sonnet-era adherence heuristic — Sonnet's attention quality dropped on longer files, so trimming raw line count was the optimisation lever. Opus 4.8 uses prompt-cache structure as the dominant cost driver: the first 30 lines must stay byte-stable across turns to keep the cache hit, and `@import` boundaries fragment the cached prefix. A 400-line CLAUDE.md with stable structure outperforms a 150-line file whose top contains a daily-rolling activity log. See `knowledge/prompt-cache-patterns.md` for detection IDs (CA-TOK-001..003).

View file

@ -2,6 +2,7 @@
> Timeline of major features, most recent first. Covers features with configuration impact.
> Source: Official Claude Code documentation, verified 2026-04-03; 2026-04 entries verified via research/03-claude-code-changes-config-surfaces.md (2026-04-19).
> 2026-06 entries (window v2.1.114v2.1.181) re-verified against the official Claude Code changelog on 2026-06-18.
---
@ -9,6 +10,12 @@
| Approx. Date | Feature | Config Impact |
|-------------|---------|---------------|
| 2026-06 (v2.1.181) | **`/config key=value` in-session settings** | Set any setting from the prompt — `/config thinking=false` — in interactive, `-p`, and Remote Control. Lets `/config-audit` settings recommendations be applied without hand-editing settings.json. |
| 2026-05 (v2.1.170) | **Claude Fable 5 (Mythos-class)** | New top-capability model family; 1M context by default (a `[1m]` id suffix is normalised away). Auto mode falls back to the best available Opus for orgs without Opus 4.8 enabled. |
| 2026-05 (v2.1.169) | **`post-session` lifecycle hook** | Runs after the session ends and before the workspace is deleted — snapshot uncommitted work or export logs (self-hosted runner). Kebab-case, distinct from `SessionEnd`. |
| 2026-04 (v2.1.154) | **Opus 4.8 — new default model** | Defaults to high effort; `/effort xhigh` (`effortLevel: "xhigh"`) is the new top tier for the hardest tasks. The lean system prompt is now the default for all models except Haiku, Sonnet, and Opus 4.7 and earlier. The settings `agent` field is honored for dispatched `claude agents` sessions (v2.1.157). |
| 2026-04 (v2.1.152) | **`MessageDisplay` hook event** | Hooks can transform or hide assistant message text as it is displayed. |
| 2026-04 (v2.1.147) | **`/simplify` renamed to `/code-review`** | Reports correctness bugs at a chosen effort (`/code-review high`); `--comment` posts inline PR comments. The old cleanup-and-fix behaviour was removed. |
| 2026-04 (v2.1.111) | **Opus 4.7 + token-efficiency surfaces** | New env vars `ENABLE_PROMPT_CACHING_1H`, `FORCE_PROMPT_CACHING_5M`, `CLAUDE_CODE_DISABLE_1M_CONTEXT`. New settings keys around `tui` / `autoScrollEnabled`. Granular commit attribution via `attribution.commit` / `attribution.pr` (replaces `includeCoAuthoredBy`). |
| 2026-04 (v2.1.83+) | **Sandbox + managed-only enterprise lockdown** | Added settings keys: `sandbox.enabled`, `sandbox.failIfUnavailable`, `sandbox.allowUnsandboxedCommands`, `sandbox.filesystem.allowRead/denyRead`, `sandbox.network.deniedDomains/allowedDomains`, `sandbox.enableWeakerNetworkIsolation`. Managed-only flags: `allowManagedHooksOnly`, `allowManagedMcpServersOnly`, `allowManagedPermissionRulesOnly`. |
| 2026-03 (v2.1.91) | **`disableSkillShellExecution`** | Blocks inline `!command` shell expansion in skills. Mitigates skill-side prompt-injection vector. |

View file

@ -73,8 +73,7 @@ Create `.mcp.json` at project root:
"memory": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-memory"],
"trust": "workspace"
"args": ["-y", "@modelcontextprotocol/server-memory"]
}
}
}

View file

@ -1,7 +1,8 @@
# Hook Events Reference
> All 26 hook events as of April 2026. Source: code.claude.com/docs/en/hooks.md
> All 28 hook events as of June 2026. Source: code.claude.com/docs/en/hooks.md
> Verified 2026-04-19 against research/03-claude-code-changes-config-surfaces.md — no new hook events introduced in v2.1.83v2.1.111. Sandbox + managed-only flags (2026-04) operate at the settings layer, not as new hook events.
> Re-verified 2026-06-18 against the official changelog (window v2.1.114v2.1.181): added `MessageDisplay` (v2.1.152) and `post-session` (v2.1.169), bringing the count from 26 to 28.
---
@ -25,6 +26,7 @@
| `StopFailure` | Turn ends in an API error | No | Error type | Alert on API errors; retry logic; log error context |
| `TeammateIdle` | An agent team member has no tasks | Yes | No matcher | Assign next task (exit 2 to keep working); log team status; rebalance work |
| `Notification` | A notification is sent (permission prompt, idle, auth) | No | `permission_prompt`, `idle_prompt`, `auth_success`, `elicitation_dialog` | Desktop notifications; Slack/webhook alerts; mobile push; audio cues |
| `MessageDisplay` | Assistant message text is about to be displayed (v2.1.152) | Yes (transforms) | No matcher | Transform or hide assistant output before it reaches the terminal — redact secrets, rewrite text |
| `ConfigChange` | A config file changes on disk | Yes | `user_settings`, `project_settings`, `local_settings`, `policy_settings`, `skills` | Validate config changes; block invalid edits; reload dependent processes |
| `CwdChanged` | Working directory changes | No | No matcher | Inject new directory context; update env vars via `$CLAUDE_ENV_FILE`; log navigation |
| `FileChanged` | A watched file changes | No | Filename pattern | Auto-reload when config changes; trigger builds on source change; sync state |
@ -35,6 +37,7 @@
| `Elicitation` | An MCP server requests user input | Yes | MCP server name | Control which servers can request input; log elicitations; pre-fill responses |
| `ElicitationResult` | User responds to MCP elicitation | Yes | MCP server name | Validate responses; log user input; transform before sending to MCP |
| `SessionEnd` | Session terminates | No | `clear`, `resume`, `logout`, `prompt_input_exit`, `other` | Final session summary; save state; cleanup temp files; send end-of-session report |
| `post-session` | Session has ended, before the workspace is deleted (v2.1.169) | No | No matcher | Snapshot uncommitted work; export logs; cleanup (self-hosted runner). Kebab-case, distinct from `SessionEnd`. |
---
@ -94,6 +97,9 @@
- `additionalContext`: string injected into context
- Or: write env vars to `$CLAUDE_ENV_FILE`
**Stop / SubagentStop** (exit 0, since v2.1.163):
- `hookSpecificOutput.additionalContext`: feedback string that keeps the turn going without being treated as a hook error
---
## Environment Variables Available in Hooks

View file

@ -1,14 +1,16 @@
# Opus 4.7 Configuration Patterns
# Prompt-Cache Configuration Patterns
> Token-efficiency patterns for Claude Opus 4.7. Detection IDs map to TOK scanner findings.
> Sources: research/01-opus-47-features-token-efficiency.md (Topic 1), research/04-prompt-caching-patterns.md (Topic 4). Last verified 2026-04-19.
> Token-efficiency patterns for current Claude Code (defaults to Opus 4.8; Fable 5 is the top-capability model). Detection IDs map to TOK scanner findings.
> Sources: research/01-opus-47-features-token-efficiency.md (Topic 1), research/04-prompt-caching-patterns.md (Topic 4). Patterns verified 2026-04-19; model-era anchor refreshed 2026-06-18.
Opus 4.7 raises the cost ceiling per turn while expanding the context window
and prompt-cache window. Net effect: cache reuse and tool-schema discipline
become the dominant levers for keeping a session affordable. The patterns
below are structural — they can be detected statically by reading config files
without running a session. Cache hit-rate measurement requires runtime
telemetry and is explicitly out of scope.
Current Claude Code defaults to Opus 4.8 (Fable 5 is the top-capability model).
On these models the cost ceiling per turn is high and both the context window
and prompt-cache window are large, so cache reuse and tool-schema discipline
are the dominant levers for keeping a session affordable. These patterns are
properties of prompt-caching itself, not of any single model — they are
structural and detectable statically by reading config files without running a
session. Cache hit-rate measurement requires runtime telemetry and is
explicitly out of scope.
| # | Pattern | Detection (ID) | Severity | Fix |
|---|---------|----------------|----------|-----|
@ -54,3 +56,14 @@ uncertainty band.
| medium | Materially inflates token cost per turn (cache miss, schema bloat) |
| low | Detectable inefficiency that compounds across long sessions |
| info | Informational signal — no action required, may indicate room for optimisation |
## Skill-listing budget lever: disableBundledSkills
When the active skill listing exceeds its token budget (SKL `CA-SKL-002`), the
`disableBundledSkills` setting is a direct token-efficiency lever: it removes the descriptions
of plugin-bundled skills from the skill listing injected into the system prompt, shrinking the
per-turn baseline. Unlike trimming individual descriptions (`CA-SKL-001`), it is a single switch
that drops the entire bundled-skill surface at once — appropriate when the bundled skills are not
in active use. feature-gap surfaces it as a conditional recommendation, the remediation companion
to `CA-SKL-002`. Sibling levers: `skillOverrides` (selectively re-enable specific skills) and
per-skill description trimming.

View file

@ -0,0 +1,116 @@
/**
* AGT Scanner Agent-listing always-loaded token budget
*
* Claude Code injects a listing of every active agent's name+description into the
* system prompt so it knows which subagents it can delegate to. That listing is
* re-sent on EVERY turn, whether or not a delegation happens so with many
* installed agents it is a large always-loaded cost (on a heavily-plugged machine
* the dominant single always-loaded source).
*
* Detection:
* CA-AGT-NNN per-agent description over the soft bloat cap (low, advisory)
* CA-AGT-NNN aggregate agent-listing estimate exceeds the listing budget (low)
*
* Per-agent advisories are emitted FIRST (mirroring SKL 001002). They are a
* soft bloat heuristic (the 500-char threshold TOK pattern F uses for SKILL.md
* descriptions), NOT a truncation finding agents have no verified per-
* description cap, so nothing is dropped; the description is simply re-sent in
* full every turn.
*
* INTELLECTUAL-HONESTY CONTRACT (the reason this is `low`, not a hard finding):
* unlike the skill listing, the agent-listing mechanism is NOT documented (agents
* are absent from Claude Code's published context breakdown). So the finding is an
* INFERRED, UPPER-BOUND ESTIMATE, the budget is a config-audit heuristic (no
* documented agent allotment) anchored on a conservative 200k window, and the
* evidence discloses all of that. The cap, budget, and enumerate-and-measure step
* live in `lib/agent-listing-budget.mjs` AGT only constructs the finding.
*
* Zero external dependencies.
*/
import { finding, scannerResult } from './lib/output.mjs';
import { SEVERITY } from './lib/severity.mjs';
import {
AGGREGATE_BUDGET_TOKENS,
BUDGET_CALIBRATION_NOTE,
PER_AGENT_DESC_SOFT_CAP,
measureActiveAgentListing,
} from './lib/agent-listing-budget.mjs';
const SCANNER = 'AGT';
/**
* Main scanner entry point.
*
* @param {string} _targetPath unused (agent listing is HOME-scoped)
* @param {object} _discovery unused (ignores project discovery)
*/
export async function scan(_targetPath, _discovery) {
const start = Date.now();
const findings = [];
const { agents, aggregate } = await measureActiveAgentListing();
// Per-agent advisory (emitted FIRST so the common "long agent + aggregate"
// case reads 001=per-agent, 002=aggregate, mirroring SKL). This is a soft
// heuristic, NOT a truncation finding — agents have no verified per-description
// cap, so the framing is "this is large and re-sent every turn", never "Claude
// Code drops the tail".
for (const agent of agents) {
if (agent.descLength <= PER_AGENT_DESC_SOFT_CAP) continue;
const sourceLabel = agent.source === 'plugin'
? `plugin:${agent.pluginName}`
: 'user';
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: 'Agent description is long (re-sent every turn in the always-loaded listing)',
description:
`Agent "${agent.name}" (${sourceLabel}) has a description of ${agent.descLength} ` +
`characters (>${PER_AGENT_DESC_SOFT_CAP}). Claude Code injects every active agent's ` +
'name+description into the agent listing on every turn so it knows which subagents it ' +
'can delegate to, so every character of this description re-enters context each turn ' +
'whether or not you delegate. Unlike the skill listing there is no verified ' +
'per-description cap, so nothing is dropped — this is a bloat advisory, not a ' +
'hard-cap finding.',
file: agent.path,
evidence:
`description_chars=${agent.descLength}; soft_cap=${PER_AGENT_DESC_SOFT_CAP} ` +
`(heuristic, same bloat threshold TOK pattern F uses for SKILL.md descriptions; agents ` +
`have NO verified per-description cap); agent="${agent.name}"; source=${sourceLabel}`,
recommendation:
'Trim the description toward its trigger phrases / "when to use this agent" cues and move ' +
'long examples into the agent body, or disable the plugin if you never delegate to this ' +
'agent (the whole agent block then leaves the always-loaded listing).',
category: 'token-efficiency',
}));
}
if (aggregate.overBudget) {
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: 'Aggregate agent listing may exceed the always-loaded budget',
description:
`The ${aggregate.scanned} active agents carry about ${aggregate.aggregateTokens} tokens of ` +
'name+description text that Claude Code injects every turn so it knows which subagents it can ' +
`delegate to — above the ${AGGREGATE_BUDGET_TOKENS}-token budget this scanner anchors on a 200k ` +
'context window. Every one of those tokens is re-sent on every turn whether or not you delegate. ' +
'Note: unlike the skill listing, the agent-listing always-loaded mechanism is inferred (not ' +
'documented), so this is an upper-bound estimate (see evidence).',
evidence:
`active_agents_scanned=${aggregate.scanned}; description_chars=${aggregate.aggregateChars}; ` +
`description_tokens~${aggregate.aggregateTokens}; budget@200k=${AGGREGATE_BUDGET_TOKENS} tok; ` +
`over_by~${aggregate.overBy} tok - ${BUDGET_CALIBRATION_NOTE}`,
recommendation:
'Shrink the always-loaded agent listing: disable plugins whose agents you do not use (the whole ' +
'agent block leaves the listing), remove dead user agents from ~/.claude/agents/, and trim long ' +
'agent descriptions toward their trigger phrases.',
category: 'token-efficiency',
}));
}
return scannerResult(SCANNER, 'ok', findings, aggregate.scanned, Date.now() - start);
}

View file

@ -5,7 +5,7 @@
* cached prefix ( CACHED_PREFIX_LINES). Distinguishes from TOK Pattern A,
* which only inspects the top 30 lines: CPS catches a `!git log` at line 60
* or a `${TIMESTAMP}` at line 100. Volatile content anywhere in the cached
* prefix breaks Opus 4.7 prompt-cache reuse from that line forward.
* prefix breaks prompt-cache reuse from that line forward.
*
* Volatile patterns extend the TOK set with shell-exec `!` prefix and
* `${VAR}` substitutions both common cache-busters in real CLAUDE.md files.
@ -15,9 +15,12 @@
* Zero external dependencies.
*/
import { resolve, dirname } from 'node:path';
import { tmpdir } from 'node:os';
import { readTextFile } from './lib/file-discovery.mjs';
import { finding, scannerResult } from './lib/output.mjs';
import { SEVERITY } from './lib/severity.mjs';
import { findImports } from './lib/yaml-parser.mjs';
const SCANNER = 'CPS';
@ -27,7 +30,28 @@ const SCANNER = 'CPS';
// hits per turn, not to chase every inline date in a long backlog file.
const CACHED_PREFIX_LINES = 150;
// Volatile-pattern set (extends token-hotspots.mjs Pattern A).
// CC-provided substitution variables that resolve to a stable per-install or
// per-project path (e.g. "${CLAUDE_PLUGIN_ROOT}/hooks/x.mjs"). CC expands them
// to the same value every turn, so they never break the prompt cache — unlike a
// runtime ${TIMESTAMP}. Excluded from the ${VAR} volatile flag (M-BUG-7).
const STABLE_CC_VARS = new Set(['CLAUDE_PLUGIN_ROOT', 'CLAUDE_PROJECT_DIR']);
// Matches every ${VAR} occurrence on a line so a line carrying only stable CC
// vars is not mistaken for a runtime cache-buster.
const VAR_RX = /\$\{([A-Z_][A-Z0-9_]*)\}/g;
/** True when a line contains at least one non-CC-stable ${VAR} substitution. */
function hasVolatileVar(line) {
VAR_RX.lastIndex = 0;
let m;
while ((m = VAR_RX.exec(line)) !== null) {
if (!STABLE_CC_VARS.has(m[1])) return true;
}
return false;
}
// Volatile-pattern set (extends token-hotspots.mjs Pattern A). The ${VAR} entry
// is `varAware` — flagged via hasVolatileVar() so CC-stable vars are excluded.
const VOLATILE_PATTERNS = [
{ rx: /\{timestamp\}/i, label: '{timestamp} placeholder' },
{ rx: /\{uuid\}/i, label: '{uuid} placeholder' },
@ -38,9 +62,25 @@ const VOLATILE_PATTERNS = [
{ rx: /^\s*\[\d{4}-\d{2}-\d{2}/, label: 'dated log line [YYYY-MM-DD ...]' },
// v5 N3 extensions:
{ rx: /^\s*!/, label: 'shell-exec line (! prefix)' },
{ rx: /\$\{[A-Z_][A-Z0-9_]*\}/, label: '${VAR} substitution' },
{ rx: /\$\{[A-Z_][A-Z0-9_]*\}/, label: '${VAR} substitution', varAware: true },
];
/**
* Resolve an @import path relative to the file that declares it.
* Mirrors import-resolver.mjs / token-hotspots.mjs path semantics.
* @param {string} importPath
* @param {string} containingFile
* @returns {string} absolute resolved path
*/
function resolveImportPath(importPath, containingFile) {
if (importPath.startsWith('~')) {
const home = process.env.HOME || process.env.USERPROFILE || tmpdir();
return resolve(importPath.replace(/^~/, home));
}
if (importPath.startsWith('/')) return importPath;
return resolve(dirname(containingFile), importPath);
}
/**
* Scan content for volatile lines within the cached prefix window.
* Returns array of {line, label, snippet}.
@ -49,16 +89,33 @@ function findVolatileLines(content) {
const out = [];
if (!content) return out;
const lines = content.split('\n').slice(0, CACHED_PREFIX_LINES);
let inFence = false;
for (let i = 0; i < lines.length; i++) {
for (const { rx, label } of VOLATILE_PATTERNS) {
if (rx.test(lines[i])) {
out.push({
line: i + 1,
label,
snippet: lines[i].length > 120 ? lines[i].slice(0, 117) + '...' : lines[i],
});
break;
}
const line = lines[i];
// Fenced code blocks (``` or ~~~) hold illustrative, byte-stable literal
// text — a ${VAR} or timestamp shown inside one is documentation, not a
// runtime cache-buster — so the fence delimiters and their content are
// skipped (M-BUG-7).
if (/^\s*(```|~~~)/.test(line)) {
inFence = !inFence;
continue;
}
if (inFence) continue;
// Strip `inline code` spans before pattern-testing: a {date} or ${VAR}
// shown inside backticks is literal documentation text, byte-stable, not a
// runtime cache-buster (M-BUG-7). The original line is still reported as the
// snippet so context is preserved.
const probe = line.replace(/`[^`]*`/g, '');
for (const { rx, label, varAware } of VOLATILE_PATTERNS) {
// The ${VAR} pattern flags only non-CC-stable substitutions; every other
// pattern keeps its plain line test.
if (varAware ? !hasVolatileVar(probe) : !rx.test(probe)) continue;
out.push({
line: i + 1,
label,
snippet: line.length > 120 ? line.slice(0, 117) + '...' : line,
});
break;
}
}
return out;
@ -75,40 +132,90 @@ export async function scan(targetPath, discovery) {
const findings = [];
let filesScanned = 0;
// Files already scanned in-file below — an @import resolving to one of these
// is reported by its own iteration, not duplicated as an import finding.
const discoveredClaudeMd = new Set(
discovery.files.filter(f => f.type === 'claude-md').map(f => f.absPath));
// @imported files reported once, even when several CLAUDE.md files import them.
const reportedImports = new Set();
for (const f of discovery.files) {
if (f.type !== 'claude-md') continue;
filesScanned++;
const content = await readTextFile(f.absPath);
if (!content) continue;
const volatile = findVolatileLines(content);
if (volatile.length === 0) continue;
// --- In-file volatility (unchanged behavior) ---
const volatile = findVolatileLines(content);
// Skip volatility that's already covered by TOK Pattern A (lines 130) —
// CPS' value is in the 31150 range. Pattern A handles 130.
const beyondTopThirty = volatile.filter(v => v.line > 30);
if (beyondTopThirty.length === 0) continue;
if (beyondTopThirty.length > 0) {
const evidence =
beyondTopThirty.slice(0, 5)
.map(v => `line ${v.line} (${v.label}): ${v.snippet}`)
.join('; ');
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: 'Volatile content inside cached prefix breaks reuse',
description:
`${f.relPath || f.absPath} contains ${beyondTopThirty.length} volatile ` +
`entr${beyondTopThirty.length === 1 ? 'y' : 'ies'} between lines 31 and ` +
`${CACHED_PREFIX_LINES}. The prompt cache covers the file's prefix; ` +
'any volatility forces a fresh cache write from that line down on every turn.',
file: f.absPath,
evidence,
recommendation:
'Move volatile sections (timestamps, !shell-exec, ${VAR} substitutions, dated logs) ' +
`below line ${CACHED_PREFIX_LINES} or extract them to an @import-ed file outside the ` +
'cached prefix. Stable content above, volatile content below.',
category: 'token-efficiency',
}));
}
const evidence =
beyondTopThirty.slice(0, 5)
.map(v => `line ${v.line} (${v.label}): ${v.snippet}`)
.join('; ');
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: 'Volatile content inside cached prefix breaks reuse',
description:
`${f.relPath || f.absPath} contains ${beyondTopThirty.length} volatile ` +
`entr${beyondTopThirty.length === 1 ? 'y' : 'ies'} between lines 31 and ` +
`${CACHED_PREFIX_LINES}. The prompt cache covers the file's prefix; ` +
'any volatility forces a fresh cache write from that line down on every turn.',
file: f.absPath,
evidence,
recommendation:
'Move volatile sections (timestamps, !shell-exec, ${VAR} substitutions, dated logs) ' +
`below line ${CACHED_PREFIX_LINES} or extract them to an @import-ed file outside the ` +
'cached prefix. Stable content above, volatile content below.',
category: 'token-efficiency',
}));
// --- v5.10 B6: volatility inside @imported files ---
// @import-ed content is inlined into the cached prefix at the import site.
// TOK Pattern A and the in-file scan above never look past the importing
// file, so volatility in an imported file is otherwise invisible. We scan
// direct imports only (one hop); IMP owns deep-chain analysis. The whole
// imported-file prefix counts (no lines-130 skip — that exclusion is
// root-file-specific to avoid Pattern A overlap, which does not reach here).
for (const imp of findImports(content)) {
if (imp.line > CACHED_PREFIX_LINES) continue; // import site outside prefix
const resolved = resolveImportPath(imp.path, f.absPath);
if (discoveredClaudeMd.has(resolved)) continue; // scanned in its own iteration
if (reportedImports.has(resolved)) continue;
reportedImports.add(resolved);
const importedContent = await readTextFile(resolved);
if (!importedContent) continue;
const importedVolatile = findVolatileLines(importedContent);
if (importedVolatile.length === 0) continue;
const importEvidence =
`imported by ${f.relPath || f.absPath} (@${imp.path} at line ${imp.line}); ` +
importedVolatile.slice(0, 5)
.map(v => `line ${v.line} (${v.label}): ${v.snippet}`)
.join('; ');
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: 'Volatile content in @imported file breaks cached prefix',
description:
`@${imp.path} (imported by ${f.relPath || f.absPath} at line ${imp.line}) contains ` +
`${importedVolatile.length} volatile entr${importedVolatile.length === 1 ? 'y' : 'ies'} ` +
`within its first ${CACHED_PREFIX_LINES} lines. @import-ed content is inlined into the ` +
'prompt-cache prefix, so volatility there forces a fresh cache write every turn — even ' +
'when the importing CLAUDE.md is itself byte-stable.',
file: resolved,
evidence: importEvidence,
recommendation:
'Move volatile content (timestamps, !shell-exec, ${VAR} substitutions, dated logs) out ' +
'of the @imported file, or import it below the cached-prefix window. Keep imported config ' +
'byte-stable so the importing file\'s cache survives.',
category: 'token-efficiency',
}));
}
}
return scannerResult(SCANNER, 'ok', findings, filesScanned, Date.now() - start);

124
scanners/campaign-cli.mjs Normal file
View file

@ -0,0 +1,124 @@
#!/usr/bin/env node
/**
* campaign-cli read-only reporter for the durable machine-wide campaign ledger
* (v5.7 Fase 2, Block 3b).
*
* Mirrors the knowledge-refresh-cli precedent: it is the DETERMINISTIC, READ-ONLY half
* of the hybrid motor. It loads the campaign ledger (the durable file that sits ABOVE
* individual config-audit sessions), validates it, and emits the repo list + a
* machine-wide roll-up as JSON. It NEVER writes the ledger initialization and every
* status transition belong to the command layer (Block 3c `/config-audit campaign`),
* which calls the lib's pure transforms + saveLedger only on explicit, human-approved
* action. A missing ledger file is reported gracefully (initialized:false), NEVER created.
*
* Naming: `-cli` suffix NOT an orchestrated scanner (the scan-orchestrator only loads
* scanner modules), so the scanner count is unchanged and the snapshot suite stays
* byte-stable.
*
* Usage:
* node campaign-cli.mjs [--ledger-file <path>] [--output-file <path>]
*
* Exit codes: 0 = initialized & valid, 1 = not initialized yet (advisory), 3 = error.
*/
import { resolve } from 'node:path';
import { writeOutputFile } from './lib/write-output.mjs';
import {
loadLedger,
validateLedger,
rollUp,
buildBacklog,
defaultLedgerPath,
} from './lib/campaign-ledger.mjs';
/**
* Usage error. Throws rather than calling process.exit(): exit() discards
* unflushed stdout when stdout is a pipe. The top-level catch prints the same
* `Error: ` text and sets the same exit code 3, so callers see no difference.
*/
class CliUsageError extends Error {}
function fail(message) {
throw new CliUsageError(message);
}
async function main() {
const args = process.argv.slice(2);
let ledgerFile = null;
let outputFile = null;
for (let i = 0; i < args.length; i++) {
const a = args[i];
if (a === '--ledger-file' && args[i + 1]) ledgerFile = args[++i];
else if (a === '--output-file' && args[i + 1]) outputFile = args[++i];
// A flag we do not understand must fail loudly. Silently dropping it lets a
// caller-side mistake — a typo'd `--ledger-file`, or a shell that did not
// word-split "--flag value" into two argv entries — produce a confident
// report about the wrong ledger.
else if (a.startsWith('--')) fail(`unknown flag "${a}"`);
}
const ledgerPath = resolve(ledgerFile || defaultLedgerPath());
let ledger;
try {
// loadLedger returns null on ENOENT (graceful first run) and throws on parse error.
ledger = await loadLedger(ledgerPath);
} catch (err) {
fail(`could not read ledger at ${ledgerPath}: ${err.message}`);
}
let payload;
let exitCode;
if (ledger === null) {
// Graceful first run — the ledger does not exist yet. We DO NOT create it; that is
// the command layer's job (Block 3c), on explicit human-approved action.
payload = {
status: 'ok',
initialized: false,
ledgerPath,
schemaVersion: null,
createdDate: null,
updatedDate: null,
repos: [],
rollUp: rollUp({ repos: [] }),
backlog: buildBacklog({ repos: [] }),
};
exitCode = 1; // advisory: there is no campaign to report yet
} else {
const { valid, errors } = validateLedger(ledger);
if (!valid) {
fail(`ledger at ${ledgerPath} is invalid:\n - ${errors.join('\n - ')}`);
}
payload = {
status: 'ok',
initialized: true,
ledgerPath,
schemaVersion: ledger.schemaVersion,
createdDate: ledger.createdDate,
updatedDate: ledger.updatedDate,
repos: ledger.repos,
rollUp: rollUp(ledger),
backlog: buildBacklog(ledger),
};
exitCode = 0;
}
const json = JSON.stringify(payload, null, 2);
if (outputFile) await writeOutputFile(outputFile, json, 'utf-8');
else process.stdout.write(json + '\n');
process.exitCode = exitCode;
}
const isDirectRun =
process.argv[1] && resolve(process.argv[1]) === resolve(new URL(import.meta.url).pathname);
if (isDirectRun) {
main().catch((err) => {
const prefix = err instanceof CliUsageError ? 'Error' : 'Fatal';
process.stderr.write(`${prefix}: ${err.message}\n`);
process.exitCode = 3;
});
}

View file

@ -0,0 +1,173 @@
#!/usr/bin/env node
/**
* campaign-export-cli export a tracked repo's action plan into that repo's own `docs/`
* (v5.7 Fase 2, Block 4c).
*
* Block 4b built the cross-repo prioritized backlog; this is the "plan export" half of Block 4c.
* Given a repo tracked in the campaign ledger, it resolves the repo's linked config-audit
* session, reads that session's `action-plan.md`, and assembles (via the pure
* `campaign-export` lib) a `docs/config-audit-plan-<sessionId>.md` document carrying a
* provenance header + the verbatim plan ("planer følger arbeidsstedet").
*
* Read-only by DEFAULT (a dry-run preview that returns the assembled `document` + `targetPath`
* so the command can show the user what will be written). The actual write happens ONLY under
* the opt-in `--write` flag which the `/config-audit campaign` command invokes solely after
* explicit human approval (Verifiseringsplikt nothing auto-written). Writing the file
* faithfully (a byte-exact copy of the assembled document) is the CLI's job, not the LLM's, so
* a 200-line plan is never re-typed and cannot drift.
*
* Execution is NOT here: Block 4c reuses the existing `/config-audit implement` (backup +
* apply + verify) + `/config-audit rollback`. This CLI only exports the durable record.
*
* Naming: `-cli` suffix NOT an orchestrated scanner, so the scanner count is unchanged and
* the snapshot suite stays byte-stable.
*
* Usage:
* node campaign-export-cli.mjs --repo <path> [--write]
* [--ledger-file <p>] [--sessions-dir <p>] [--reference-date <YYYY-MM-DD>] [--output-file <p>]
*
* Exit codes: 0 = exportable (preview ready, or written under --write),
* 1 = advisory: repo tracked but not exportable yet (no linked session / no plan),
* 3 = error (missing --repo, untracked repo, no/corrupt ledger, unreadable plan).
*/
import { resolve, join, dirname } from 'node:path';
import { homedir } from 'node:os';
import { readFile, writeFile, mkdir } from 'node:fs/promises';
import { writeOutputFile } from './lib/write-output.mjs';
import {
loadLedger,
validateLedger,
defaultLedgerPath,
} from './lib/campaign-ledger.mjs';
import { planExportPath, buildPlanExportDocument } from './lib/campaign-export.mjs';
const DATE_RE = /^\d{4}-\d{2}-\d{2}$/;
/**
* Usage error. Throws rather than calling process.exit(): exit() discards
* unflushed stdout when stdout is a pipe. The top-level catch prints the same
* `Error: ` text and sets the same exit code 3, so callers see no difference.
*/
class CliUsageError extends Error {}
function fail(message) {
throw new CliUsageError(message);
}
/** Default session store: next to the ledger, OUTSIDE the plugin dir. */
function defaultSessionsDir() {
return join(homedir(), '.claude', 'config-audit', 'sessions');
}
function parseArgs(argv) {
const flags = { repo: null, ledgerFile: null, sessionsDir: null, referenceDate: null, outputFile: null, write: false };
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
if (a === '--repo' && argv[i + 1] !== undefined) flags.repo = argv[++i];
else if (a === '--ledger-file' && argv[i + 1] !== undefined) flags.ledgerFile = argv[++i];
else if (a === '--sessions-dir' && argv[i + 1] !== undefined) flags.sessionsDir = argv[++i];
else if (a === '--reference-date' && argv[i + 1] !== undefined) flags.referenceDate = argv[++i];
else if (a === '--output-file' && argv[i + 1] !== undefined) flags.outputFile = argv[++i];
else if (a === '--write') flags.write = true;
else if (a.startsWith('--')) fail(`unknown flag "${a}"`);
else fail(`unexpected argument "${a}"`);
}
return flags;
}
async function emit(payload, outputFile, exitCode) {
const json = JSON.stringify(payload, null, 2);
if (outputFile) await writeOutputFile(outputFile, json, 'utf-8');
else process.stdout.write(json + '\n');
process.exitCode = exitCode;
}
async function main() {
const flags = parseArgs(process.argv.slice(2));
if (!flags.repo) fail('--repo <path> is required');
if (flags.referenceDate && !DATE_RE.test(flags.referenceDate)) fail('--reference-date must be YYYY-MM-DD');
const ledgerPath = resolve(flags.ledgerFile || defaultLedgerPath());
const sessionsDir = resolve(flags.sessionsDir || defaultSessionsDir());
const repoPath = resolve(flags.repo);
// The clock is read here ONLY — passed to the pure lib as the injected `now`.
const now = flags.referenceDate || new Date().toISOString().slice(0, 10);
let ledger;
try {
ledger = await loadLedger(ledgerPath);
} catch (err) {
fail(`could not read ledger at ${ledgerPath}: ${err.message}`);
}
if (ledger === null) fail(`no campaign ledger at ${ledgerPath} — run "/config-audit campaign init" first`);
const { valid, errors } = validateLedger(ledger);
if (!valid) fail(`ledger at ${ledgerPath} is invalid:\n - ${errors.join('\n - ')}`);
const repo = ledger.repos.find((r) => r.path === repoPath);
if (!repo) fail(`repo "${repoPath}" is not tracked in the campaign — add it first`);
const repoInfo = { path: repo.path, name: repo.name, status: repo.status, sessionId: repo.sessionId ?? null };
// Gate 1: the repo must have a linked session (set via `set-status … --session <id>`).
if (typeof repo.sessionId !== 'string' || repo.sessionId.trim() === '') {
return emit(
{ status: 'ok', action: 'export', repo: repoInfo, exportable: false, problems: ['no-session-linked'],
written: false, targetPath: null, document: null },
flags.outputFile,
1,
);
}
// Gate 2: that session must carry an action-plan.md (i.e. `/config-audit plan` has run).
const sourcePlanPath = join(sessionsDir, repo.sessionId, 'action-plan.md');
let planMarkdown;
try {
planMarkdown = await readFile(sourcePlanPath, 'utf-8');
} catch (err) {
if (err && err.code === 'ENOENT') {
return emit(
{ status: 'ok', action: 'export', repo: repoInfo, sessionId: repo.sessionId, sourcePlanPath,
exportable: false, problems: ['no-action-plan'], written: false, targetPath: null, document: null },
flags.outputFile,
1,
);
}
fail(`could not read action plan at ${sourcePlanPath}: ${err.message}`);
}
const targetPath = planExportPath(repo.path, repo.sessionId);
const document = buildPlanExportDocument({
repoName: repo.name,
repoPath: repo.path,
sessionId: repo.sessionId,
planMarkdown,
now,
});
let written = false;
if (flags.write) {
await mkdir(dirname(targetPath), { recursive: true });
await writeFile(targetPath, document, 'utf-8');
written = true;
}
return emit(
{ status: 'ok', action: 'export', repo: repoInfo, sessionId: repo.sessionId, sourcePlanPath,
exportable: true, problems: [], written, targetPath, document },
flags.outputFile,
0,
);
}
const isDirectRun =
process.argv[1] && resolve(process.argv[1]) === resolve(new URL(import.meta.url).pathname);
if (isDirectRun) {
main().catch((err) => {
const prefix = err instanceof CliUsageError ? 'Error' : 'Fatal';
process.stderr.write(`${prefix}: ${err.message}\n`);
process.exitCode = 3;
});
}

View file

@ -0,0 +1,289 @@
#!/usr/bin/env node
/**
* campaign-write-cli the human-approved WRITE half of the durable campaign ledger
* (v5.7 Fase 2, Block 3c).
*
* Sibling of the read-only `campaign-cli`: where that one only reports, this one mutates.
* Every mutation is routed through the pure, invariant-enforcing lib transforms
* (`createLedger`/`addRepo`/`setRepoStatus`) + `saveLedger` so path normalization/dedup,
* idempotent add, the status-lifecycle guard, and the `updatedDate` bump are never
* re-implemented by hand. The `/config-audit campaign` command is a thin opus orchestrator:
* it reports (via campaign-cli), proposes a change, and only on explicit human approval
* invokes a single subcommand here (Verifiseringsplikt). It NEVER auto-writes.
*
* Determinism mirrors the lib + the knowledge-refresh CLI: `--reference-date` is the only
* place the clock is read (defaulting to today), and it is passed to the transforms as the
* injected `now`, so the persisted stamps are fully testable.
*
* Naming: `-cli` suffix NOT an orchestrated scanner (the scan-orchestrator only loads
* scanner modules), so the scanner count is unchanged and the snapshot suite stays
* byte-stable.
*
* Usage:
* node campaign-write-cli.mjs init [--ledger-file <p>] [--reference-date <YYYY-MM-DD>]
* node campaign-write-cli.mjs add <path>... [--name <n>] [--ledger-file <p>] [--reference-date <d>]
* node campaign-write-cli.mjs set-status <path> <status>
* [--findings '<json>'] [--session <id>]
* [--ledger-file <p>] [--reference-date <d>]
* node campaign-write-cli.mjs refresh-tokens [--ledger-file <p>] [--reference-date <d>]
* (all accept [--output-file <p>] to write the result payload to a file instead of stdout)
*
* `refresh-tokens` is the live cross-repo token sweep (v5.9 B2b): for every tracked
* repo it runs the manifest's always-loaded accounting (readActiveConfig buildManifest)
* and splits each source into the shared global layer vs the repo's per-repo delta
* (splitManifestByOwnership). The shared layer is HOME-derived and identical across
* repos, so it is captured ONCE (from the first successful read) and stored at the
* ledger root; each repo gets only its delta. This is the IO half of the machine-wide
* token roll-up whose pure data model shipped in B2a.
*
* Exit codes: 0 = write performed (or no repos to sweep, a benign no-op), 1 = advisory
* no-op (init when already initialized), 3 = error (unknown subcommand, bad args,
* invalid status, untracked repo, no/corrupt ledger).
*/
import { resolve } from 'node:path';
import { stat } from 'node:fs/promises';
import { writeOutputFile } from './lib/write-output.mjs';
import {
createLedger,
addRepo,
setRepoStatus,
setSharedGlobal,
setRepoTokens,
rollUp,
loadLedger,
saveLedger,
defaultLedgerPath,
} from './lib/campaign-ledger.mjs';
import { readActiveConfig } from './lib/active-config-reader.mjs';
import { buildManifest, splitManifestByOwnership } from './manifest.mjs';
const DATE_RE = /^\d{4}-\d{2}-\d{2}$/;
/**
* Usage error. Throws rather than calling process.exit(): exit() discards
* unflushed stdout when stdout is a pipe. The top-level catch prints the same
* `Error: ` text and sets the same exit code 3, so callers see no difference.
*/
class CliUsageError extends Error {}
function fail(message) {
throw new CliUsageError(message);
}
/** Parse argv into a subcommand, positional args, and the flag map. */
function parseArgs(argv) {
const positionals = [];
const flags = { ledgerFile: null, referenceDate: null, outputFile: null, name: null, findings: null, session: null };
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
if (a === '--ledger-file' && argv[i + 1] !== undefined) flags.ledgerFile = argv[++i];
else if (a === '--reference-date' && argv[i + 1] !== undefined) flags.referenceDate = argv[++i];
else if (a === '--output-file' && argv[i + 1] !== undefined) flags.outputFile = argv[++i];
else if (a === '--name' && argv[i + 1] !== undefined) flags.name = argv[++i];
else if (a === '--findings' && argv[i + 1] !== undefined) flags.findings = argv[++i];
else if (a === '--session' && argv[i + 1] !== undefined) flags.session = argv[++i];
else if (a.startsWith('--')) fail(`unknown flag "${a}"`);
else positionals.push(a);
}
return { subcommand: positionals[0], rest: positionals.slice(1), flags };
}
/**
* Can this path actually be read as a repo right now?
*
* `readActiveConfig` resolves any string and its sub-readers all tolerate ENOENT, so a repo
* that does not exist yields an EMPTY config instead of an error. Without this check the
* sweep records a phantom repo as successfully swept with 0 tokens, and the machine-wide
* bill claims coverage it does not have. A missing path is reported, never rejected an
* unmounted volume is a legitimate reason for a tracked repo to be absent today.
*/
async function isReadableRepoDir(path) {
try {
return (await stat(path)).isDirectory();
} catch {
return false;
}
}
/** Load an existing ledger, treating a parse error as a hard failure (never clobber corrupt data). */
async function loadOrFail(path) {
try {
return await loadLedger(path); // null on ENOENT (no ledger yet)
} catch (err) {
fail(`could not read ledger at ${path}: ${err.message}`);
}
}
async function emit(payload, outputFile, exitCode) {
const json = JSON.stringify(payload, null, 2);
if (outputFile) await writeOutputFile(outputFile, json, 'utf-8');
else process.stdout.write(json + '\n');
process.exitCode = exitCode;
}
async function main() {
const { subcommand, rest, flags } = parseArgs(process.argv.slice(2));
if (!subcommand) fail('a subcommand is required: init | add | set-status | refresh-tokens');
const ledgerPath = resolve(flags.ledgerFile || defaultLedgerPath());
if (flags.referenceDate && !DATE_RE.test(flags.referenceDate)) fail('--reference-date must be YYYY-MM-DD');
// The clock is read here ONLY — the transforms take this injected `now` and stay pure.
const now = flags.referenceDate || new Date().toISOString().slice(0, 10);
if (subcommand === 'init') {
const existing = await loadOrFail(ledgerPath);
if (existing !== null) {
// Advisory no-op: never wipe an existing campaign.
return emit(
{ status: 'ok', action: 'init', written: false, alreadyInitialized: true, ledgerPath },
flags.outputFile,
1,
);
}
const ledger = createLedger({ now });
await saveLedger(ledgerPath, ledger);
return emit(
{
status: 'ok', action: 'init', written: true, alreadyInitialized: false, ledgerPath,
schemaVersion: ledger.schemaVersion, createdDate: ledger.createdDate, updatedDate: ledger.updatedDate,
repos: ledger.repos, rollUp: rollUp(ledger),
},
flags.outputFile,
0,
);
}
if (subcommand === 'add') {
const paths = rest;
if (paths.length === 0) fail('add requires at least one repo path');
const loaded = await loadOrFail(ledgerPath);
const autoInitialized = loaded === null;
let ledger = loaded === null ? createLedger({ now }) : loaded;
const added = [];
const addedUnverified = [];
const skipped = [];
for (const p of paths) {
const resolved = resolve(p);
const present = ledger.repos.some((r) => r.path === resolved);
// --name applies only to a lone path; multi-add lets the lib derive each basename.
const name = paths.length === 1 ? flags.name || undefined : undefined;
ledger = addRepo(ledger, { path: p, name }, { now });
if (present) skipped.push(resolved);
else if (await isReadableRepoDir(resolved)) added.push(resolved);
// Tracked either way, but never silently vouched for: the command reports these
// separately so a typo does not become a permanent phantom row in the backlog.
else addedUnverified.push(resolved);
}
await saveLedger(ledgerPath, ledger);
return emit(
{
status: 'ok', action: 'add', written: true, autoInitialized, ledgerPath,
added, addedUnverified, skipped, repos: ledger.repos, rollUp: rollUp(ledger),
},
flags.outputFile,
0,
);
}
if (subcommand === 'set-status') {
const [path, status] = rest;
if (!path || !status) fail('set-status requires <path> <status>');
const ledger = await loadOrFail(ledgerPath);
if (ledger === null) fail(`no ledger at ${ledgerPath} — run "init" or "add" first`);
let findingsBySeverity;
if (flags.findings !== null) {
try {
findingsBySeverity = JSON.parse(flags.findings);
} catch (err) {
fail(`--findings must be valid JSON: ${err.message}`);
}
}
const opts = { now };
if (findingsBySeverity !== undefined) opts.findingsBySeverity = findingsBySeverity;
if (flags.session !== null) opts.sessionId = flags.session;
let next;
try {
next = setRepoStatus(ledger, path, status, opts);
} catch (err) {
// RangeError (bad status) or Error (untracked repo) → caller error.
fail(err.message);
}
await saveLedger(ledgerPath, next);
const resolved = resolve(path);
return emit(
{
status: 'ok', action: 'set-status', written: true, ledgerPath,
repo: next.repos.find((r) => r.path === resolved), repos: next.repos, rollUp: rollUp(next),
},
flags.outputFile,
0,
);
}
if (subcommand === 'refresh-tokens') {
const ledger0 = await loadOrFail(ledgerPath);
if (ledger0 === null) fail(`no ledger at ${ledgerPath} — run "init" or "add" first`);
let ledger = ledger0;
const swept = [];
const skipped = [];
// The shared global layer is HOME-derived and identical across every repo, so it
// is captured ONCE (from the first repo that reads cleanly) and stored at the
// ledger root — the structural guard against the historic shared-layer double-count.
let sharedSummary = null;
for (const repo of ledger0.repos) {
// Check readability FIRST: readActiveConfig returns an empty config for a path that
// does not exist rather than throwing, so the catch below would never see it and the
// repo would be recorded as swept with a 0-token delta — a bill that looks complete.
if (!(await isReadableRepoDir(repo.path))) {
skipped.push({ path: repo.path, reason: 'repo path is not readable (does not exist or is not a directory)' });
continue;
}
let split;
try {
const activeConfig = await readActiveConfig(repo.path, { verbose: false });
const { sources } = buildManifest(activeConfig);
split = splitManifestByOwnership(sources);
} catch (err) {
skipped.push({ path: repo.path, reason: err.message });
continue;
}
if (sharedSummary === null) sharedSummary = split.shared;
ledger = setRepoTokens(ledger, repo.path, split.delta, { now });
swept.push(repo.path);
}
if (sharedSummary !== null) ledger = setSharedGlobal(ledger, sharedSummary, { now });
const written = swept.length > 0;
if (written) await saveLedger(ledgerPath, ledger);
return emit(
{
status: 'ok', action: 'refresh-tokens', written, ledgerPath,
swept, skipped, sharedGlobal: ledger.sharedGlobal ?? null, rollUp: rollUp(ledger),
},
flags.outputFile,
0,
);
}
fail(`unknown subcommand "${subcommand}" — expected init | add | set-status | refresh-tokens`);
}
const isDirectRun =
process.argv[1] && resolve(process.argv[1]) === resolve(new URL(import.meta.url).pathname);
if (isDirectRun) {
main().catch((err) => {
const prefix = err instanceof CliUsageError ? 'Error' : 'Fatal';
process.stderr.write(`${prefix}: ${err.message}\n`);
process.exitCode = 3;
});
}

View file

@ -9,11 +9,27 @@ import { finding, scannerResult, resetCounter } from './lib/output.mjs';
import { SEVERITY } from './lib/severity.mjs';
import { parseFrontmatter, extractSections, findImports } from './lib/yaml-parser.mjs';
import { lineCount, truncate } from './lib/string-utils.mjs';
import { CONTEXT_WINDOW_ANCHOR, LARGE_CONTEXT_WINDOW, LARGE_CONTEXT_SCALE, scaleForWindow, withCommas } from './lib/context-window.mjs';
import { dirname } from 'node:path';
const SCANNER = 'CML';
const MAX_RECOMMENDED_LINES = 200;
const MAX_ABSOLUTE_LINES = 500;
// Shared remediation for the char-budget finding (byte-identical across the
// default and the B8 window-calibrated branches).
const CHAR_BUDGET_RECOMMENDATION =
'Split detail into @imports and .claude/rules/ files so only the relevant rules load, and keep the top of CLAUDE.md byte-stable for cache hits.';
// Claude Code's own startup warning ("Large CLAUDE.md will impact performance
// (X chars > 40.0k)") fires once a CLAUDE.md passes ~40.0k chars on a
// 200k-context model. CC 2.1.169 made that threshold scale with the model's
// context window. We mirror it in the same unit CC uses (chars, not lines):
// anchor on the conservative 200k window (we cannot observe the user's window,
// and the anchor fires earliest) and disclose the relaxed 1M figure.
const CLAUDE_MD_CHAR_WARN_ANCHOR = 40_000; // chars @ 200k context (CC startup warning)
const CLAUDE_MD_CHAR_WARN_LARGE = CLAUDE_MD_CHAR_WARN_ANCHOR * LARGE_CONTEXT_SCALE; // 200,000 @ 1M
/** Recommended sections for a project CLAUDE.md */
const RECOMMENDED_SECTIONS = [
{ pattern: /project|overview|description|what/i, label: 'Project overview' },
@ -28,10 +44,20 @@ const RECOMMENDED_SECTIONS = [
* @param {{ files: import('./lib/file-discovery.mjs').ConfigFile[] }} discovery
* @returns {Promise<object>}
*/
export async function scan(targetPath, discovery) {
export async function scan(targetPath, discovery, opts = {}) {
const start = Date.now();
const claudeFiles = discovery.files.filter(f => f.type === 'claude-md');
// B8 — calibrate the char-budget threshold to the resolved context window. The
// default (no opts) is the conservative 200k anchor (40k chars) at full
// severity — byte-identical to the pre-B8 finding. An unknown (advisory) window
// keeps the anchor but downgrades the finding to info instead of a breach.
const cw = opts.contextWindow;
const window = (cw && typeof cw.window === 'number') ? cw.window : CONTEXT_WINDOW_ANCHOR;
const advisory = !!(cw && cw.advisory);
const isDefaultWindow = window === CONTEXT_WINDOW_ANCHOR && !advisory;
const charThreshold = scaleForWindow(CLAUDE_MD_CHAR_WARN_ANCHOR, window);
if (claudeFiles.length === 0) {
return scannerResult(SCANNER, 'ok', [
finding({
@ -58,16 +84,37 @@ export async function scan(targetPath, discovery) {
const sections = extractSections(body);
const imports = findImports(content);
// A nested (subdirectory) CLAUDE.md is NOT re-injected after a context
// compaction — only the project-root CLAUDE.md is (context-window.md). Its
// instructions silently drop until a file in that directory is read again.
const relDir = dirname(file.relPath);
if (file.scope === 'project' && relDir !== '.' && relDir !== '.claude' && lines > 5) {
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: 'Nested CLAUDE.md is not re-injected after compaction',
description: `${file.relPath} is a nested (subdirectory) CLAUDE.md. It loads when Claude reads a file in that directory, but after a context compaction it is not re-injected (only the project-root CLAUDE.md is) — its instructions silently drop until a file in that directory is read again.`,
file: file.absPath,
evidence: `${lines} lines, nested (scope=project, dir="${relDir}")`,
recommendation: 'If these instructions must always apply, move the must-hold parts to the project-root CLAUDE.md (re-injected after compaction). Keep nested CLAUDE.md for guidance only needed when working in that directory.',
autoFixable: false,
}));
}
// --- Length checks ---
// Raw line count is no longer an absolute adherence threshold: CC 2.1.169
// scales the "too long" warning by context window, and cache-prefix
// stability (not line count) is the dominant cost driver on large-context
// models. These are MEDIUM token-cost signals, not a HIGH adherence cliff.
if (lines > MAX_ABSOLUTE_LINES) {
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.high,
severity: SEVERITY.medium,
title: 'CLAUDE.md exceeds 500 lines',
description: `${file.relPath} has ${lines} lines. Files over 500 lines significantly reduce Claude's adherence to instructions.`,
description: `${file.relPath} has ${lines} lines. A file this size loads in full on every turn (token cost) and, on smaller-context models, can crowd out instructions. Large-context models tolerate longer files when the cache prefix stays stable — raw line count is no longer an absolute adherence threshold (CC 2.1.169 scales it by context window).`,
file: file.absPath,
evidence: `${lines} lines`,
recommendation: 'Split into @imports and .claude/rules/ files. Keep CLAUDE.md under 200 lines.',
recommendation: 'Split into @imports and .claude/rules/ files, and keep the top of CLAUDE.md byte-stable for cache hits (see token / cache-prefix findings). Under ~200 lines stays safest across models.',
autoFixable: false,
}));
} else if (lines > MAX_RECOMMENDED_LINES) {
@ -75,7 +122,7 @@ export async function scan(targetPath, discovery) {
scanner: SCANNER,
severity: SEVERITY.medium,
title: 'CLAUDE.md exceeds recommended 200 lines',
description: `${file.relPath} has ${lines} lines. Best practice is under 200 lines for optimal adherence.`,
description: `${file.relPath} has ${lines} lines. Under ~200 lines is the safe default across models; larger is fine on large-context models when the cache prefix stays stable. A long file still costs tokens every turn.`,
file: file.absPath,
evidence: `${lines} lines`,
recommendation: 'Consider using @imports or .claude/rules/ for detailed content.',
@ -83,6 +130,44 @@ export async function scan(targetPath, discovery) {
}));
}
// --- Char budget (mirrors Claude Code's own startup warning) ---
// Keyed on chars, not lines: CC's "Large CLAUDE.md will impact performance"
// warning is char-based (~40.0k @ 200k context) and CC 2.1.169 scales that
// threshold with the context window. A file can be long by lines yet under
// this budget (short lines), or short by lines yet over it (long lines), so
// this is complementary to the line-count checks above.
const chars = content.length;
if (chars > charThreshold) {
if (isDefaultWindow) {
// Conservative 200k anchor — byte-identical to the pre-B8 finding.
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: 'CLAUDE.md exceeds Claude Code\'s performance-warning threshold',
description: `${file.relPath} is ${withCommas(chars)} chars. Claude Code shows a startup warning ("Large CLAUDE.md will impact performance ... chars > 40.0k") once a CLAUDE.md passes ~40.0k chars on a 200k-context model — it loads in full on every turn. CC 2.1.169 scales that threshold with the context window, so on a ${withCommas(LARGE_CONTEXT_WINDOW)}-token model it relaxes to ~${withCommas(CLAUDE_MD_CHAR_WARN_LARGE)} chars and you are likely within it.`,
file: file.absPath,
evidence: `${withCommas(chars)} chars > 40.0k (200k-context anchor; ~${withCommas(CLAUDE_MD_CHAR_WARN_LARGE)} at ${withCommas(LARGE_CONTEXT_WINDOW)} context). This is an estimate, not measured telemetry.`,
recommendation: CHAR_BUDGET_RECOMMENDATION,
autoFixable: false,
}));
} else {
// B8 — window-calibrated. Advisory (unknown window) downgrades to info.
const winLabel = withCommas(window);
const threshLabel = withCommas(charThreshold);
findings.push(finding({
scanner: SCANNER,
severity: advisory ? SEVERITY.info : SEVERITY.medium,
title: 'CLAUDE.md exceeds Claude Code\'s performance-warning threshold',
description: `${file.relPath} is ${withCommas(chars)} chars, over the ~${threshLabel}-char performance-warning threshold Claude Code applies at a ${winLabel}-token context window (it scales the ~40.0k-char @ 200k warning by the context window, CC 2.1.169) — it loads in full on every turn.` +
(advisory ? ' Your context window is unknown, so this anchors on the conservative 200k window — advisory.' : ''),
file: file.absPath,
evidence: `${withCommas(chars)} chars > ${threshLabel} (calibrated to a ${winLabel}-token context window). This is an estimate, not measured telemetry.`,
recommendation: CHAR_BUDGET_RECOMMENDATION,
autoFixable: false,
}));
}
}
// --- Empty file ---
if (lines < 3) {
findings.push(finding({

View file

@ -5,27 +5,33 @@
* Finding IDs: CA-CNF-NNN
*/
import { sep } from 'node:path';
import { readTextFile } from './lib/file-discovery.mjs';
import { finding, scannerResult } from './lib/output.mjs';
import { SEVERITY } from './lib/severity.mjs';
import { parseJson } from './lib/yaml-parser.mjs';
import { truncate } from './lib/string-utils.mjs';
import { rulesIntersect } from './lib/permission-rules.mjs';
const SCANNER = 'CNF';
// Keys checked separately or not meaningful to compare
const SKIP_KEYS = new Set(['$schema', 'hooks', 'permissions']);
// Files under `.claude/plugins/` are shipped by installed plugins — the plugin's
// own settings.json/hooks.json plus bundled test fixtures and examples. They are
// not the user's authored cascade and a "conflict" between them is not something
// the user can resolve, so they must be excluded from cross-scope conflict
// analysis. (Other scanners still need active plugin config, so this exclusion is
// CNF-local, not a discovery-level skip. M-BUG-2.)
const PLUGIN_TREE_MARKER = `.claude${sep}plugins${sep}`;
/**
* Extract the tool name prefix from a permission rule.
* e.g., "Bash(npm run *)" "Bash", "Read(src/**)" "Read"
* @param {string} rule
* @returns {{ tool: string, pattern: string }}
* @param {import('./lib/file-discovery.mjs').ConfigFile} file
* @returns {boolean} true if the file is shipped by an installed plugin
*/
function parsePermissionRule(rule) {
const match = rule.match(/^(\w+)\((.+)\)$/);
if (match) return { tool: match[1], pattern: match[2] };
return { tool: rule, pattern: '*' };
function isPluginBundled(file) {
return file.absPath.includes(PLUGIN_TREE_MARKER);
}
/**
@ -74,10 +80,10 @@ export async function scan(targetPath, discovery) {
const start = Date.now();
const findings = [];
// Collect settings files
const settingsFiles = discovery.files.filter(f => f.type === 'settings-json');
// Collect hooks files
const hooksFiles = discovery.files.filter(f => f.type === 'hooks-json');
// Collect settings files (excluding plugin-bundled — see PLUGIN_TREE_MARKER)
const settingsFiles = discovery.files.filter(f => f.type === 'settings-json' && !isPluginBundled(f));
// Collect hooks files (excluding plugin-bundled)
const hooksFiles = discovery.files.filter(f => f.type === 'hooks-json' && !isPluginBundled(f));
const totalFiles = settingsFiles.length + hooksFiles.length;
@ -150,10 +156,8 @@ export async function scan(targetPath, discovery) {
// Check: allow in A, deny in B (and vice versa)
for (const allowRule of aAllow) {
const { tool: aTool, pattern: aPattern } = parsePermissionRule(allowRule);
for (const denyRule of bDeny) {
const { tool: dTool, pattern: dPattern } = parsePermissionRule(denyRule);
if (aTool === dTool && (aPattern === dPattern || aPattern === '*' || dPattern === '*')) {
if (rulesIntersect(allowRule, denyRule)) {
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.high,
@ -169,10 +173,8 @@ export async function scan(targetPath, discovery) {
// Reverse: allow in B, deny in A
for (const allowRule of bAllow) {
const { tool: bTool, pattern: bPattern } = parsePermissionRule(allowRule);
for (const denyRule of aDeny) {
const { tool: dTool, pattern: dPattern } = parsePermissionRule(denyRule);
if (bTool === dTool && (bPattern === dPattern || bPattern === '*' || dPattern === '*')) {
if (rulesIntersect(allowRule, denyRule)) {
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.high,

View file

@ -7,9 +7,11 @@
* intent. Often arises from copy-paste edits where one list was updated and
* the other was forgotten.
*
* Compares tool identity by the bare tool name (everything before the first
* `(`). `Bash(npm:*)` and `Bash` are treated as the same tool for collision
* purposes a deny on `Bash` blocks all `Bash(...)` allows.
* Compares rule identity param-aware (CC 2.1.178 `Tool(param:value)`,
* 2.1.172 `domain:` rules). An allow entry is dead only when some deny entry
* fully COVERS it: a bare `Bash` deny blocks all `Bash(...)` allows, but
* `Agent(model:opus)` deny does NOT kill an `Agent(model:sonnet)` allow.
* Coverage logic lives in `lib/permission-rules.mjs` (shared with CNF).
*
* Finding ID: CA-DIS-NNN. Severity: low.
*
@ -20,21 +22,13 @@ import { readTextFile } from './lib/file-discovery.mjs';
import { finding, scannerResult } from './lib/output.mjs';
import { SEVERITY } from './lib/severity.mjs';
import { parseJson } from './lib/yaml-parser.mjs';
import { dominates, parseRule, isIneffectiveAllowGlob, forbiddenParamRule } from './lib/permission-rules.mjs';
const SCANNER = 'DIS';
/**
* Bare tool name = everything before the first `(`. `Bash(npm:*)` `Bash`.
*/
function bareTool(entry) {
if (typeof entry !== 'string') return null;
const idx = entry.indexOf('(');
return (idx === -1 ? entry : entry.slice(0, idx)).trim();
}
/**
* Find tools whose bare name appears in both deny and allow within the same
* settings.json. Returns array of { tool, allowEntry, denyEntry }.
* Find allow entries that are dead config because some deny entry fully covers
* them. Returns array of { tool, allowEntry, denyEntry }.
*/
function findDenyAllowOverlaps(settings) {
if (!settings || typeof settings !== 'object') return [];
@ -45,25 +39,54 @@ function findDenyAllowOverlaps(settings) {
const denyList = Array.isArray(perms.deny) ? perms.deny : [];
if (allowList.length === 0 || denyList.length === 0) return [];
const denyByBare = new Map();
for (const d of denyList) {
const bare = bareTool(d);
if (bare && !denyByBare.has(bare)) denyByBare.set(bare, d);
}
const overlaps = [];
const seen = new Set();
for (const a of allowList) {
const bare = bareTool(a);
if (!bare) continue;
if (denyByBare.has(bare) && !seen.has(bare)) {
overlaps.push({ tool: bare, allowEntry: a, denyEntry: denyByBare.get(bare) });
seen.add(bare);
if (typeof a !== 'string' || seen.has(a)) continue;
const dominator = denyList.find(d => dominates(d, a));
if (dominator) {
overlaps.push({ tool: parseRule(a).tool, allowEntry: a, denyEntry: dominator });
seen.add(a);
}
}
return overlaps;
}
/**
* Find `permissions.allow` entries that are unanchored tool-name globs Claude
* Code silently skips (e.g. `mcp__*`, `B*`, `*`). They auto-approve nothing but
* the author usually believes they grant access. Returns array of entry strings.
*/
function findIneffectiveAllowGlobs(settings) {
if (!settings || typeof settings !== 'object') return [];
const perms = settings.permissions;
if (!perms || typeof perms !== 'object') return [];
const allowList = Array.isArray(perms.allow) ? perms.allow : [];
return allowList.filter(e => isIneffectiveAllowGlob(e));
}
/**
* Find permission rules CC silently ignores because their `Tool(param:value)`
* key is the tool's own canonicalizing field (`command`, `file_path`, `path`,
* `notebook_path`, `url`). Scans allow + deny + ask so severity can split:
* deny/ask hits are false security, allow hits are dead config. Returns array
* of { list, entry, tool, key, hint }.
*/
function findForbiddenParamRules(settings) {
if (!settings || typeof settings !== 'object') return [];
const perms = settings.permissions;
if (!perms || typeof perms !== 'object') return [];
const results = [];
for (const list of ['allow', 'deny', 'ask']) {
const arr = Array.isArray(perms[list]) ? perms[list] : [];
for (const entry of arr) {
const hit = forbiddenParamRule(entry);
if (hit) results.push({ list, entry, ...hit });
}
}
return results;
}
/**
* Main scanner entry point.
*
@ -82,28 +105,101 @@ export async function scan(targetPath, discovery) {
if (!content) continue;
const parsed = parseJson(content);
if (!parsed) continue;
const overlaps = findDenyAllowOverlaps(parsed);
if (overlaps.length === 0) continue;
const evidence = overlaps.slice(0, 5)
.map(o => `${o.tool}: allow="${o.allowEntry}" + deny="${o.denyEntry}"`)
.join('; ');
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: 'Tool listed in both permissions.deny and permissions.allow',
description:
`${f.relPath || f.absPath} contains ${overlaps.length} tool` +
`${overlaps.length === 1 ? '' : 's'} present in both deny and allow lists. ` +
'The deny list wins — the allow entries are dead config but still load on ' +
'every turn and may confuse future readers about intent.',
file: f.absPath,
evidence,
recommendation:
'Remove the redundant allow entries. If you actually want this tool enabled, ' +
'remove it from the deny list instead. Settings should express intent clearly.',
category: 'permissions-hygiene',
}));
const overlaps = findDenyAllowOverlaps(parsed);
if (overlaps.length > 0) {
const evidence = overlaps.slice(0, 5)
.map(o => `${o.tool}: allow="${o.allowEntry}" + deny="${o.denyEntry}"`)
.join('; ');
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: 'Tool listed in both permissions.deny and permissions.allow',
description:
`${f.relPath || f.absPath} contains ${overlaps.length} tool` +
`${overlaps.length === 1 ? '' : 's'} present in both deny and allow lists. ` +
'The deny list wins — the allow entries are dead config but still load on ' +
'every turn and may confuse future readers about intent.',
file: f.absPath,
evidence,
recommendation:
'Remove the redundant allow entries. If you actually want this tool enabled, ' +
'remove it from the deny list instead. Settings should express intent clearly.',
category: 'permissions-hygiene',
}));
}
const ineffective = findIneffectiveAllowGlobs(parsed);
if (ineffective.length > 0) {
const evidence = `allow: ${ineffective.slice(0, 5).map(e => `"${e}"`).join(', ')}`;
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: 'Ineffective allow wildcard — Claude Code ignores this rule',
description:
`${f.relPath || f.absPath} has ${ineffective.length} permissions.allow ` +
`entr${ineffective.length === 1 ? 'y' : 'ies'} that Claude Code skips: an ` +
'unanchored tool-name wildcard auto-approves nothing. CC accepts allow ' +
'wildcards only after a literal `mcp__<server>__` prefix.',
file: f.absPath,
evidence,
recommendation:
'Replace `*`/`mcp__*` with explicit tool names, or anchor MCP wildcards to ' +
'a server (`mcp__<server>__*`). As written these entries grant nothing.',
category: 'permissions-hygiene',
}));
}
const forbidden = findForbiddenParamRules(parsed);
const falseSecurity = forbidden.filter(x => x.list === 'deny' || x.list === 'ask');
const deadAllow = forbidden.filter(x => x.list === 'allow');
if (falseSecurity.length > 0) {
const evidence = falseSecurity.slice(0, 5)
.map(x => `${x.list}: "${x.entry}" → use ${x.hint}`)
.join('; ');
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: 'Permission rule silently ignored — deny/ask uses a forbidden param key',
description:
`${f.relPath || f.absPath} has ${falseSecurity.length} deny/ask ` +
`rule${falseSecurity.length === 1 ? '' : 's'} whose \`Tool(param:value)\` key is ` +
'the tool\'s own canonicalizing field (`command`/`file_path`/`path`/`notebook_path`/' +
'`url`). Claude Code ignores these and emits a startup warning, so the guard you ' +
'intended does NOT apply — the action you meant to block or gate is effectively ' +
'unrestricted.',
file: f.absPath,
evidence,
recommendation:
'Rewrite each rule with the tool\'s own specifier syntax (e.g. `Bash(rm *)`, ' +
'`Read(./path)`, `WebFetch(domain:host)`). As written these rules block nothing.',
category: 'permissions-hygiene',
}));
}
if (deadAllow.length > 0) {
const evidence = deadAllow.slice(0, 5)
.map(x => `allow: "${x.entry}" → use ${x.hint}`)
.join('; ');
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: 'Permission rule silently ignored — allow uses a forbidden param key (dead config)',
description:
`${f.relPath || f.absPath} has ${deadAllow.length} permissions.allow ` +
`rule${deadAllow.length === 1 ? '' : 's'} using \`Tool(param:value)\` on the tool's ` +
'own canonicalizing field. `param:value` matching applies only to deny/ask rules; ' +
'allow rules use each tool\'s own specifier syntax. Claude Code ignores these and ' +
'emits a startup warning — they grant nothing.',
file: f.absPath,
evidence,
recommendation:
'Replace with the tool\'s specifier syntax (e.g. `Read(./path)`), or remove the ' +
'entry. As written it auto-approves nothing.',
category: 'permissions-hygiene',
}));
}
}
return scannerResult(SCANNER, 'ok', findings, filesScanned, Date.now() - start);

View file

@ -5,17 +5,22 @@
* Compare current configuration against a saved baseline.
* Usage:
* node drift-cli.mjs <path> --save [--name my-baseline]
* node drift-cli.mjs <path> [--baseline my-baseline] [--json]
* node drift-cli.mjs <path> [--baseline my-baseline] [--json] [--output-file path]
* node drift-cli.mjs --list
* Unknown options and value-less --name/--baseline/--output-file exit 3.
* Zero external dependencies.
*/
import { resolve } from 'node:path';
import { writeOutputFile } from './lib/write-output.mjs';
import { runAllScanners } from './scan-orchestrator.mjs';
import { diffEnvelopes, formatDiffReport } from './lib/diff-engine.mjs';
import { saveBaseline, loadBaseline, listBaselines } from './lib/baseline.mjs';
import { humanizeFindings } from './lib/humanizer.mjs';
const BOOL_FLAGS = ['--save', '--list', '--json', '--raw', '--global'];
const VALUE_FLAGS = ['--name', '--baseline', '--output-file'];
async function main() {
const args = process.argv.slice(2);
let targetPath = '.';
@ -25,30 +30,50 @@ async function main() {
let jsonMode = false;
let rawMode = false;
let includeGlobal = false;
let outputFile = null;
// M-BUG-21: this loop used to end in `else if (!arg.startsWith('-')) targetPath = arg`,
// with no unknown-flag branch. An unrecognised flag was dropped silently and its
// VALUE fell through to targetPath — so `--output-file /tmp/x.json` scanned
// /tmp/x.json, a path that does not exist, yielding a near-empty scan and
// therefore permanent phantom drift. A missing value for --name was equally
// silent and destructive: it left baselineName at 'default' and OVERWROTE the
// default baseline. Both now fail loudly (exit 3) instead.
for (let i = 0; i < args.length; i++) {
if (args[i] === '--save') {
save = true;
} else if (args[i] === '--name' && args[i + 1]) {
baselineName = args[++i];
} else if (args[i] === '--baseline' && args[i + 1]) {
baselineName = args[++i];
} else if (args[i] === '--list') {
list = true;
} else if (args[i] === '--json') {
jsonMode = true;
} else if (args[i] === '--raw') {
rawMode = true;
} else if (args[i] === '--global') {
includeGlobal = true;
} else if (!args[i].startsWith('-')) {
targetPath = args[i];
const arg = args[i];
if (BOOL_FLAGS.includes(arg)) {
if (arg === '--save') save = true;
else if (arg === '--list') list = true;
else if (arg === '--json') jsonMode = true;
else if (arg === '--raw') rawMode = true;
else if (arg === '--global') includeGlobal = true;
} else if (VALUE_FLAGS.includes(arg)) {
const value = args[i + 1];
if (value === undefined || value.startsWith('-')) {
throw new Error(`Option ${arg} requires a value.`);
}
if (arg === '--name' || arg === '--baseline') baselineName = value;
else outputFile = value;
i++;
} else if (arg.startsWith('-')) {
throw new Error(
`Unknown option: ${arg}\n` +
`Valid options: ${[...BOOL_FLAGS, ...VALUE_FLAGS].join(' ')}`
);
} else {
targetPath = arg;
}
}
// --- List mode ---
if (list) {
const result = await listBaselines();
// commands/drift.md runs this with `2>/dev/null` (ux-rules rule 2). The
// human listing below goes to stderr, so without --output-file the command
// received 0 bytes and could render nothing. The flag was already accepted
// by the arg parser; only list mode ignored it.
if (outputFile) await writeOutputFile(outputFile, JSON.stringify(result, null, 2) + '\n', 'utf-8');
if (jsonMode || rawMode) {
process.stdout.write(JSON.stringify(result, null, 2) + '\n');
} else {
@ -65,7 +90,7 @@ async function main() {
process.stderr.write('\n━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\n');
}
}
process.exit(0);
return;
}
// --- Save mode ---
@ -84,7 +109,7 @@ async function main() {
process.stderr.write(`\nBaseline "${result.name}" saved to ${result.path}\n`);
process.stderr.write(`Findings: ${envelope.aggregate.total_findings}\n`);
}
process.exit(0);
return;
}
// --- Drift mode (default) ---
@ -103,7 +128,28 @@ async function main() {
process.stderr.write(`Baseline "${baselineName}" not found.\n`);
process.stderr.write(`Save one first: node drift-cli.mjs <path> --save\n`);
}
process.exit(1);
process.exitCode = 1;
return;
}
// M-BUG-27: a baseline carries the path it was saved from, but nothing ever
// compared it against the current target. Diffing a repo against a baseline
// anchored elsewhere produced 100% phantom drift — every baseline finding
// "resolved", every current finding "new" — and reported it as trend
// "improving": a reassuring and entirely false signal, on the DEFAULT
// baseline. The warning goes to stderr in every mode; stdout stays
// byte-identical to the frozen v5.0.0 shape.
const baselineTarget = baseline._baseline?.target_path || '';
const currentTarget = resolve(targetPath);
const anchorMatches = !baselineTarget || baselineTarget === currentTarget;
if (baselineTarget && baselineTarget !== currentTarget) {
process.stderr.write(
`\nWarning: baseline "${baselineName}" was saved from a different target path.\n` +
` baseline: ${baselineTarget}\n` +
` current: ${currentTarget}\n` +
` The two scans cover different trees, so this diff is not a drift signal.\n` +
` Re-anchor with: drift-cli.mjs ${currentTarget} --save --name ${baselineName}\n\n`
);
}
// Run current scan
@ -115,25 +161,41 @@ async function main() {
// Diff
const diff = diffEnvelopes(baseline, current);
// Default mode: humanize finding-bearing diff fields before report rendering.
// `_baselineAnchor` rides here and NOT in the raw shape: commands/drift.md runs
// the CLI under `2>/dev/null`, so the stderr warning above never reaches the
// caller that has to act on it. --json/--raw stay v5.0.0-shaped.
const humanizedDiff = {
...diff,
_baselineAnchor: { matches: anchorMatches, baselineTarget, currentTarget },
newFindings: humanizeFindings(diff.newFindings || []),
resolvedFindings: humanizeFindings(diff.resolvedFindings || []),
unchangedFindings: humanizeFindings(diff.unchangedFindings || []),
movedFindings: humanizeFindings(diff.movedFindings || []),
};
if (jsonMode || rawMode) {
// --json and --raw both write the raw v5.0.0-shape diff (byte-identical).
process.stdout.write(JSON.stringify(diff, null, 2) + '\n');
} else {
// Default mode: humanize finding-bearing diff fields before report rendering.
const humanizedDiff = {
...diff,
newFindings: humanizeFindings(diff.newFindings || []),
resolvedFindings: humanizeFindings(diff.resolvedFindings || []),
unchangedFindings: humanizeFindings(diff.unchangedFindings || []),
movedFindings: humanizeFindings(diff.movedFindings || []),
};
const report = formatDiffReport(humanizedDiff);
process.stderr.write('\n' + report + '\n');
}
// ux-rules rule 2: every scanner Bash call uses `--output-file <path>` and the
// command reads the file with the Read tool. drift-cli had no such flag, and
// its default-mode report goes to stderr — which commands/drift.md discarded
// via `2>/dev/null` while instructing the agent to "read stdout". The command
// captured nothing. Matches posture.mjs: raw diff in --json/--raw, humanized
// otherwise; stdout is unaffected.
if (outputFile) {
const fileDiff = (jsonMode || rawMode) ? diff : humanizedDiff;
await writeOutputFile(outputFile, JSON.stringify(fileDiff, null, 2), 'utf-8');
process.stderr.write(`\nResults written to ${outputFile}\n`);
}
// Exit code: 0=stable/improving, 1=degrading
if (diff.summary.trend === 'degrading') process.exit(1);
process.exit(0);
process.exitCode = diff.summary.trend === 'degrading' ? 1 : 0;
}
// Only run CLI if invoked directly
@ -141,6 +203,6 @@ const isDirectRun = process.argv[1] && resolve(process.argv[1]) === resolve(new
if (isDirectRun) {
main().catch(err => {
process.stderr.write(`Fatal: ${err.message}\n`);
process.exit(3);
process.exitCode = 3;
});
}

View file

@ -1,15 +1,20 @@
/**
* GAP Scanner Feature Gap Scanner
* Compares actual configuration against complete Claude Code feature register.
* 25 gap dimensions across 4 tiers. Always runs with includeGlobal: true.
* 25 gap dimensions across 4 tiers, plus a conditional disableBundledSkills
* budget-lever check (remediation companion to SKL CA-SKL-002, fires only under
* measured skill-listing pressure). Always runs with includeGlobal: true.
* Finding IDs: CA-GAP-NNN
*/
import { resolve } from 'node:path';
import { resolve, join, sep } from 'node:path';
import { readTextFile, discoverConfigFiles } from './lib/file-discovery.mjs';
import { finding, scannerResult } from './lib/output.mjs';
import { SEVERITY } from './lib/severity.mjs';
import { findImports, parseJson, parseFrontmatter } from './lib/yaml-parser.mjs';
import { measureActiveSkillListing, isBundledSkillsDisabled, BUDGET_CALIBRATION_NOTE } from './lib/skill-listing-budget.mjs';
import { assessMcpDeferralForRepo } from './lib/mcp-deferral.mjs';
import { assessHookContextForRepo } from './lib/hook-additional-context.mjs';
const SCANNER = 'GAP';
@ -41,6 +46,68 @@ function isTargetLocal(ctx, f) {
return f.absPath.startsWith(ctx.targetPath);
}
// Files that are test/demo/vendored config — NOT part of the user's authored
// cascade — must not satisfy "is feature X present?" checks, or they mask real
// gaps. The canonical case: this plugin's own examples/optimal-setup sets
// outputStyle/statusLine/worktree/model/keybindings/.lsp.json, and (because GAP
// always runs includeGlobal) its copies vendored under ~/.claude/plugins/cache
// drive every tier-3 presence check to "present" — hiding the user's real gaps
// on ANY target. Two classes to exclude:
// - plugin-bundled: anything under ~/.claude/plugins/ (absPath marker, mirrors
// the CNF conflict-detector exclusion from M-BUG-2).
// - nested demo/test data: a file whose path RELATIVE TO THE SCAN TARGET sits
// under an examples/ or tests/fixtures/ subtree. relPath (not absPath) is
// deliberate: a fixture scanned AS the target keeps its own files, so the
// frozen v5.0.0 byte-snapshots (scanned from tests/fixtures/marketplace-medium)
// are untouched. (M-BUG-13)
const PLUGIN_TREE_MARKER = `.claude${sep}plugins${sep}`;
/**
* @param {import('./lib/file-discovery.mjs').ConfigFile} file
* @returns {boolean} true if the file is part of the user's authored config
*/
function isAuthoredConfig(file) {
if (file.absPath.includes(PLUGIN_TREE_MARKER)) return false;
const segs = (file.relPath || '').split(sep);
if (segs.includes('examples')) return false;
const ti = segs.indexOf('tests');
if (ti !== -1 && segs[ti + 1] === 'fixtures') return false;
return true;
}
/**
* Read the userprojectlocal settings cascade directly from the filesystem.
* The settings-key gap checks ask "does the USER's resolved config set X?" a
* question the includeGlobal discovery answers unreliably on a real machine: the
* top-level ~/.claude/settings.json is missed (its relPath carries no `.claude`
* segment when the walk root IS ~/.claude) and, when many vendored plugins flood
* the walk, dropped by the discovery file cap. Reading the canonical cascade
* paths directly is immune to both. Merged INTO (not replacing) the discovery
* settings so any non-canonical project settings still count and the frozen
* snapshots stay byte-stable. (M-BUG-13)
* @param {string} targetPath
* @returns {Promise<Array<{ key: string, parsed: object }>>}
*/
async function readSettingsCascade(targetPath) {
const home = process.env.HOME || process.env.USERPROFILE || '';
const paths = [];
if (home) {
paths.push(['user', join(home, '.claude', 'settings.json')]);
paths.push(['user-local', join(home, '.claude', 'settings.local.json')]);
}
paths.push(['project', join(targetPath, '.claude', 'settings.json')]);
paths.push(['local', join(targetPath, '.claude', 'settings.local.json')]);
const out = [];
for (const [scope, p] of paths) {
const content = await readTextFile(p);
if (!content) continue;
const parsed = parseJson(content);
if (parsed && typeof parsed === 'object') out.push({ key: `cascade:${scope}:${p}`, parsed });
}
return out;
}
const TIER_SEVERITY = {
t1: SEVERITY.medium,
t2: SEVERITY.low,
@ -87,6 +154,125 @@ function getSettingsValue(ctx, key) {
return undefined;
}
/**
* Remediation companion to SKL CA-SKL-002: when the active skill listing is over
* its budget and the `disableBundledSkills` lever is un-pulled, recommend it.
*
* Bundled (built-in) skills /code-review, /batch, /debug, /loop, /claude-api
* and more live in the Claude Code binary, not on disk, so their exact listing
* cost cannot be measured here. But they draw on the SAME budget the SKL scanner
* measures; when that budget is already exceeded, dropping them is a zero-cost
* lever that does not touch the user's own skills. We fire ONLY under measured
* pressure (SKL's overflow signal) so this stays an opportunity, not noise.
*
* Pure and exported for unit testing.
*
* @param {{ leverPulled: boolean, aggregate: (import('./lib/skill-listing-budget.mjs').BudgetAssessment|null) }} args
* @returns {object|null} a GAP finding, or null when the lever is pulled or the listing is within budget
*/
export function bundledSkillsLeverFinding({ leverPulled, aggregate }) {
if (leverPulled) return null;
if (!aggregate || !aggregate.overBudget) return null;
return finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: 'Bundled skills add to an over-budget skill listing',
description:
`Your ${aggregate.scanned} active skills already carry ~${aggregate.aggregateTokens} tokens of ` +
`description text, over the ${aggregate.budgetTokens}-token listing budget Claude Code allots the ` +
'skill listing on a 200k context window (~2% of context, CC 2.1.32). Claude Code also loads its ' +
'bundled (built-in) skills — /code-review, /batch, /debug, /loop, /claude-api and more — into that ' +
'same listing. They are not on disk, so their exact cost cannot be measured here, but they draw on ' +
'the same budget. `disableBundledSkills: true` drops them from the listing, reclaiming space without ' +
'touching your own skills.',
evidence:
`description_tokens~${aggregate.aggregateTokens}; budget@200k=${aggregate.budgetTokens} tok; over_by~` +
`${aggregate.overBy} tok; lever=disableBundledSkills (unset) - ${BUDGET_CALIBRATION_NOTE}`,
recommendation:
'Set `disableBundledSkills: true` in settings.json (or the CLAUDE_CODE_DISABLE_BUNDLED_SKILLS env var) ' +
'to hide built-in skills and slash commands from the model and reclaim skill-listing budget (CC 2.1.169+). ' +
'Keep it off if you rely on bundled skills like /code-review — in that case trim your own skill ' +
'descriptions or use `skillOverrides` instead.',
category: 'token-efficiency',
});
}
/**
* CLI-over-MCP lever remediation companion to CA-TOK-006 (v5.10 B4).
*
* Fires ONLY when MCP tool schemas are forced into the always-loaded prefix
* (tool search disabled, or a per-server alwaysLoad), i.e. when MCP is actually
* costing always-loaded tokens. When schemas are deferred (the default), MCP is
* effectively free until used, so there is nothing to recommend and we stay
* silent opportunity, not noise (mirrors the bundledSkills lever's "fire only
* under measured pressure" contract). CLI tools (gh / aws / gcloud) add zero
* context tokens until invoked, so they are the lever the deferral mechanism
* cannot reach for the forced-upfront servers.
*
* Pure and exported for unit testing.
*
* @param {{ assessment: (import('./lib/mcp-deferral.mjs').assessMcpDeferral)|null }} args
* @returns {object|null} a GAP finding, or null when nothing is forced upfront
*/
export function cliOverMcpLeverFinding({ assessment } = {}) {
if (!assessment || !assessment.forcedUpfront) return null;
const names = (assessment.affectedServers || []).map((m) => m.name).join(', ');
return finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: 'Prefer CLI over MCP for common operations',
description:
`Your active project MCP tool schemas (~${assessment.aggregateTokens} tokens) are forced into the ` +
'always-loaded prefix every turn rather than deferred (see CA-TOK-006). CLI tools (gh, aws, gcloud, …) ' +
'add ZERO context tokens until you actually call them, so moving common operations off MCP and onto a ' +
'CLI reclaims always-loaded budget the deferral mechanism cannot.',
evidence:
`forced_schema_tokens~${assessment.aggregateTokens}; servers=${names}; ` +
`reason=${assessment.reason || 'alwaysLoad'} (companion to CA-TOK-006)`,
recommendation:
'For operations a CLI already covers (GitHub → gh, AWS → aws, GCP → gcloud), prefer the CLI over an ' +
'MCP server — CLI output enters context only when invoked. Keep MCP for capabilities with no CLI ' +
'equivalent, disable unused servers via /mcp, and re-enable tool-search deferral so the rest stay names-only.',
category: 'token-efficiency',
});
}
/**
* filter-before-Claude-reads lever remediation companion to HKV's B5 advisory
* (v5.10). Fires ONLY when 1 active hook was detected injecting unfiltered
* command output into additionalContext, i.e. when there is a measured chatty
* hook to fix. When no such hook exists there is nothing to recommend and we
* stay silent opportunity, not noise (same "fire only under measured pressure"
* contract as the cliOverMcp / bundledSkills levers).
*
* Pure and exported for unit testing.
*
* @param {{ flaggedHooks: Array<{event:string, scriptPath:string}> }} args
* @returns {object|null} a GAP finding, or null when no chatty hook was detected
*/
export function filterHookLeverFinding({ flaggedHooks } = {}) {
const hooks = Array.isArray(flaggedHooks) ? flaggedHooks : [];
if (hooks.length === 0) return null;
const scripts = hooks.map((h) => h.scriptPath.split('/').slice(-1)[0]).join(', ');
return finding({
scanner: SCANNER,
severity: SEVERITY.info,
title: 'Filter hook output before it enters context',
description:
`${hooks.length} active hook${hooks.length === 1 ? '' : 's'} build hookSpecificOutput.additionalContext ` +
"from un-grepped command output (see HKV advisory). That field enters Claude's context on every fire, " +
'so filtering verbose output down to what matters BEFORE Claude reads it reclaims per-turn tokens — the ' +
'documented filter-test-output.sh pattern (grep ERROR and return only matches instead of a 10,000-line log).',
evidence: `chatty_hooks=${hooks.length}; scripts=${scripts} (companion to HKV additionalContext advisory)`,
recommendation:
'In each flagged hook, pipe the command output through grep/head/jq to keep only the actionable lines ' +
'before assigning additionalContext. Reserve additionalContext for concise signals; leave bulk diagnostics ' +
'on plain stdout (exit 0) so they go to the debug log, not context.',
category: 'token-efficiency',
});
}
/** @type {GapCheck[]} */
const GAP_CHECKS = [
// --- Tier 1: Foundation ---
@ -355,18 +541,30 @@ export async function scan(targetPath, sharedDiscovery) {
? sharedDiscovery
: await discoverConfigFiles(resolve(targetPath), { includeGlobal: true });
// Parse all settings files upfront
// Presence checks ("does the user have feature X?") must see only the user's
// authored cascade — not bundled/vendored/demo config, which masks real gaps
// (M-BUG-13, see isAuthoredConfig).
const authoredFiles = discovery.files.filter(isAuthoredConfig);
// Parse all settings files upfront (authored discovery files) ...
const parsedSettings = new Map();
for (const file of discovery.files.filter(f => f.type === 'settings-json')) {
for (const file of authoredFiles.filter(f => f.type === 'settings-json')) {
const content = await readTextFile(file.absPath);
if (content) {
const parsed = parseJson(content);
parsedSettings.set(`${file.scope}:${file.relPath}`, parsed);
}
}
// ... plus the real user→project→local cascade read directly, so settings-key
// checks see the true resolved config regardless of the discovery cap/gotcha
// (M-BUG-13). Merged, not replacing — keeps non-canonical project settings and
// the frozen byte-snapshots unchanged.
for (const { key, parsed } of await readSettingsCascade(resolve(targetPath))) {
parsedSettings.set(key, parsed);
}
const ctx = {
files: discovery.files,
files: authoredFiles,
targetPath: resolve(targetPath),
parsedSettings,
fileContents: new Map(),
@ -386,6 +584,27 @@ export async function scan(targetPath, sharedDiscovery) {
}
}
// disableBundledSkills lever — fires only under measured skill-listing pressure
// (SKL's CA-SKL-002 overflow signal). HOME-scoped: the listing and the lever
// cascade both resolve via process.env.HOME, independent of project discovery.
const leverPulled = await isBundledSkillsDisabled(ctx.targetPath);
const { aggregate } = await measureActiveSkillListing();
const leverFinding = bundledSkillsLeverFinding({ leverPulled, aggregate });
if (leverFinding) findings.push(leverFinding);
// CLI-over-MCP lever — companion to CA-TOK-006: fires only when project-local
// MCP tool schemas are forced into the always-loaded prefix (tool search
// disabled or a per-server alwaysLoad). Reuses the same static assessment.
const mcpAssessment = await assessMcpDeferralForRepo(ctx.targetPath);
const cliLever = cliOverMcpLeverFinding({ assessment: mcpAssessment });
if (cliLever) findings.push(cliLever);
// filter-before-Claude-reads lever — companion to HKV's B5 advisory: fires
// only when an active hook injects unfiltered output into additionalContext.
const flaggedHooks = await assessHookContextForRepo(discovery);
const hookLever = filterHookLeverFinding({ flaggedHooks });
if (hookLever) findings.push(hookLever);
const filesScanned = discovery.files.length;
return scannerResult(SCANNER, 'ok', findings, filesScanned, Date.now() - start);
}

View file

@ -9,11 +9,18 @@
*/
import { resolve } from 'node:path';
import { writeOutputFile } from './lib/write-output.mjs';
import { runAllScanners } from './scan-orchestrator.mjs';
import { planFixes, applyFixes, verifyFixes } from './fix-engine.mjs';
import { createBackup } from './lib/backup.mjs';
import { humanizeFinding } from './lib/humanizer.mjs';
// `--dry-run` is a no-op alias: dry-run is already the default. It exists because
// commands/fix.md documents it in argument-hint, and a documented flag that the
// CLI silently drops is the same fail-silent class as the unknown-flag sink below.
const BOOL_FLAGS = ['--apply', '--dry-run', '--json', '--raw', '--global'];
const VALUE_FLAGS = ['--output-file'];
async function main() {
const args = process.argv.slice(2);
let targetPath = '.';
@ -21,18 +28,36 @@ async function main() {
let jsonMode = false;
let rawMode = false;
let includeGlobal = false;
let outputFile = null;
// Same defect class as M-BUG-21 in drift-cli: this loop used to end in
// `else if (!args[i].startsWith('-')) targetPath = args[i]` with no
// unknown-flag branch, so an unrecognised flag was dropped silently and its
// VALUE became the scan target. Here that is worse than in drift: combined
// with --apply it silently moves the WRITE target to another tree.
for (let i = 0; i < args.length; i++) {
if (args[i] === '--apply') {
apply = true;
} else if (args[i] === '--json') {
jsonMode = true;
} else if (args[i] === '--raw') {
rawMode = true;
} else if (args[i] === '--global') {
includeGlobal = true;
} else if (!args[i].startsWith('-')) {
targetPath = args[i];
const arg = args[i];
if (BOOL_FLAGS.includes(arg)) {
if (arg === '--apply') apply = true;
else if (arg === '--json') jsonMode = true;
else if (arg === '--raw') rawMode = true;
else if (arg === '--global') includeGlobal = true;
// --dry-run: default behaviour, accepted so it is not silently dropped.
} else if (VALUE_FLAGS.includes(arg)) {
const value = args[i + 1];
if (value === undefined || value.startsWith('-')) {
throw new Error(`Option ${arg} requires a value.`);
}
outputFile = value;
i++;
} else if (arg.startsWith('-')) {
throw new Error(
`Unknown option: ${arg}\n` +
`Valid options: ${[...BOOL_FLAGS, ...VALUE_FLAGS].join(' ')}`
);
} else {
targetPath = arg;
}
}
@ -105,16 +130,21 @@ async function main() {
let backupId = null;
if (fixes.length === 0) {
const output = { planned: [], applied: [], failed: [], verified: [], regressions: [], manual, backupId: null };
if (machineMode) {
const output = { planned: [], applied: [], failed: [], verified: [], regressions: [], manual, backupId: null };
process.stdout.write(JSON.stringify(output, null, 2) + '\n');
}
process.exit(0);
if (outputFile) await writeOutputFile(outputFile, JSON.stringify(output, null, 2) + '\n', 'utf-8');
return;
}
if (apply) {
// Create backup first
const filesToBackup = [...new Set(fixes.filter(f => f.type !== 'file-rename').map(f => f.file))];
// Create backup first. file-rename used to be excluded here, so a rule file
// whose only defect was its extension was renamed with NO backup entry —
// while commands/fix.md promised "every fix creates a backup first" and
// handed the user a backupId that could not restore it. The source file is
// backed up like any other; rollback recreates it at its original path.
const filesToBackup = [...new Set(fixes.map(f => f.file))];
const backup = createBackup(filesToBackup);
backupId = backup.backupId;
@ -142,7 +172,10 @@ async function main() {
process.stderr.write(`\n Verifying...\n`);
}
const verification = await verifyFixes(envelope, applied);
// Verification must re-scan the scope the fix run used. It hardcoded
// includeGlobal:false, so with --global every untouched global-scope
// finding fell out of the re-scan and was reported as verified.
const verification = await verifyFixes(envelope, applied, { includeGlobal });
verified = verification.verified;
regressions = verification.regressions;
@ -151,7 +184,11 @@ async function main() {
if (regressions.length > 0) {
process.stderr.write(` Regressions: ${regressions.join(', ')}\n`);
}
process.stderr.write(`\n Rollback: node scanners/rollback-cli.mjs ${backupId}\n`);
// There is no rollback-cli.mjs — the restore path is the command, which
// drives rollback-engine.mjs. Pointing at a nonexistent script in the
// one message a user reaches for after a bad fix is the worst place for
// a dead reference.
process.stderr.write(`\n Rollback: /config-audit rollback ${backupId}\n`);
}
}
} else {
@ -165,7 +202,7 @@ async function main() {
}
// JSON output (both --json and --raw write byte-equal v5.0.0-shape stdout)
if (machineMode) {
{
const output = {
planned: fixes.map(f => ({
findingId: f.findingId,
@ -193,7 +230,17 @@ async function main() {
})),
backupId,
};
process.stdout.write(JSON.stringify(output, null, 2) + '\n');
const serialized = JSON.stringify(output, null, 2) + '\n';
if (machineMode) process.stdout.write(serialized);
// --output-file carries the same payload to disk. ux-rules rule 2 requires
// it: commands run scanners with `2>/dev/null`, so anything the command has
// to act on must ride in a file, not in stdout or stderr.
if (outputFile) await writeOutputFile(outputFile, serialized, 'utf-8');
// Exit code follows the convention the other scanners use: 0 PASS,
// 2 FAIL, 3 tool error. A failed fix used to exit 0, so a caller could not
// tell a clean run from one that silently lost a fix.
if (failed.length > 0) process.exitCode = 2;
}
}
@ -202,6 +249,6 @@ const isDirectRun = process.argv[1] && resolve(process.argv[1]) === resolve(new
if (isDirectRun) {
main().catch(err => {
process.stderr.write(`Fatal: ${err.message}\n`);
process.exit(3);
process.exitCode = 3;
});
}

View file

@ -56,9 +56,21 @@ export function planFixes(envelope) {
}
}
// Sort fixes by severity weight (critical first)
// Sort fixes by severity weight (critical first), but a file-rename always
// sorts after every other fix. A rename moves the file out from under any
// later fix that still addresses the old path: a rule file with both
// `globs:` and a non-.md extension had the rename applied first, and the
// frontmatter fix then failed with ENOENT while the run still exited 0.
// `?? 4`, not `|| 4`: critical weighs 0, and `0 || 4` evaluates to 4 — so
// critical fixes sorted LAST, the exact opposite of this function's contract
// (M-BUG-30). The old test used the same falsy fallback and agreed with the bug.
const severityOrder = { critical: 0, high: 1, medium: 2, low: 3, info: 4 };
fixes.sort((a, b) => (severityOrder[a.severity] || 4) - (severityOrder[b.severity] || 4));
fixes.sort((a, b) => {
const aRename = a.type === FIX_TYPES.FILE_RENAME ? 1 : 0;
const bRename = b.type === FIX_TYPES.FILE_RENAME ? 1 : 0;
if (aRename !== bRename) return aRename - bRename;
return (severityOrder[a.severity] ?? 4) - (severityOrder[b.severity] ?? 4);
});
return { fixes, skipped, manual };
}
@ -180,7 +192,7 @@ function createFixPlan(finding) {
// --- RUL scanner fixes ---
if (scanner === 'RUL') {
if (title === 'Rule uses deprecated "globs" field') {
if (title === 'Rule uses "globs" instead of documented "paths"') {
return {
...base,
type: FIX_TYPES.FRONTMATTER_RENAME,
@ -600,16 +612,21 @@ function extractEventFromDescription(description) {
* Verify fixes by re-running affected scanners.
* @param {object} originalEnvelope - Original scanner envelope
* @param {object[]} appliedResults - Results from applyFixes()
* @param {object} [opts]
* @param {boolean} [opts.includeGlobal=false] - Must match the scope the fix run scanned
* @returns {Promise<{ verified: string[], regressions: string[], newFindings: object[] }>}
*/
export async function verifyFixes(originalEnvelope, appliedResults) {
export async function verifyFixes(originalEnvelope, appliedResults, opts = {}) {
const targetPath = originalEnvelope.meta.target;
const verified = [];
const regressions = [];
const newFindings = [];
// Re-scan the target
const newEnvelope = await runAllScanners(targetPath, { includeGlobal: false });
// Re-scan the target in the SAME scope the fix run used. This was hardcoded
// to includeGlobal:false: after a --global run, every global-scope finding
// was absent from the re-scan and therefore counted as verified — a clean
// "fixed" report for files nothing had touched.
const newEnvelope = await runAllScanners(targetPath, { includeGlobal: opts.includeGlobal === true });
// Build set of original finding IDs that were fixed
const fixedIds = new Set(

View file

@ -8,16 +8,18 @@ import { readTextFile, discoverConfigFiles } from './lib/file-discovery.mjs';
import { finding, scannerResult } from './lib/output.mjs';
import { SEVERITY } from './lib/severity.mjs';
import { parseJson } from './lib/yaml-parser.mjs';
import { assessHookAdditionalContext } from './lib/hook-additional-context.mjs';
import { stat } from 'node:fs/promises';
import { resolve, dirname } from 'node:path';
const SCANNER = 'HKV';
/** All valid hook events (as of April 2026) */
/** All valid hook events — verified against code.claude.com/docs/en/hooks.md (2026-06-19) */
const VALID_EVENTS = new Set([
'SessionStart', 'InstructionsLoaded', 'UserPromptSubmit',
'SessionStart', 'Setup', 'InstructionsLoaded',
'UserPromptSubmit', 'UserPromptExpansion',
'PreToolUse', 'PermissionRequest', 'PermissionDenied',
'PostToolUse', 'PostToolUseFailure',
'PostToolUse', 'PostToolUseFailure', 'PostToolBatch',
'SubagentStart', 'SubagentStop',
'TaskCreated', 'TaskCompleted',
'Stop', 'StopFailure',
@ -26,7 +28,11 @@ const VALID_EVENTS = new Set([
'WorktreeCreate', 'WorktreeRemove',
'PreCompact', 'PostCompact',
'Elicitation', 'ElicitationResult',
'SessionEnd',
'SessionEnd', 'MessageDisplay',
// 'post-session' deliberately EXCLUDED: the 2.1.169 changelog `post-session`
// is a self-hosted-runner workspace-lifecycle hook, NOT a settings.json hook
// event (absent from hooks.md; all settings.json events are PascalCase).
// Verified 2026-06-20.
]);
/** Valid hook handler types */
@ -134,7 +140,7 @@ async function validateHooksObject(hooks, file, findings, baseDir) {
description: `${file.relPath}: "${event}" is not a valid hook event. This hook will never fire.`,
file: file.absPath,
evidence: event,
recommendation: `Valid events: ${[...VALID_EVENTS].slice(0, 8).join(', ')}... (26 total)`,
recommendation: `Valid events: ${[...VALID_EVENTS].slice(0, 8).join(', ')}... (${VALID_EVENTS.size} total)`,
autoFixable: false,
}));
continue;
@ -243,6 +249,35 @@ async function validateHooksObject(hooks, file, findings, baseDir) {
autoFixable: false,
}));
}
// v5.10 B5: advisory (info) — a hook that injects unfiltered
// command output into hookSpecificOutput.additionalContext pays
// that whole payload into Claude's context on every fire (plain
// stdout does not). Low-precision static heuristic, so info only.
const scriptContent = await readTextFile(scriptPath);
const ac = assessHookAdditionalContext({ scriptContent });
if (ac.flagged) {
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.info,
title: 'Hook injects unfiltered output into context',
description:
`${file.relPath}: "${event}" runs ${scriptPath.split('/').slice(-2).join('/')} ` +
'which builds hookSpecificOutput.additionalContext from un-grepped command ' +
"output. That field enters Claude's context every time the hook fires (plain " +
'stdout does not), so an unfiltered payload is a recurring per-turn token cost. ' +
'Advisory only — low-precision static heuristic; verify the real payload size.',
file: scriptPath,
evidence:
'additional_context_unfiltered=true; ' +
`verbose_capture=${ac.hasVerboseCapture}; filter_applied=${ac.hasFilter}`,
recommendation:
'Filter before Claude reads: grep/head the command output down to what matters ' +
'before putting it in additionalContext (the documented filter-test-output.sh ' +
'pattern), or keep large diagnostics on plain stdout so they stay out of context.',
autoFixable: false,
}));
}
}
}
}

View file

@ -0,0 +1,115 @@
#!/usr/bin/env node
/**
* knowledge-refresh CLI feeds the v5.7 `/config-audit knowledge-refresh` command
* (Chunk 3: the "living" half of the living knowledge base).
*
* This is the DETERMINISTIC half of the hybrid motor: it loads the best-practices
* register and classifies every entry as `fresh` or `stale` by the age of its
* `source.verified` stamp (via the pure `assessFreshness` core). It is READ-ONLY
* it NEVER writes the register and NEVER touches the network. Candidate discovery
* (polling the CC changelog + Anthropic blog) and the human-approved writes live in
* the command layer (Verifiseringsplikt). `--dry-run` is implicit and the only mode;
* the flag is accepted for explicitness and echoed back.
*
* Naming: `-cli` suffix NOT an orchestrated scanner (the scan-orchestrator only
* loads scanner modules), so the scanner count is unchanged and the snapshot suite
* stays byte-stable.
*
* Usage:
* node knowledge-refresh-cli.mjs [--output-file <path>] [--stale-after <N>]
* [--reference-date <YYYY-MM-DD>] [--dry-run]
*
* Exit codes: 0 = every entry fresh, 1 = one or more stale (advisory), 3 = error.
*/
import { resolve } from 'node:path';
import { writeOutputFile } from './lib/write-output.mjs';
import { loadRegister, REGISTER_PATH } from './lib/best-practices-register.mjs';
import { assessFreshness, STALE_AFTER_DAYS_DEFAULT } from './lib/knowledge-refresh.mjs';
const DATE_RE = /^\d{4}-\d{2}-\d{2}$/;
/**
* Usage error. Throws rather than calling process.exit(): exit() discards
* unflushed stdout when stdout is a pipe. The top-level catch prints the same
* `Error: ` text and sets the same exit code 3, so callers see no difference.
*/
class CliUsageError extends Error {}
function fail(message) {
throw new CliUsageError(message);
}
async function main() {
const args = process.argv.slice(2);
let outputFile = null;
let staleAfterDays = STALE_AFTER_DAYS_DEFAULT;
let referenceDate = null; // null → today
let dryRun = false;
for (let i = 0; i < args.length; i++) {
const a = args[i];
if (a === '--dry-run') dryRun = true;
else if (a === '--output-file' && args[i + 1]) outputFile = args[++i];
else if (a === '--stale-after' && args[i + 1] !== undefined) {
const n = Number.parseInt(args[++i], 10);
if (!Number.isInteger(n) || n < 0) fail('--stale-after must be a non-negative integer (days)');
staleAfterDays = n;
} else if (a === '--reference-date' && args[i + 1]) {
referenceDate = args[++i];
if (!DATE_RE.test(referenceDate)) fail('--reference-date must be YYYY-MM-DD');
}
// A flag we do not understand must fail loudly. Silently dropping it is how
// `--stale-after 30` — arriving as ONE argv entry from a shell that does not
// word-split — became "all 14 entries are fresh within 90 days": a confident
// answer to a question the caller did not ask.
else if (a.startsWith('--')) fail(`unknown flag "${a}"`);
}
// The clock is read here ONLY — the core takes an injected date and stays pure.
const ref = referenceDate || new Date();
let register;
try {
register = loadRegister();
} catch (err) {
fail(`could not load register at ${REGISTER_PATH}: ${err.message}`);
}
let assessment;
try {
assessment = assessFreshness(register, { referenceDate: ref, staleAfterDays });
} catch (err) {
fail(err.message);
}
const payload = {
status: 'ok',
registerPath: REGISTER_PATH,
version: register.version,
dryRun: true, // this CLI never writes; the flag is informational
requestedDryRun: dryRun,
referenceDate: assessment.referenceDate,
staleAfterDays: assessment.staleAfterDays,
counts: assessment.counts,
stale: assessment.stale,
fresh: assessment.fresh,
};
const json = JSON.stringify(payload, null, 2);
if (outputFile) await writeOutputFile(outputFile, json, 'utf-8');
else process.stdout.write(json + '\n');
process.exitCode = assessment.counts.stale > 0 ? 1 : 0;
}
const isDirectRun =
process.argv[1] && resolve(process.argv[1]) === resolve(new URL(import.meta.url).pathname);
if (isDirectRun) {
main().catch((err) => {
const prefix = err instanceof CliUsageError ? 'Error' : 'Fatal';
process.stderr.write(`${prefix}: ${err.message}\n`);
process.exitCode = 3;
});
}

View file

@ -53,6 +53,134 @@ export function estimateTokens(bytes, kind = 'markdown', opts = {}) {
return Math.ceil(bytes / 4);
}
/**
* Strip block-level HTML comments (`<!-- ... -->`) that lie OUTSIDE fenced code
* blocks. Claude Code strips these before injecting a CLAUDE.md / memory file
* into context (code.claude.com/docs/en/memory: "block-level HTML comments are
* stripped before the content is injected"), preserving them only inside fenced
* code blocks (``` / ~~~). A byte-accurate token estimate must therefore discount
* them. (M-BUG-6)
*
* Conservative scope only *block-level* comments are removed (a comment that
* occupies its own line(s)); inline comments sharing a line with other text are
* retained, since the verified CC behavior covers block-level stripping only.
*
* @param {string} content
* @returns {string} content with out-of-fence block comments removed
*/
export function stripInjectedHtmlComments(content) {
if (typeof content !== 'string' || content === '') return '';
const lines = content.split('\n');
const out = [];
let inFence = false;
let inComment = false;
for (const line of lines) {
if (inComment) {
// Inside a multi-line block comment: drop lines until the closing `-->`,
// keeping any real content that trails the close on the same line.
const end = line.indexOf('-->');
if (end !== -1) {
inComment = false;
const rest = line.slice(end + 3);
if (rest.trim() !== '') out.push(rest);
}
continue;
}
// Fence delimiters (``` / ~~~) toggle a preserve-verbatim region.
if (/^\s*(```|~~~)/.test(line)) {
inFence = !inFence;
out.push(line);
continue;
}
if (inFence) {
out.push(line);
continue;
}
// Whole line is a single self-contained block comment → CC strips it.
if (/^\s*<!--[\s\S]*?-->\s*$/.test(line)) continue;
// Block comment opening with nothing but whitespace before it and no close
// on this line → runs onto following lines.
const openIdx = line.indexOf('<!--');
if (openIdx !== -1 && line.indexOf('-->', openIdx) === -1 && line.slice(0, openIdx).trim() === '') {
inComment = true;
continue;
}
out.push(line);
}
return out.join('\n');
}
/**
* Effective injected byte length of a CLAUDE.md / memory source: raw UTF-8 bytes
* minus the block-level HTML comments CC strips before injection. Used wherever a
* CLAUDE.md token estimate must reflect what actually enters context. (M-BUG-6)
*
* @param {string} content
* @returns {number}
*/
export function effectiveMemoryBytes(content) {
if (typeof content !== 'string') return 0;
return Buffer.byteLength(stripInjectedHtmlComments(content), 'utf8');
}
// ─────────────────────────────────────────────────────────────────────────
// Load-pattern model (v5.6 Foundation)
// ─────────────────────────────────────────────────────────────────────────
/**
* Derive how a config source loads into context and whether it survives a
* `/compact`, from the published Claude Code loading model. Deterministic,
* side-effect-free. `derivationConfidence` is 'confirmed' when a primary-doc
* row nails the row (V-rows in docs/v5.5-steering-model-plan.md), 'inferred'
* when reasoned from an analogue (so a renderer can choose to mark it).
*
* loadPattern { 'always', 'on-demand', 'external', 'unknown' }
* survivesCompaction { 'yes', 'no', 'n/a' }
*
* @param {string} kind - source kind (see switch)
* @param {{scoped?: boolean}} [opts] - kind-specific discriminators
* @returns {{loadPattern:string, survivesCompaction:string, derivationConfidence:string}}
*/
export function deriveLoadPattern(kind, opts = {}) {
const mk = (loadPattern, survivesCompaction, derivationConfidence) =>
({ loadPattern, survivesCompaction, derivationConfidence });
switch (kind) {
// CLAUDE.md cascade
case 'claude-md-root': return mk('always', 'yes', 'confirmed'); // V1
case 'claude-md-nested': return mk('on-demand', 'no', 'confirmed'); // V3
case 'claude-md-user':
case 'claude-md-managed':
case 'claude-md-import': return mk('always', 'yes', 'inferred');
// Rules
case 'rule':
return opts.scoped
? mk('on-demand', 'no', 'confirmed') // V2, V4 (loads on Read of a match)
: mk('always', 'yes', 'confirmed'); // V1, V6 (unscoped = always-on)
// Skills
case 'skill-listing': return mk('always', 'n/a', 'confirmed'); // V7 (name+desc every turn)
case 'skill-body': return mk('on-demand', 'n/a', 'confirmed'); // V7 (body on invoke)
// Agents — name+description load for delegation each turn (skill analogue;
// no primary-doc row pins it, so 'inferred').
case 'agent': return mk('always', 'n/a', 'inferred');
// Output styles modify the system prompt, re-sent every turn (V10, V12).
case 'output-style': return mk('always', 'yes', 'confirmed');
// Hooks run outside context (V18); the hook itself is external.
case 'hook': return mk('external', 'n/a', 'confirmed');
// MCP tool schemas are part of the per-turn payload (no explicit
// compaction-survival row → 'inferred').
case 'mcp': return mk('always', 'yes', 'inferred');
// Slash-command body loads when the command is invoked (on-demand). No
// primary-doc row pins the always-loaded command listing cost → 'inferred'.
case 'command': return mk('on-demand', 'n/a', 'inferred');
// Harness-config files (settings.json, keybindings.json, .mcp.json, hooks.json,
// plugin.json, ~/.claude.json) are read by the CLI to configure the harness —
// they are NOT injected into the model context, so they cost no per-turn
// context tokens. 'external' = outside the context window (like hooks).
case 'harness-config': return mk('external', 'n/a', 'inferred');
default: return mk('unknown', 'n/a', 'inferred');
}
}
// ─────────────────────────────────────────────────────────────────────────
// Git root detection
// ─────────────────────────────────────────────────────────────────────────
@ -144,7 +272,11 @@ export async function walkClaudeMdCascade(repoPath) {
const totalBytes = files.reduce((sum, f) => sum + f.bytes, 0);
const totalLines = files.reduce((sum, f) => sum + f.lines, 0);
const estimatedTokens = estimateTokens(totalBytes, 'markdown');
// Token estimate is computed from the *effective* (injected) byte count — CC
// strips block-level HTML comments before injection — while totalBytes stays
// the honest on-disk figure. (M-BUG-6)
const effectiveBytes = files.reduce((sum, f) => sum + (f.effectiveBytes ?? f.bytes), 0);
const estimatedTokens = estimateTokens(effectiveBytes, 'markdown');
return { files, totalBytes, totalLines, estimatedTokens };
}
@ -159,6 +291,7 @@ async function tryAddClaudeMd(absPath, scope, parent, files, seen) {
path: absPath,
scope,
bytes: s.size,
effectiveBytes: effectiveMemoryBytes(content),
lines: lineCount(content),
parent,
};
@ -269,19 +402,120 @@ export async function readClaudeJsonProjectSlice(repoPath) {
// ─────────────────────────────────────────────────────────────────────────
/**
* Enumerate all plugins installed under ~/.claude/plugins/marketplaces.
* For each plugin: counts commands, agents, skills, hooks, rules; reads version from plugin.json.
* Enumerate the plugins Claude Code actually injects for a repo.
*
* Authoritative source is `~/.claude/plugins/installed_plugins.json` (the install
* manifest) gated by the `enabledPlugins` toggle map. Only plugins that are both
* installed AND `enabledPlugins[key] === true` are injected, so only those are
* counted each resolved to its ACTIVE `installPath`, which for polyrepo plugins
* lives under `plugins/cache` (never under `plugins/marketplaces`, so the historic
* marketplaces walk missed them entirely while also counting disabled/uninstalled
* marketplaces plugins). Mirrors file-discovery.mjs's "trust installed_plugins.json"
* contract: when the manifest is absent (test fixtures, pre-v2 installs) we cannot
* tell enabled from installed, so we fall back to discovering everything under
* `plugins/marketplaces` rather than silently dropping config. (M-BUG-1)
*
* @param {string} [repoPath] - when given, project/local-scoped installs and
* project-level `enabledPlugins` overrides are resolved relative to it; omit for
* HOME/global scope (only user-scope installs + user `enabledPlugins`).
* @returns {Promise<Array<{name:string, path:string, version:string|null, commands:number, agents:number, skills:number, hooks:number, rules:number, totalBytes:number, estimatedTokens:number}>>}
*/
export async function enumeratePlugins() {
export async function enumeratePlugins(repoPath) {
const home = process.env.HOME || process.env.USERPROFILE || '';
if (!home) return [];
const marketplacesRoot = join(home, '.claude', 'plugins', 'marketplaces');
const pluginRoots = await discoverAllPluginsUnder(marketplacesRoot);
const installed = await readInstalledPluginsManifest(home);
// Dedupe via realpath (symlinks are common)
let pluginRoots;
if (installed) {
// Manifest present → inject only ENABLED plugins, from their active installPath.
const enabled = await readEnabledPluginsMap(home, repoPath);
pluginRoots = [];
for (const [key, recs] of Object.entries(installed)) {
if (enabled[key] !== true) continue; // not explicitly enabled → not injected
const rec = pickActivePluginRecord(recs, repoPath);
if (!rec || !rec.installPath) continue;
try {
await stat(rec.installPath); // skip enabled-but-missing installPaths
pluginRoots.push(rec.installPath);
} catch { /* installPath gone → not loadable */ }
}
} else {
// No manifest → cannot tell enabled from installed → discover all on disk.
pluginRoots = await discoverAllPluginsUnder(join(home, '.claude', 'plugins', 'marketplaces'));
}
return buildPluginRecords(pluginRoots);
}
/**
* Read the install manifest's `plugins` map ({ "name@marketplace": [record, ] }).
* Returns null when absent/unparseable so callers fall back to disk discovery.
*/
async function readInstalledPluginsManifest(home) {
const p = join(home, '.claude', 'plugins', 'installed_plugins.json');
let raw;
try { raw = await readFile(p, 'utf-8'); } catch { return null; }
const parsed = parseJson(raw);
if (!parsed || !parsed.plugins || typeof parsed.plugins !== 'object') return null;
return parsed.plugins;
}
/**
* Merge the `enabledPlugins` toggle map across the scopes Claude Code reads:
* user settings.json, then (when repoPath given) project settings + local + the
* ~/.claude.json project slice. Later scopes override earlier ones.
*/
async function readEnabledPluginsMap(home, repoPath) {
const merged = {};
const sources = [join(home, '.claude', 'settings.json')];
if (repoPath) {
sources.push(join(repoPath, '.claude', 'settings.json'));
sources.push(join(repoPath, '.claude', 'settings.local.json'));
}
for (const s of sources) {
try {
const parsed = parseJson(await readFile(s, 'utf-8'));
if (parsed && parsed.enabledPlugins && typeof parsed.enabledPlugins === 'object') {
Object.assign(merged, parsed.enabledPlugins);
}
} catch { /* missing/unreadable scope */ }
}
if (repoPath) {
try {
const slice = await readClaudeJsonProjectSlice(repoPath);
if (slice && slice.enabledPlugins && typeof slice.enabledPlugins === 'object') {
Object.assign(merged, slice.enabledPlugins);
}
} catch { /* ignore */ }
}
return merged;
}
/**
* Pick the applicable install record for a plugin. User-scope records apply
* everywhere; project/local-scope records only when repoPath is within their
* projectPath (so a project-scoped plugin never leaks into HOME/global scope).
*/
function pickActivePluginRecord(recs, repoPath) {
if (!Array.isArray(recs) || recs.length === 0) return null;
const applicable = recs.filter((r) => {
if (!r || !r.installPath) return false;
const scope = r.scope || 'user';
if (scope === 'user') return true;
if (!repoPath || !r.projectPath) return false;
const target = normalizePath(resolve(repoPath));
const pp = normalizePath(resolve(r.projectPath));
return target === pp || target.startsWith(pp + sep);
});
return applicable.find((r) => (r.scope || 'user') === 'user') || applicable[0] || null;
}
/**
* Build plugin records from a list of plugin root paths: dedupe via realpath,
* count items, read plugin.json name/version.
*/
async function buildPluginRecords(pluginRoots) {
const seen = new Set();
const results = [];
for (const root of pluginRoots) {
@ -413,14 +647,20 @@ async function countPluginItems(pluginRoot) {
return counts;
}
async function listMarkdownFiles(dir) {
async function listMarkdownFiles(dir, recursive = false) {
const out = [];
let entries;
try { entries = await readdir(dir, { withFileTypes: true }); } catch { return out; }
for (const e of entries) {
const full = join(dir, e.name);
if (e.isDirectory()) {
// Opt-in recursion (M-BUG-3): CC scans agents dirs recursively, so agents
// organized into subfolders must be enumerated too. Other callers stay flat.
if (recursive) out.push(...await listMarkdownFiles(full, true));
continue;
}
if (!e.isFile()) continue;
if (!e.name.endsWith('.md')) continue;
const full = join(dir, e.name);
try {
const s = await stat(full);
out.push({ path: full, size: s.size });
@ -499,6 +739,148 @@ export async function enumerateSkills(pluginList = []) {
return out;
}
// ─────────────────────────────────────────────────────────────────────────
// Rules, agents, output styles (v5.6 Foundation enumeration)
// ─────────────────────────────────────────────────────────────────────────
/** True when `v` is a non-empty, non-whitespace string (a usable frontmatter field). */
function hasText(v) {
return typeof v === 'string' && v.trim().length > 0;
}
/**
* Build the project/user/plugin directory list for a per-kind enumerator.
* Project + user dirs live under `.claude/<dir>`; plugins under each of the
* given subpaths relative to the plugin root.
*/
function configDirs(repoPath, pluginList, subdir, pluginSubdirs = [subdir]) {
const home = process.env.HOME || process.env.USERPROFILE || '';
const projectDir = join(repoPath, '.claude', subdir);
const userDir = home ? join(home, '.claude', subdir) : null;
const dirs = [];
// M-BUG-4: when repoPath === $HOME (the `manifest --global` self-scan), the
// project dir resolves to the same path as the user dir. Count it once, as
// user scope, instead of enumerating the same directory twice.
if (!(userDir && userDir === projectDir)) {
dirs.push({ dir: projectDir, source: 'project', pluginName: null });
}
if (userDir) dirs.push({ dir: userDir, source: 'user', pluginName: null });
for (const p of pluginList) {
for (const sub of pluginSubdirs) {
dirs.push({ dir: join(p.path, sub), source: 'plugin', pluginName: p.name });
}
}
return dirs;
}
/**
* Enumerate rule files: `<repo>/.claude/rules/`, `~/.claude/rules/`, and each
* plugin's `rules/` + `.claude/rules/`. A rule is path-scoped when its
* frontmatter declares `paths:` (the only documented scoping field, V5) which
* determines its load pattern (scoped = on-demand, unscoped = always, V1/V2/V6).
*
* @param {string} repoPath
* @param {Array<{name:string, path:string}>} [pluginList]
* @returns {Promise<Array<{name:string, source:string, pluginName:string|null, path:string, scoped:boolean, bytes:number, estimatedTokens:number, loadPattern:string, survivesCompaction:string, derivationConfidence:string}>>}
*/
export async function enumerateRules(repoPath, pluginList = []) {
const out = [];
const dirs = configDirs(repoPath, pluginList, 'rules', ['rules', join('.claude', 'rules')]);
for (const { dir, source, pluginName } of dirs) {
const files = await listMarkdownFiles(dir);
for (const f of files) {
let scoped = false;
try {
const content = await readFile(f.path, 'utf-8');
const { frontmatter } = parseFrontmatter(content);
scoped = !!(frontmatter && frontmatter.paths);
} catch { /* unreadable → treat as unscoped */ }
out.push({
name: basename(f.path),
source,
pluginName,
path: f.path,
scoped,
bytes: f.size,
estimatedTokens: estimateTokens(f.size, 'markdown'),
...deriveLoadPattern('rule', { scoped }),
});
}
}
return out;
}
/**
* Enumerate agent definitions: `<repo>/.claude/agents/`, `~/.claude/agents/`,
* and each plugin's `agents/`. Only name+description load for delegation each
* turn, so cost is estimated like other frontmatter-only sources.
*
* @param {string} repoPath
* @param {Array<{name:string, path:string}>} [pluginList]
* @returns {Promise<Array<{name:string, source:string, pluginName:string|null, path:string, bytes:number, estimatedTokens:number, loadPattern:string, survivesCompaction:string, derivationConfidence:string}>>}
*/
export async function enumerateAgents(repoPath, pluginList = []) {
const out = [];
const lp = deriveLoadPattern('agent');
const dirs = configDirs(repoPath, pluginList, 'agents');
for (const { dir, source, pluginName } of dirs) {
const files = await listMarkdownFiles(dir, true); // M-BUG-3: CC scans agents dirs recursively
for (const f of files) {
// M-BUG-5: CC registers a subagent only when its frontmatter declares both
// `name` and `description` (docs: identity comes only from `name`; both are
// required). Frontmatter-less / incomplete files are registration no-ops
// that cost zero always-loaded tokens — don't count them as agents.
let frontmatter;
try {
({ frontmatter } = parseFrontmatter(await readFile(f.path, 'utf-8')));
} catch { continue; }
if (!hasText(frontmatter && frontmatter.name) || !hasText(frontmatter && frontmatter.description)) continue;
out.push({
name: basename(f.path).replace(/\.md$/, ''),
source,
pluginName,
path: f.path,
bytes: f.size,
estimatedTokens: estimateTokens(f.size, 'frontmatter'),
...lp,
});
}
}
return out;
}
/**
* Enumerate output styles: `<repo>/.claude/output-styles/`,
* `~/.claude/output-styles/`, and each plugin's `output-styles/`. An output
* style modifies the system prompt and is re-sent every turn (V10, V12).
* Foundation only enumerates them; the `keep-coding-instructions` /
* `force-for-plugin` checks are the CA-OST scanner (v5.6 C).
*
* @param {string} repoPath
* @param {Array<{name:string, path:string}>} [pluginList]
* @returns {Promise<Array<{name:string, source:string, pluginName:string|null, path:string, bytes:number, estimatedTokens:number, loadPattern:string, survivesCompaction:string, derivationConfidence:string}>>}
*/
export async function enumerateOutputStyles(repoPath, pluginList = []) {
const out = [];
const lp = deriveLoadPattern('output-style');
const dirs = configDirs(repoPath, pluginList, 'output-styles');
for (const { dir, source, pluginName } of dirs) {
const files = await listMarkdownFiles(dir);
for (const f of files) {
out.push({
name: basename(f.path).replace(/\.md$/, ''),
source,
pluginName,
path: f.path,
bytes: f.size,
estimatedTokens: estimateTokens(f.size, 'markdown'),
...lp,
});
}
}
return out;
}
// ─────────────────────────────────────────────────────────────────────────
// Hooks (user + project + plugin)
// ─────────────────────────────────────────────────────────────────────────
@ -610,6 +992,7 @@ export async function readActiveMcpServers(repoPath, claudeJsonSlice = null, plu
toolCount,
toolCountUnknown: detected.toolCountUnknown,
estimatedTokens: estimateTokens(0, 'mcp', { toolCount: toolCount ?? 0 }),
alwaysLoad: def?.alwaysLoad === true,
});
}
@ -639,6 +1022,7 @@ async function collectMcpFromFile(path, source, disabled, out, repoPath) {
toolCount,
toolCountUnknown: detected.toolCountUnknown,
estimatedTokens: estimateTokens(0, 'mcp', { toolCount: toolCount ?? 0 }),
alwaysLoad: def?.alwaysLoad === true,
});
}
}
@ -837,15 +1221,18 @@ export async function readActiveConfig(repoPath, opts = {}) {
detectGitRoot(absRepoPath),
walkClaudeMdCascade(absRepoPath),
readClaudeJsonProjectSlice(absRepoPath),
enumeratePlugins(),
enumeratePlugins(absRepoPath),
readSettingsCascade(absRepoPath),
]);
// Skills depend on plugins
const [skills, hooks, mcpServers] = await Promise.all([
// Skills, hooks, MCP, and the v5.6 enumerations all depend on plugins
const [skills, hooks, mcpServers, rules, agents, outputStyles] = await Promise.all([
enumerateSkills(plugins),
readActiveHooks(absRepoPath, plugins),
readActiveMcpServers(absRepoPath, claudeJsonSlice, plugins),
enumerateRules(absRepoPath, plugins),
enumerateAgents(absRepoPath, plugins),
enumerateOutputStyles(absRepoPath, plugins),
]);
// Totals
@ -854,6 +1241,9 @@ export async function readActiveConfig(repoPath, opts = {}) {
skills: skills.length,
mcpServers: mcpServers.length,
hooks: hooks.length,
rules: rules.length,
agents: agents.length,
outputStyles: outputStyles.length,
claudeMdFiles: claudeMd.files.length,
estimatedTokens: {
claudeMd: claudeMd.estimatedTokens,
@ -861,6 +1251,9 @@ export async function readActiveConfig(repoPath, opts = {}) {
skills: skills.reduce((s, k) => s + k.estimatedTokens, 0),
mcpServers: mcpServers.reduce((s, m) => s + m.estimatedTokens, 0),
hooks: hooks.reduce((s, h) => s + h.estimatedTokens, 0),
rules: rules.reduce((s, r) => s + r.estimatedTokens, 0),
agents: agents.reduce((s, a) => s + a.estimatedTokens, 0),
outputStyles: outputStyles.reduce((s, o) => s + o.estimatedTokens, 0),
grandTotal: 0,
},
};
@ -869,7 +1262,10 @@ export async function readActiveConfig(repoPath, opts = {}) {
totals.estimatedTokens.plugins +
totals.estimatedTokens.skills +
totals.estimatedTokens.mcpServers +
totals.estimatedTokens.hooks;
totals.estimatedTokens.hooks +
totals.estimatedTokens.rules +
totals.estimatedTokens.agents +
totals.estimatedTokens.outputStyles;
const warnings = [];
@ -898,6 +1294,9 @@ export async function readActiveConfig(repoPath, opts = {}) {
skills,
mcpServers,
hooks,
rules,
agents,
outputStyles,
settings: { cascade: settingsCascade },
totals,
suggestDisables,

View file

@ -0,0 +1,51 @@
/**
* Active-model resolution for the `--context-window auto` probe (B8b).
*
* Reads the configured model the way Claude Code itself resolves it, so the
* window probe (context-window.mjs `modelToContextWindow`) sees the real model:
* 1. the shell `ANTHROPIC_MODEL` override (applies to the launched session);
* 2. otherwise the settings cascade `model` field user `~/.claude`, then
* project `.claude`, then project-local `.claude` (local > project > user).
*
* Reads the cascade files directly (like isBundledSkillsDisabled) rather than via
* config-discovery classification, and takes an injectable `env` so it is
* deterministic and hermetic under the test HOME. Returns null when no model is
* pinned anywhere the honest signal that `auto` must fall back to advisory.
*
* Zero external dependencies (repo invariant).
*/
import { join } from 'node:path';
import { readTextFile } from './file-discovery.mjs';
import { parseJson } from './yaml-parser.mjs';
/**
* @param {string|null|undefined} projectPath - project root, to also read project + local settings
* @param {{ env?: Record<string,string|undefined> }} [opts]
* @returns {Promise<string|null>} the resolved model id/alias, or null if unset
*/
export async function resolveActiveModel(projectPath, { env = process.env } = {}) {
// 1. Shell ANTHROPIC_MODEL overrides settings (CC: applies to the session).
const envModel = typeof env?.ANTHROPIC_MODEL === 'string' ? env.ANTHROPIC_MODEL.trim() : '';
if (envModel) return envModel;
// 2. Settings cascade: user -> project -> project-local, later wins.
const home = (env && (env.HOME || env.USERPROFILE)) || '';
const candidates = [];
if (home) candidates.push(join(home, '.claude', 'settings.json'));
if (projectPath) {
candidates.push(join(projectPath, '.claude', 'settings.json'));
candidates.push(join(projectPath, '.claude', 'settings.local.json'));
}
let model = null;
for (const p of candidates) {
const content = await readTextFile(p);
if (!content) continue;
const parsed = parseJson(content);
if (parsed && typeof parsed.model === 'string' && parsed.model.trim()) {
model = parsed.model.trim();
}
}
return model;
}

View file

@ -0,0 +1,132 @@
/**
* Agent-listing budget single source of truth for the AGT scanner.
*
* Claude Code injects a listing of every active agent's name+description into the
* system prompt so the model knows which subagents it can delegate to. With many
* installed agents that listing is a large always-loaded cost: re-sent every turn
* whether or not a delegation actually happens.
*
* CRUCIAL HONESTY CAVEAT this is why AGT differs from SKL. Unlike the skill
* listing (documented ~2% allotment, CC 2.1.32; verified 1,536-char truncation
* cap, CC 2.1.105), the agent-listing mechanism is NOT documented agents are
* absent from Claude Code's published context breakdown. Therefore:
* - "always-loaded" is INFERRED (reasoned from the skill analogue + agents'
* absence from the deferred/on-demand list), not documented.
* - the token figure is an UPPER-BOUND ESTIMATE, not measured telemetry.
* - there is no documented per-listing allotment, so the aggregate budget is a
* config-audit heuristic anchored by analogy to the skill listing on a
* conservative 200k window, and the evidence says so loudly.
* - there is no verified per-description truncation cap, so (unlike CA-SKL-001)
* each description contributes its FULL length to the aggregate.
*
* Zero external dependencies.
*/
import { estimateTokens, enumeratePlugins, enumerateAgents } from './active-config-reader.mjs';
import { readTextFile } from './file-discovery.mjs';
import { parseFrontmatter } from './yaml-parser.mjs';
import { CONTEXT_WINDOW_ANCHOR, LARGE_CONTEXT_WINDOW, withCommas } from './context-window.mjs';
// Heuristic budget by analogy to the skill listing's documented ~2% allotment.
// Agents have NO documented allotment of their own — disclosed loudly in
// BUDGET_CALIBRATION_NOTE so the number is never mistaken for a CC guarantee.
export const BUDGET_FRACTION = 0.02;
export const AGGREGATE_BUDGET_TOKENS = Math.round(BUDGET_FRACTION * CONTEXT_WINDOW_ANCHOR); // 4000
export const LARGE_CONTEXT_BUDGET_TOKENS = Math.round(BUDGET_FRACTION * LARGE_CONTEXT_WINDOW); // 20000
export { CONTEXT_WINDOW_ANCHOR, LARGE_CONTEXT_WINDOW, withCommas };
// Per-agent soft cap for the description-bloat advisory. This is a HEURISTIC,
// NOT a truncation cap: agents have no verified per-description limit (unlike
// the skill listing's 1,536-char cap, CC 2.1.105). We reuse the 500-char
// bloat threshold TOK pattern F already applies to SKILL.md descriptions, so a
// long agent description is flagged for the same reason — every char re-enters
// context in the always-loaded agent listing on every turn — without ever
// claiming Claude Code drops the tail.
export const PER_AGENT_DESC_SOFT_CAP = 500;
// The honest framing required because (a) the mechanism is inferred and (b) the
// budget depends on a context window we cannot observe. Appended to overflow evidence.
export const BUDGET_CALIBRATION_NOTE =
'the agent-listing always-loaded mechanism is INFERRED, not documented (agents are absent from ' +
"Claude Code's published context breakdown), so this token figure is an UPPER-BOUND ESTIMATE, not " +
'measured telemetry. unlike the skill listing there is no documented per-listing allotment, so this ' +
'budget is a config-audit heuristic anchored on a conservative 200k window; at ' +
`${withCommas(LARGE_CONTEXT_WINDOW)} context the budget is ~${withCommas(LARGE_CONTEXT_BUDGET_TOKENS)} ` +
'tok and you are likely within it';
/**
* @typedef {object} AgentBudgetAssessment
* @property {number} scanned - number of agent descriptions assessed
* @property {number} aggregateChars - sum of description lengths (no cap; no verified truncation)
* @property {number} aggregateTokens - estimateTokens(aggregateChars, 'markdown')
* @property {number} budgetTokens - AGGREGATE_BUDGET_TOKENS (the 200k-anchored heuristic)
* @property {boolean} overBudget - aggregateTokens strictly greater than budgetTokens
* @property {number} overBy - tokens over budget (0 when not over)
*/
/**
* Pure aggregate-budget assessment. Each description contributes its full length:
* agents have no verified truncation cap, so nothing is dropped from the estimate.
*
* @param {number[]} descLengths - one entry per active agent (description char count)
* @returns {AgentBudgetAssessment}
*/
export function assessAgentListingBudget(descLengths) {
let aggregateChars = 0;
for (const len of descLengths) {
const safe = (typeof len === 'number' && Number.isFinite(len) && len > 0) ? len : 0;
aggregateChars += safe;
}
const aggregateTokens = estimateTokens(aggregateChars, 'markdown');
const overBudget = aggregateTokens > AGGREGATE_BUDGET_TOKENS;
return {
scanned: descLengths.length,
aggregateChars,
aggregateTokens,
budgetTokens: AGGREGATE_BUDGET_TOKENS,
overBudget,
overBy: overBudget ? aggregateTokens - AGGREGATE_BUDGET_TOKENS : 0,
};
}
/**
* @typedef {object} ActiveAgentEntry
* @property {string} name
* @property {'user'|'plugin'} source
* @property {string|null} pluginName
* @property {string} path
* @property {number} descLength
*/
/**
* Enumerate every active agent (user + plugin) and measure the listing budget.
* HOME-scoped (mirrors the skill listing): resolves ~/.claude via process.env.HOME
* and excludes project-scoped agents those load only inside their own repo.
* Callers that run under test MUST override HOME (runScannerWithHome pattern).
*
* @returns {Promise<{ agents: ActiveAgentEntry[], aggregate: AgentBudgetAssessment }>}
*/
export async function measureActiveAgentListing() {
const plugins = await enumeratePlugins();
// enumerateAgents yields project + user + plugin; drop project (repo-local).
const allAgents = await enumerateAgents('', plugins);
const agents = [];
for (const agent of allAgents) {
if (!agent || agent.source === 'project' || typeof agent.path !== 'string') continue;
const content = await readTextFile(agent.path);
if (!content) continue;
const fm = parseFrontmatter(content)?.frontmatter || null;
const desc = (fm && typeof fm.description === 'string') ? fm.description : '';
agents.push({
name: agent.name,
source: agent.source,
pluginName: agent.pluginName,
path: agent.path,
descLength: desc.length,
});
}
const aggregate = assessAgentListingBudget(agents.map((a) => a.descLength));
return { agents, aggregate };
}

View file

@ -10,15 +10,29 @@ import { join, basename } from 'node:path';
import { createHash } from 'node:crypto';
import { homedir } from 'node:os';
const BACKUP_ROOT = join(homedir(), '.config-audit', 'backups');
const MAX_BACKUPS = 10;
/**
* Get the backup root directory path.
*
* Canonical location is `~/.claude/config-audit/backups` the path every
* command, agent and doc uses. `CONFIG_AUDIT_BACKUP_ROOT` overrides it so tests
* never write into the operator's real home.
* @returns {string}
*/
export function getBackupDir() {
return BACKUP_ROOT;
return process.env.CONFIG_AUDIT_BACKUP_ROOT
|| join(homedir(), '.claude', 'config-audit', 'backups');
}
/**
* Get the pre-v2.2.0 backup root. Read-only: nothing writes here any more, but
* backups made before the move must stay listable and restorable.
* @returns {string}
*/
export function getLegacyBackupDir() {
return process.env.CONFIG_AUDIT_LEGACY_BACKUP_ROOT
|| join(homedir(), '.config-audit', 'backups');
}
/**
@ -63,7 +77,7 @@ export function checksum(content) {
*/
export function createBackup(files, opts = {}) {
const backupId = opts.backupId || generateBackupId();
const backupPath = join(BACKUP_ROOT, backupId);
const backupPath = join(getBackupDir(), backupId);
const filesDir = join(backupPath, 'files');
mkdirSync(filesDir, { recursive: true });
@ -128,7 +142,7 @@ function serializeManifest(manifest) {
* @returns {object}
*/
export function parseManifest(content) {
const result = { created_at: '', backup_id: '', files: [] };
const result = { created_at: '', backup_id: '', files: [], created: [] };
const createdMatch = content.match(/created_at:\s*"([^"]+)"/);
if (createdMatch) result.created_at = createdMatch[1];
@ -136,7 +150,7 @@ export function parseManifest(content) {
const idMatch = content.match(/backup_id:\s*"([^"]+)"/);
if (idMatch) result.backup_id = idMatch[1];
// Parse file entries
// Parse file entries — engine format (quoted `original_path:` …).
const fileBlocks = content.split(/\n\s+-\s+original_path:/).slice(1);
for (const block of fileBlocks) {
const origMatch = block.match(/^\s*"([^"]+)"/);
@ -154,6 +168,45 @@ export function parseManifest(content) {
}
}
// Parse file entries — implement-flow format. `commands/implement.md` has the
// agent hand-build the backup dir, so real manifests on disk use unquoted
// `- backup:` / `original:` / `sha256:`. Reading only the engine format made
// restoreBackup a success-shaped no-op on every backup implement produced.
if (result.files.length === 0) {
const implBlocks = content.split(/\n\s+-\s+backup:/).slice(1);
for (const block of implBlocks) {
const bpMatch = block.match(/^\s*(\S+)/);
const origMatch = block.match(/original:\s*(\S+)/);
const csMatch = block.match(/sha256:\s*(\S+)/);
if (origMatch && bpMatch && csMatch) {
result.files.push({
originalPath: origMatch[1],
backupPath: bpMatch[1],
checksum: csMatch[1],
sizeBytes: 0,
});
}
}
if (!result.backup_id) {
const implId = content.match(/^created:\s*(\S+)\s*$/m);
if (implId) result.backup_id = implId[1];
}
}
// Files the implement step CREATED. A backup cannot hold a file that did not
// exist, so rollback can never restore these — but it must be able to say so.
const lines = content.split('\n');
const createdAt = lines.findIndex(l => /^created:[ \t]*$/.test(l));
if (createdAt !== -1) {
for (const line of lines.slice(createdAt + 1)) {
const item = line.match(/^[ \t]+-[ \t]+(\S+)[ \t]*$/);
if (!item) break;
result.created.push(item[1]);
}
}
return result;
}
@ -161,9 +214,10 @@ export function parseManifest(content) {
* Remove old backups beyond MAX_BACKUPS.
*/
function cleanupOldBackups() {
if (!existsSync(BACKUP_ROOT)) return;
const backupRoot = getBackupDir();
if (!existsSync(backupRoot)) return;
const dirs = readdirSync(BACKUP_ROOT, { withFileTypes: true })
const dirs = readdirSync(backupRoot, { withFileTypes: true })
.filter(d => d.isDirectory())
.map(d => d.name)
.sort();
@ -171,7 +225,7 @@ function cleanupOldBackups() {
if (dirs.length > MAX_BACKUPS) {
const toDelete = dirs.slice(0, dirs.length - MAX_BACKUPS);
for (const dir of toDelete) {
rmSync(join(BACKUP_ROOT, dir), { recursive: true, force: true });
rmSync(join(backupRoot, dir), { recursive: true, force: true });
}
}
}

View file

@ -0,0 +1,113 @@
/**
* best-practices-register loader + schema validator for the machine-readable
* best-practices register (knowledge/best-practices.json).
*
* The register is the SOURCE OF TRUTH for the v5.7 optimization lens (CA-OPT). Each entry
* is provenance-stamped (source.url + source.verified) and carries a confidence; only
* CONFIRMED claims are surfaced user-facing (Verifiseringsplikt). Zero-dependency: native
* JSON, validated by hand here. See docs/v5.7-optimization-lens-plan.md.
*/
import { readFileSync } from 'node:fs';
import { fileURLToPath } from 'node:url';
import { SEVERITY } from './severity.mjs';
/** Absolute path to the bundled register. */
export const REGISTER_PATH = fileURLToPath(
new URL('../../knowledge/best-practices.json', import.meta.url)
);
/** Confidence levels. Only `confirmed` is consumed user-facing by the lens. */
export const CONFIDENCE_LEVELS = Object.freeze(['confirmed', 'inferred', 'unverified']);
const VALID_SEVERITIES = new Set(Object.keys(SEVERITY));
const ID_RE = /^BP-[A-Z]+-\d{3}$/;
const DATE_RE = /^\d{4}-\d{2}-\d{2}$/;
/**
* Load + parse a register file (defaults to the bundled one). Throws on missing file or
* invalid JSON callers that want graceful handling should try/catch.
* @param {string} [path]
* @returns {{version:number, entries:object[]}}
*/
export function loadRegister(path = REGISTER_PATH) {
return JSON.parse(readFileSync(path, 'utf8'));
}
/**
* Validate a parsed register against the schema. Never throws returns a result so the
* caller (and tests) can inspect every problem at once.
* @param {unknown} data
* @returns {{valid:boolean, errors:string[]}}
*/
export function validateRegister(data) {
const errors = [];
if (!data || typeof data !== 'object' || Array.isArray(data)) {
return { valid: false, errors: ['register must be an object'] };
}
if (typeof data.version !== 'number') errors.push('version must be a number');
if (!Array.isArray(data.entries)) {
errors.push('entries must be an array');
return { valid: false, errors };
}
const seen = new Set();
data.entries.forEach((e, i) => {
const at = `entry[${i}]${e && typeof e === 'object' && e.id ? ` (${e.id})` : ''}`;
if (!e || typeof e !== 'object' || Array.isArray(e)) {
errors.push(`${at}: must be an object`);
return;
}
// id — required, BP-TOPIC-NNN, unique
if (typeof e.id !== 'string' || !ID_RE.test(e.id)) {
errors.push(`${at}: id must match BP-TOPIC-NNN`);
} else if (seen.has(e.id)) {
errors.push(`${at}: duplicate id`);
} else {
seen.add(e.id);
}
// claim — required, non-empty
if (typeof e.claim !== 'string' || e.claim.trim() === '') {
errors.push(`${at}: claim is required`);
}
// confidence — required, enum
if (!CONFIDENCE_LEVELS.includes(e.confidence)) {
errors.push(`${at}: confidence must be one of ${CONFIDENCE_LEVELS.join('|')}`);
}
// source — required object with url + verified date (provenance)
if (!e.source || typeof e.source !== 'object' || Array.isArray(e.source)) {
errors.push(`${at}: source is required`);
} else {
if (typeof e.source.url !== 'string' || e.source.url.trim() === '') {
errors.push(`${at}: source.url is required`);
}
if (typeof e.source.verified !== 'string' || !DATE_RE.test(e.source.verified)) {
errors.push(`${at}: source.verified must be a YYYY-MM-DD date`);
}
}
// optional fields — typed only when present
if (e.severity !== undefined && !VALID_SEVERITIES.has(e.severity)) {
errors.push(`${at}: severity must be one of ${[...VALID_SEVERITIES].join('|')}`);
}
for (const key of ['mechanism', 'appliesTo', 'recommendation', 'category']) {
if (e[key] !== undefined && typeof e[key] !== 'string') {
errors.push(`${at}: ${key} must be a string`);
}
}
if (e.lensCheck !== undefined && e.lensCheck !== null && typeof e.lensCheck !== 'string') {
errors.push(`${at}: lensCheck must be a string or null`);
}
});
return { valid: errors.length === 0, errors };
}
/**
* Look up an entry by id.
* @param {{entries:object[]}} register
* @param {string} id
* @returns {object|undefined}
*/
export function getEntry(register, id) {
if (!register || !Array.isArray(register.entries)) return undefined;
return register.entries.find((e) => e.id === id);
}

View file

@ -0,0 +1,78 @@
/**
* campaign-export plan-export transforms (v5.7 Fase 2, Block 4c).
*
* The second half of Block 4 ("durable backlog + execution"). Block 4b built the cross-repo
* prioritized backlog the user picks from; this exports a picked repo's per-repo action plan
* into the TARGET repo's OWN `docs/` directory, so the plan gets a durable, human-readable
* home where the work is done ("planer følger arbeidsstedet" the operator's continuity rule
* that plans live next to the workplace, in `docs/`).
*
* Design mirrors campaign-ledger: PURE, deterministic transforms `now` is injected as a
* YYYY-MM-DD string, never read from the clock here, so they are fully unit-testable. The IO
* (loading the ledger, reading the session's action-plan.md, writing the exported file) lives
* in the thin `campaign-export-cli` shell. The transforms throw on programmer error
* (missing/blank required field), consistent with the ledger transforms.
*
* NOTE on execution: Block 4c deliberately adds NO new execution machinery. Execution reuses
* the existing per-repo `/config-audit implement` (which backs up every changed file, applies
* the plan from the session, and verifies) + `/config-audit rollback`. The exported `docs/`
* copy is the repo's durable record of the plan, NOT the execution input `implement` still
* reads the canonical plan from the session directory. See docs/v5.7-optimization-lens-plan.md
* §Fase 2 (Block 4).
*/
import { join } from 'node:path';
/**
* The exported plan's destination inside the TARGET repo's own `docs/`. Keyed on the source
* `sessionId` (timestamp-unique per audit) rather than the calendar date, so two audits of the
* same repo on the same day produce distinct files (history is preserved, never silently
* overwritten) and the filename ties the export back to the audit that produced it.
*
* @param {string} repoPath - absolute path to the target repo (the ledger stores it resolved)
* @param {string} sessionId - the config-audit session that produced the plan
* @returns {string} `<repoPath>/docs/config-audit-plan-<sessionId>.md`
*/
export function planExportPath(repoPath, sessionId) {
if (typeof repoPath !== 'string' || repoPath.trim() === '') {
throw new TypeError('repoPath is required');
}
if (typeof sessionId !== 'string' || sessionId.trim() === '') {
throw new TypeError('sessionId is required');
}
return join(repoPath, 'docs', `config-audit-plan-${sessionId}.md`);
}
/**
* Assemble the exported document: a provenance header (who/when/where this came from + how to
* execute and undo it) followed by the verbatim session plan body. Pure given the same inputs
* it always produces the same bytes, so it is snapshot-testable.
*
* @param {{repoName:string, repoPath:string, sessionId:string, planMarkdown:string, now:string}} input
* @returns {string} the full markdown to write into the repo's docs/
*/
export function buildPlanExportDocument({ repoName, repoPath, sessionId, planMarkdown, now } = {}) {
for (const [k, v] of Object.entries({ repoName, repoPath, sessionId, planMarkdown, now })) {
if (typeof v !== 'string' || v.trim() === '') {
throw new TypeError(`${k} is required`);
}
}
const header = [
`# Config-Audit Action Plan — ${repoName}`,
'',
`> Exported from the config-audit machine-wide campaign on ${now}.`,
`> **Repo:** \`${repoPath}\``,
`> **Source session:** \`${sessionId}\``,
'>',
'> Generated by `/config-audit plan`. To **execute**: run `/config-audit implement` in this',
'> repo — it backs up every changed file, applies the plan, then verifies the result. To',
'> **undo**: `/config-audit rollback`. Record progress back in the campaign with',
`> \`/config-audit campaign set-status ${repoPath} implemented\`.`,
'',
'---',
'',
].join('\n');
return `${header}${planMarkdown.trimEnd()}\n`;
}

View file

@ -0,0 +1,369 @@
/**
* campaign-ledger durable, machine-wide campaign ledger (v5.7 Fase 2, Block 3a THIN).
*
* The ledger sits ABOVE individual config-audit sessions: it tracks which repos are part
* of a machine-wide audit campaign, each repo's lifecycle status (pending audited
* planned implemented), and a machine-wide roll-up by status + severity. It is resumable
* across sessions because it persists to a single JSON file OUTSIDE the plugin dir
* (`~/.claude/config-audit/campaign-ledger.json`, next to `sessions/` and `mcp-cache/`),
* so it survives plugin uninstall/reinstall/upgrade.
*
* Design mirrors the knowledge-refresh precedent: the transformations are PURE and
* deterministic (every "now" is injected as a YYYY-MM-DD string, never read from the clock
* here), so they are fully unit-testable; a thin IO shell (load/save, explicit path) does
* the only filesystem work. `validateLedger` is soft (returns a result, never throws) for
* externally-loaded data; the transforms throw on programmer error (invalid status, unknown
* path). `schemaVersion` is stamped from the start so a future Block 4 migration is cheap.
*
* THIN scope (Block 3a): ledger core + roll-up + persistence only NOT execution,
* orchestration, or a command surface (those are Blocks 3b/3c/4). Zero dependencies.
* See docs/v5.7-optimization-lens-plan.md §Fase 2.
*/
import { readFile, writeFile, mkdir } from 'node:fs/promises';
import { dirname, join, resolve } from 'node:path';
import { homedir } from 'node:os';
/** Ledger schema version — bump + add a migration (Block 4) on any breaking shape change. */
export const CAMPAIGN_SCHEMA_VERSION = 1;
/** Per-repo lifecycle, in order. A repo advances through these as the campaign progresses. */
export const STATUSES = Object.freeze(['pending', 'audited', 'planned', 'implemented']);
/** Severity buckets aggregated by the machine-wide roll-up. */
const SEVERITIES = Object.freeze(['critical', 'high', 'medium', 'low']);
/**
* Order-of-magnitude severity weights for the cross-repo backlog priority score. Each tier
* dominates the next so a single higher-severity finding outranks many lower ones; exact
* score collisions are still broken deterministically by the lexicographic + name tie-break
* in `buildBacklog`. Exported so the score is documented, not a magic number.
*/
export const SEVERITY_WEIGHTS = Object.freeze({ critical: 1000, high: 100, medium: 10, low: 1 });
const DATE_RE = /^\d{4}-\d{2}-\d{2}$/;
/** Validate an injected `now` (required, YYYY-MM-DD). Throws — callers pass today's date. */
function requireNow(now) {
if (typeof now !== 'string' || !DATE_RE.test(now)) {
throw new TypeError('now must be a YYYY-MM-DD string');
}
return now;
}
/** Canonicalize a repo path so the same repo never appears twice under different spellings. */
function normalizePath(path) {
if (typeof path !== 'string' || path.trim() === '') {
throw new TypeError('repo path is required');
}
return resolve(path);
}
/**
* Build an empty, versioned ledger stamped with `now`.
* @param {{now: string}} opts
* @returns {{schemaVersion:number, createdDate:string, updatedDate:string, repos:object[]}}
*/
export function createLedger({ now } = {}) {
requireNow(now);
return { schemaVersion: CAMPAIGN_SCHEMA_VERSION, createdDate: now, updatedDate: now, repos: [] };
}
/**
* Add a repo to the campaign (status `pending`). Idempotent on the normalized path a repo
* already present is left untouched (its progress is NOT reset). Returns a NEW ledger.
* @param {object} ledger
* @param {{path: string, name?: string}} repo
* @param {{now: string}} opts
*/
export function addRepo(ledger, { path, name } = {}, { now } = {}) {
requireNow(now);
const resolved = normalizePath(path);
if (ledger.repos.some((r) => r.path === resolved)) {
return ledger; // idempotent: already tracked, preserve its status
}
const repo = {
path: resolved,
name: typeof name === 'string' && name.trim() !== '' ? name : resolved.split('/').pop(),
status: 'pending',
sessionId: null,
findingsBySeverity: null,
tokens: null,
updatedDate: now,
};
return { ...ledger, updatedDate: now, repos: [...ledger.repos, repo] };
}
/**
* Transition a tracked repo to a new status, optionally attaching the audit's
* findings-by-severity and the producing sessionId. Returns a NEW ledger.
* @param {object} ledger
* @param {string} path
* @param {string} status - one of STATUSES
* @param {{now: string, findingsBySeverity?: object|null, sessionId?: string|null}} opts
*/
export function setRepoStatus(ledger, path, status, { now, findingsBySeverity, sessionId } = {}) {
requireNow(now);
if (!STATUSES.includes(status)) {
throw new RangeError(`invalid status "${status}" — must be one of ${STATUSES.join(', ')}`);
}
const resolved = normalizePath(path);
const idx = ledger.repos.findIndex((r) => r.path === resolved);
if (idx === -1) {
throw new Error(`repo "${resolved}" is not in the ledger — addRepo first`);
}
const updated = { ...ledger.repos[idx], status, updatedDate: now };
if (findingsBySeverity !== undefined) updated.findingsBySeverity = findingsBySeverity;
if (sessionId !== undefined) updated.sessionId = sessionId;
const repos = ledger.repos.slice();
repos[idx] = updated;
return { ...ledger, updatedDate: now, repos };
}
/** Load-pattern buckets carried by a token summary (manifest `summarizeByLoadPattern` shape). */
const LOAD_PATTERNS = Object.freeze(['always', 'onDemand', 'external', 'unknown']);
/**
* Set the machine-wide SHARED global always-loaded layer (global CLAUDE.md + agent listing +
* global MCP + unscoped global rules + active plugins' always-loaded components). Stored ONCE
* at the ledger root never per repo so the machine-wide roll-up counts it exactly once
* (the structural guard against the historic double-count). The `summary` is the shape
* manifest's `summarizeByLoadPattern` emits: `{always|onDemand|external|unknown: {tokens,count}}`.
* Returns a NEW ledger. (B2b populates this from a live cross-repo sweep.)
* @param {object} ledger
* @param {object} summary
* @param {{now: string}} opts
*/
export function setSharedGlobal(ledger, summary, { now } = {}) {
requireNow(now);
return { ...ledger, updatedDate: now, sharedGlobal: summary };
}
/**
* Attach a tracked repo's PER-REPO always-loaded token delta (its project-scoped contribution
* beyond the shared global layer project CLAUDE.md / rules / agents / MCP). Same `summary`
* shape as `setSharedGlobal`. Mirrors `setRepoStatus`: throws if the repo is untracked. Returns
* a NEW ledger.
* @param {object} ledger
* @param {string} path
* @param {object} tokens
* @param {{now: string}} opts
*/
export function setRepoTokens(ledger, path, tokens, { now } = {}) {
requireNow(now);
const resolved = normalizePath(path);
const idx = ledger.repos.findIndex((r) => r.path === resolved);
if (idx === -1) {
throw new Error(`repo "${resolved}" is not in the ledger — addRepo first`);
}
const repos = ledger.repos.slice();
repos[idx] = { ...ledger.repos[idx], tokens, updatedDate: now };
return { ...ledger, updatedDate: now, repos };
}
/** Tolerantly read the four load-pattern token numbers from a stored summary (or null/old data). */
function bucketTokens(summary) {
const out = {};
for (const k of LOAD_PATTERNS) {
const v = summary && summary[k];
out[k] = v && typeof v.tokens === 'number' ? v.tokens : 0;
}
return out;
}
/**
* Machine-wide roll-up: repo counts by status, a severity total across every repo carrying
* `findingsBySeverity`, AND a machine-wide always-loaded token bill. The token aggregate adds
* the SHARED global layer (counted once, from the ledger root) to the SUM of per-repo deltas,
* so `tokens.machineWide` is the honest "context spent every turn across the whole machine".
* Pure derivation never mutates; tolerant of old ledgers without token fields.
* @param {object} ledger
* @returns {{totalRepos:number, byStatus:object, bySeverity:object, reposWithFindings:number, tokens:object}}
*/
export function rollUp(ledger) {
const byStatus = Object.fromEntries(STATUSES.map((s) => [s, 0]));
const bySeverity = Object.fromEntries(SEVERITIES.map((s) => [s, 0]));
let reposWithFindings = 0;
const sharedGlobal = bucketTokens(ledger.sharedGlobal);
const perRepoDelta = Object.fromEntries(LOAD_PATTERNS.map((k) => [k, 0]));
const byRepo = [];
let reposWithTokens = 0;
for (const repo of ledger.repos) {
if (byStatus[repo.status] !== undefined) byStatus[repo.status] += 1;
const f = repo.findingsBySeverity;
if (f && typeof f === 'object') {
reposWithFindings += 1;
for (const sev of SEVERITIES) {
if (typeof f[sev] === 'number') bySeverity[sev] += f[sev];
}
}
if (repo.tokens && typeof repo.tokens === 'object') {
reposWithTokens += 1;
const b = bucketTokens(repo.tokens);
for (const k of LOAD_PATTERNS) perRepoDelta[k] += b[k];
byRepo.push({ name: repo.name, path: repo.path, always: b.always, onDemand: b.onDemand, external: b.external });
}
}
// DESC by always-loaded cost ("most expensive repos"); deterministic name tie-break.
byRepo.sort((x, y) => y.always - x.always || x.name.localeCompare(y.name));
const machineWide = Object.fromEntries(
LOAD_PATTERNS.map((k) => [k, sharedGlobal[k] + perRepoDelta[k]]),
);
return {
totalRepos: ledger.repos.length,
byStatus,
bySeverity,
reposWithFindings,
tokens: { sharedGlobal, perRepoDelta, machineWide, reposWithTokens, byRepo },
};
}
/**
* Build the single, machine-wide PRIORITIZED backlog the user picks from. Pure derivation
* never mutates. The actionable unit is a REPO (the ledger tracks per-repo severity counts,
* not individual findings it tracks state, it does not re-run audits), so each backlog item
* is one repo with outstanding work.
*
* Inclusion: a repo is in the backlog iff it is NOT yet `implemented` AND has at least one
* outstanding finding (`totalFindings > 0`). `implemented` repos are done; `pending` and
* zero-finding repos have nothing known to fix (they still surface in `rollUp.byStatus`).
*
* Order: DESC by `weightedScore` (SEVERITY_WEIGHTS), tie-broken lexicographically by
* criticalhighmediumlow count, then ascending by `name` fully deterministic, and the
* tie-break preserves "criticals always win" even when two repos share a weighted score.
*
* @param {object} ledger
* @returns {Array<{path:string,name:string,status:string,sessionId:string|null,findingsBySeverity:object,totalFindings:number,weightedScore:number,rank:number}>}
*/
export function buildBacklog(ledger) {
const items = [];
for (const repo of ledger.repos) {
if (repo.status === 'implemented') continue;
const f = repo.findingsBySeverity;
if (!f || typeof f !== 'object') continue;
const findingsBySeverity = Object.fromEntries(
SEVERITIES.map((s) => [s, typeof f[s] === 'number' ? f[s] : 0]),
);
const totalFindings = SEVERITIES.reduce((sum, s) => sum + findingsBySeverity[s], 0);
if (totalFindings === 0) continue;
const weightedScore = SEVERITIES.reduce(
(score, s) => score + findingsBySeverity[s] * SEVERITY_WEIGHTS[s],
0,
);
items.push({
path: repo.path,
name: repo.name,
status: repo.status,
sessionId: repo.sessionId ?? null,
findingsBySeverity,
totalFindings,
weightedScore,
});
}
items.sort(
(x, y) =>
y.weightedScore - x.weightedScore ||
y.findingsBySeverity.critical - x.findingsBySeverity.critical ||
y.findingsBySeverity.high - x.findingsBySeverity.high ||
y.findingsBySeverity.medium - x.findingsBySeverity.medium ||
y.findingsBySeverity.low - x.findingsBySeverity.low ||
x.name.localeCompare(y.name),
);
return items.map((item, i) => ({ ...item, rank: i + 1 }));
}
/**
* Validate a parsed ledger against the schema. Never throws returns every problem at once
* so the caller (and tests) can inspect them. Soft by design (loaded data may be corrupt).
* @param {unknown} data
* @returns {{valid:boolean, errors:string[]}}
*/
export function validateLedger(data) {
const errors = [];
if (!data || typeof data !== 'object' || Array.isArray(data)) {
return { valid: false, errors: ['ledger must be an object'] };
}
if (data.schemaVersion !== CAMPAIGN_SCHEMA_VERSION) {
errors.push(`schemaVersion must be ${CAMPAIGN_SCHEMA_VERSION}`);
}
for (const field of ['createdDate', 'updatedDate']) {
if (typeof data[field] !== 'string' || !DATE_RE.test(data[field])) {
errors.push(`${field} must be YYYY-MM-DD`);
}
}
if (!Array.isArray(data.repos)) {
errors.push('repos must be an array');
return { valid: false, errors };
}
const seen = new Set();
data.repos.forEach((r, i) => {
const at = `repos[${i}]${r && typeof r === 'object' && r.path ? ` (${r.path})` : ''}`;
if (!r || typeof r !== 'object' || Array.isArray(r)) {
errors.push(`${at}: must be an object`);
return;
}
if (typeof r.path !== 'string' || r.path.trim() === '') {
errors.push(`${at}: path is required`);
} else if (seen.has(r.path)) {
errors.push(`${at}: duplicate path`);
} else {
seen.add(r.path);
}
if (!STATUSES.includes(r.status)) {
errors.push(`${at}: status must be one of ${STATUSES.join(', ')}`);
}
});
return { valid: errors.length === 0, errors };
}
// ── Persistence (thin IO shell) ────────────────────────────────────────────────
/**
* Default on-disk location: next to `sessions/`, OUTSIDE the plugin dir, so the campaign
* survives plugin uninstall/reinstall/upgrade.
* @returns {string}
*/
export function defaultLedgerPath() {
return join(homedir(), '.claude', 'config-audit', 'campaign-ledger.json');
}
/**
* Load + parse a ledger file. Returns `null` if the file does not exist (graceful first run);
* other read/parse errors propagate so corruption is not silently swallowed.
* @param {string} [path]
* @returns {Promise<object|null>}
*/
export async function loadLedger(path = defaultLedgerPath()) {
let content;
try {
content = await readFile(path, 'utf8');
} catch (err) {
if (err && err.code === 'ENOENT') return null;
throw err;
}
return JSON.parse(content);
}
/**
* Persist a ledger as human-readable JSON, creating parent directories as needed.
* @param {string} path
* @param {object} ledger
* @returns {Promise<{path: string}>}
*/
export async function saveLedger(path = defaultLedgerPath(), ledger) {
await mkdir(dirname(path), { recursive: true });
await writeFile(path, `${JSON.stringify(ledger, null, 2)}\n`, 'utf8');
return { path };
}

View file

@ -0,0 +1,132 @@
/**
* Context-window constants single source of truth.
*
* Several Claude Code budgets scale with the model's context window:
* - the skill listing is allotted ~2% of context (CC 2.1.32, changelog L2860);
* - the "CLAUDE.md is too long" warning threshold scales with it (CC 2.1.169).
*
* We cannot observe the user's actual context window, so scanners anchor on a
* conservative 200k window (the smallest common size it fires earliest, the
* safe default when the window is unknown) and disclose the relaxed 1M figure.
*
* Zero external dependencies.
*/
// Conservative anchor: the smallest common context window. Budgets anchored
// here fire earliest, which is the safe default when the window is unknown.
export const CONTEXT_WINDOW_ANCHOR = 200_000;
// Large context window (Opus/Sonnet 1M tier). Used to disclose how a budget
// relaxes on large-context models.
export const LARGE_CONTEXT_WINDOW = 1_000_000;
// How much larger the 1M window is than the 200k anchor (= 5). A budget that
// scales linearly with the window relaxes by this factor at 1M.
export const LARGE_CONTEXT_SCALE = LARGE_CONTEXT_WINDOW / CONTEXT_WINDOW_ANCHOR;
// Dependency-free thousands separator (repo invariant: zero external deps).
export const withCommas = (n) => String(n).replace(/\B(?=(\d{3})+(?!\d))/g, ',');
// Model families whose context window is the large (1M) tier. Verified June 2026
// (platform.claude.com models overview): Fable 5, Opus 4.8/4.7/4.6 and Sonnet 4.6
// all run a 1M context window. Matched by substring so dated IDs
// (claude-opus-4-8-20260528) and provider-prefixed IDs
// (us.anthropic.claude-opus-4-8) resolve too. Models we cannot confirm (e.g.
// Haiku, older 200k-era IDs) are deliberately left out: the caller then keeps the
// conservative anchor rather than guess a relaxed budget.
export const LARGE_CONTEXT_MODEL_IDS = [
'claude-fable-5',
'claude-opus-4-8',
'claude-opus-4-7',
'claude-opus-4-6',
'claude-sonnet-4-6',
];
// Short aliases Claude Code accepts in the `model` setting / ANTHROPIC_MODEL that
// currently resolve to a 1M-tier model (`opusplan` plans on an Opus-tier model).
export const LARGE_CONTEXT_MODEL_ALIASES = new Set(['opus', 'sonnet', 'fable', 'opusplan']);
/**
* Map a configured model id/alias to its context window, or null when we cannot
* confirm it. Pure: no IO. Used by the `--context-window auto` probe (B8b) so
* known 1M-tier models calibrate budgets instead of falling back to the
* conservative advisory anchor.
*
* @param {string} modelId - e.g. "claude-opus-4-8[1m]", "claude-sonnet-4-6", "opus"
* @returns {number|null} the context window, or null if unrecognized
*/
export function modelToContextWindow(modelId) {
if (typeof modelId !== 'string') return null;
const id = modelId.trim().toLowerCase();
if (!id) return null;
// Explicit tier tag wins — the running session model surfaces as e.g.
// "claude-opus-4-8[1m]". This is the strongest, most future-proof signal.
if (id.includes('[1m]')) return LARGE_CONTEXT_WINDOW;
// Known 1M-tier families (substring → tolerant of date/provider-prefix variants).
for (const fam of LARGE_CONTEXT_MODEL_IDS) {
if (id.includes(fam)) return LARGE_CONTEXT_WINDOW;
}
// Short aliases.
if (LARGE_CONTEXT_MODEL_ALIASES.has(id)) return LARGE_CONTEXT_WINDOW;
// Unknown: cannot confirm the window — keep the conservative anchor (null).
return null;
}
/**
* @typedef {object} ResolvedContextWindow
* @property {number} window - the context window budgets calibrate against
* @property {boolean} advisory - true when the window is unknown: keep the anchor
* but downgrade budget findings to info instead of
* firing them as a breach
* @property {'default'|'explicit'|'auto-probed'|'auto-unresolved'} source
*/
/**
* Resolve the raw `--context-window` CLI value into a window + advisory flag.
*
* Design (B8): the DEFAULT (no flag) is byte-identical to the pre-B8 behavior
* the conservative 200k anchor at full severity. Only an explicit value changes
* calibration. `auto` asks the tool to figure out the window.
*
* B8b: `auto` now probes the configured model (`opts.model`, resolved from the
* settings cascade / ANTHROPIC_MODEL by the orchestrator). A recognized 1M-tier
* model calibrates to its window (source `auto-probed`, not advisory). When the
* model is unknown or unpinned, it keeps the conservative anchor but marks the
* result advisory (source `auto-unresolved`) so SKL/CML downgrade their budget
* findings to info rather than "crying wolf" on a window we cannot confirm.
*
* @param {string|number|null|undefined} arg
* @param {{ model?: string|null }} [opts] - probe input for `auto` (ignored on the
* default/explicit paths, which stay byte-stable).
* @returns {ResolvedContextWindow}
*/
export function resolveContextWindow(arg, opts = {}) {
if (arg == null) {
return { window: CONTEXT_WINDOW_ANCHOR, advisory: false, source: 'default' };
}
if (String(arg).trim().toLowerCase() === 'auto') {
const probed = modelToContextWindow(opts.model);
if (probed) {
return { window: probed, advisory: false, source: 'auto-probed' };
}
return { window: CONTEXT_WINDOW_ANCHOR, advisory: true, source: 'auto-unresolved' };
}
const n = typeof arg === 'number' ? arg : parseInt(String(arg).trim(), 10);
if (Number.isFinite(n) && n > 0) {
return { window: n, advisory: false, source: 'explicit' };
}
// Unparseable / non-positive: fall back to the conservative default (no advisory).
return { window: CONTEXT_WINDOW_ANCHOR, advisory: false, source: 'default' };
}
/**
* Scale a 200k-anchored budget to a given context window. Linear in the window,
* so it is the identity at the anchor (keeps the default byte-stable).
*
* @param {number} anchorValue - the budget/threshold defined at the 200k anchor
* @param {number} window - the target context window
* @returns {number}
*/
export function scaleForWindow(anchorValue, window) {
return Math.round(anchorValue * (window / CONTEXT_WINDOW_ANCHOR));
}

View file

@ -11,8 +11,148 @@ const SKIP_DIRS = new Set([
'node_modules', '.git', 'dist', 'build', 'coverage', '__pycache__',
'.next', '.nuxt', '.output', '.cache', '.turbo', '.parcel-cache',
'vendor', 'venv', '.venv', '.tox',
// A `backups` dir holds backup COPIES, not live config — auditing it as if
// live produces stale findings. config-audit's own session backups
// (~/.claude/config-audit/backups/<ts>/files/.../CLAUDE.md) are the canonical
// case (M-BUG-8), but the rule is general: backups are never live config.
'backups',
]);
// Path marker for the plugin install cache (~/.claude/plugins/cache).
// Structure: <...>/plugins/cache/<marketplace>/<plugin>/<version>/...
// installed_plugins.json's installPath points INTO this tree at the ACTIVE
// version; any other version dir is a stale leftover (superseded install).
const PLUGIN_CACHE_MARKER = `plugins${sep}cache${sep}`;
/**
* Extract the `<marketplace>/<plugin>/<version>` key for a path inside
* ~/.claude/plugins/cache. Returns null when the path is not under
* plugins/cache, or is shallower than the version directory.
* @param {string} absPath
* @returns {string | null}
*/
export function cacheVersionKey(absPath) {
const i = absPath.indexOf(PLUGIN_CACHE_MARKER);
if (i === -1) return null;
const rest = absPath.slice(i + PLUGIN_CACHE_MARKER.length);
const segs = rest.split(sep).filter(Boolean);
if (segs.length < 3) return null; // need marketplace/plugin/version
return segs.slice(0, 3).join('/');
}
/**
* Given a path inside plugins/cache, return the absolute `<...>/plugins`
* directory that owns it (where installed_plugins.json lives).
* @param {string} absPath
* @returns {string | null}
*/
function pluginsDirForCachePath(absPath) {
const i = absPath.indexOf(PLUGIN_CACHE_MARKER);
if (i === -1) return null;
return absPath.slice(0, i + 'plugins'.length);
}
/**
* Read the set of ACTIVE cache version-keys from a `<...>/plugins` directory's
* installed_plugins.json. Each record's installPath points at the version
* Claude Code loads. Returns null when the manifest is absent or unparseable
* callers must then NOT filter (we cannot safely tell active from stale, and
* silently dropping active config is the worse failure).
* @param {string} pluginsDir
* @returns {Promise<Set<string> | null>}
*/
async function readActiveCacheVersions(pluginsDir) {
let raw;
try {
raw = await readFile(join(pluginsDir, 'installed_plugins.json'), 'utf-8');
} catch {
return null;
}
let parsed;
try {
parsed = JSON.parse(raw);
} catch {
return null;
}
const active = new Set();
const plugins = parsed && parsed.plugins;
if (plugins && typeof plugins === 'object') {
for (const recs of Object.values(plugins)) {
if (!Array.isArray(recs)) continue;
for (const r of recs) {
const key = r && r.installPath ? cacheVersionKey(resolve(r.installPath)) : null;
if (key) active.add(key);
}
}
}
return active;
}
/**
* Identify stale plugin-cache versions among discovered files and (optionally)
* filter them out. "Stale" = a cache version-key not referenced by the owning
* installed_plugins.json. Active versions are always kept installPaths point
* INTO the cache, so a blunt "skip all of plugins/cache" would drop live config.
*
* @param {Array} files
* @param {boolean} excludeCache - when true, stale-version files are removed
* @returns {Promise<{ files: Array, staleCacheVersions: Array<{key:string, fileCount:number, estimatedBytes:number}> }>}
*/
async function applyCacheFilter(files, excludeCache) {
const cacheFiles = [];
for (const f of files) {
const key = cacheVersionKey(f.absPath);
if (key) cacheFiles.push({ f, key });
}
if (cacheFiles.length === 0) return { files, staleCacheVersions: [] };
// Union active version-keys across every installed_plugins.json adjacent to
// the cache (a full-machine sweep only ever sees one, but be robust).
const pluginsDirs = new Set();
for (const { f } of cacheFiles) {
const pd = pluginsDirForCachePath(f.absPath);
if (pd) pluginsDirs.add(pd);
}
let active = null;
for (const pd of pluginsDirs) {
const keys = await readActiveCacheVersions(pd);
if (keys) {
if (active === null) active = new Set();
for (const k of keys) active.add(k);
}
}
// No readable manifest → cannot distinguish active from stale → do nothing.
if (active === null) return { files, staleCacheVersions: [] };
const byKey = new Map();
for (const { f, key } of cacheFiles) {
if (!byKey.has(key)) byKey.set(key, []);
byKey.get(key).push(f);
}
const staleCacheVersions = [];
for (const [key, group] of byKey) {
if (!active.has(key)) {
staleCacheVersions.push({
key,
fileCount: group.length,
estimatedBytes: group.reduce((s, f) => s + (f.size || 0), 0),
});
}
}
staleCacheVersions.sort((a, b) => (a.key < b.key ? -1 : a.key > b.key ? 1 : 0));
let outFiles = files;
if (excludeCache && staleCacheVersions.length > 0) {
const staleKeys = new Set(staleCacheVersions.map(s => s.key));
outFiles = files.filter(f => {
const k = cacheVersionKey(f.absPath);
return k === null || !staleKeys.has(k);
});
}
return { files: outFiles, staleCacheVersions };
}
/** Config file patterns to discover */
const CONFIG_PATTERNS = {
claudeMd: /^CLAUDE\.md$|^CLAUDE\.local\.md$/i,
@ -34,7 +174,8 @@ const CONFIG_PATTERNS = {
* @param {object} [opts]
* @param {number} [opts.maxFiles=500] - max files to return
* @param {boolean} [opts.includeGlobal=false] - also scan ~/.claude/
* @returns {Promise<{ files: ConfigFile[], skipped: number }>}
* @param {boolean} [opts.excludeCache=false] - drop stale ~/.claude/plugins/cache versions (B3)
* @returns {Promise<{ files: ConfigFile[], skipped: number, staleCacheVersions: Array }>}
*
* @typedef {{ absPath: string, relPath: string, type: string, scope: string, size: number }} ConfigFile
*/
@ -68,7 +209,8 @@ export async function discoverConfigFiles(targetPath, opts = {}) {
} catch { /* doesn't exist */ }
}
return { files, skipped: skippedRef.count };
const { files: outFiles, staleCacheVersions } = await applyCacheFilter(files, opts.excludeCache || false);
return { files: outFiles, skipped: skippedRef.count, staleCacheVersions };
}
/**
@ -245,7 +387,8 @@ export async function discoverFullMachinePaths() {
* @param {Array<{ path: string, maxDepth: number }>} roots
* @param {object} [opts]
* @param {number} [opts.maxFiles=2000] - global max across all roots
* @returns {Promise<{ files: ConfigFile[], skipped: number }>}
* @param {boolean} [opts.excludeCache=false] - drop stale ~/.claude/plugins/cache versions (B3)
* @returns {Promise<{ files: ConfigFile[], skipped: number, staleCacheVersions: Array }>}
*/
export async function discoverConfigFilesMulti(roots, opts = {}) {
const maxFiles = opts.maxFiles || 2000;
@ -287,7 +430,8 @@ export async function discoverConfigFilesMulti(roots, opts = {}) {
} catch { /* doesn't exist */ }
}
return { files: allFiles, skipped: totalSkipped };
const { files: outFiles, staleCacheVersions } = await applyCacheFilter(allFiles, opts.excludeCache || false);
return { files: outFiles, skipped: totalSkipped, staleCacheVersions };
}
/**

View file

@ -0,0 +1,126 @@
/**
* floor-exclusion the deterministic veto that runs BEFORE the subtraction
* judge sees anything (v5.13, brief §6.0 / §7 q2).
*
* The subtraction lens asks "what no longer earns its always-loaded rent?", and
* that question is only safe to ask because this module answers a prior one
* first: **which blocks are not eligible to be asked about at all?**
*
* Brief §6.0 splits CLAUDE.md content on one axis:
*
* compensatory corrects model *behaviour* ("think before you code").
* A more capable model does it unprompted. Deletable.
* load-bearing a *local fact* no amount of intelligence derives from the
* codebase ("only Forgejo at git.example.test", "system bash
* is 3.2"). FLOOR. Never a deletion candidate.
*
* Why this is deterministic and not the judge's call: precision is asymmetric.
* A missed dead line costs a few tokens per turn; a deleted load-bearing line
* costs a wrong remote, a broken script, or a lost afternoon. A blocking
* guarantee must not rest on a probabilistic prose judgement, so the floor is
* decided here, in code, and the judge only ever ranks what survives.
*
* The veto keys on **underivable local literals** the textual fingerprints of
* a fact that came from this machine rather than from general engineering
* knowledge: an inline code span, a path, a domain, a version pin, a concrete
* filename. Plus §6.0's explicit carve-out: policy invariants (secrets,
* credentials, production, destructive operations) are floor *by decision, not
* by classification* the model would probably honour them unprompted, but the
* cost of being wrong is asymmetric and their token cost is trivial.
*
* Deliberately over-broad. A false veto costs recall (a dead line survives
* another turn); a false clearance costs the guarantee. When in doubt: floor.
*
* Zero external dependencies. Pure: input text boolean.
*/
/** An inline code span — the single strongest local-literal signal. */
const CODE_SPAN_RE = /`[^`\n]+`/;
/** A URL or a bare hostname. `.test`/`.local` included for fixtures. */
const URL_RE = /https?:\/\/|\b[a-z0-9][a-z0-9-]*(?:\.[a-z0-9-]+)+\.(?:com|org|net|io|dev|sh|no|test|local|ai)\b/i;
/**
* A rooted path (`~/x`, `./x`, `/Users/x`) or a glob. Deliberately does NOT
* match a bare `word/word`: "pros/cons" is not a path, and treating it as one
* vetoed the single largest deletable block in the dogfood run. Real filenames
* are FILENAME_RE's job, and backticked paths are CODE_SPAN_RE's.
*/
const PATH_RE = /(?:^|[\s(«"'])(?:~|\.{1,2})?\/[\w.~/*-]+|\*\*?\//;
/** A concrete filename with a known extension (STATE.md, .zshenv, foo.sh). */
const FILENAME_RE =
/\b[\w.-]+\.(?:ts|tsx|js|jsx|mjs|cjs|py|md|json|ya?ml|toml|go|rs|java|rb|php|c|cpp|h|hpp|sh|sql|css|scss|html|env|template|lock)\b|(?:^|\s)\.\w+rc\b|\bzshenv\b/;
/**
* A version pin "bash 3.2", "Opus 4.8", "v5.12.5". A number that specific is
* a fact about this machine's world, not general knowledge.
*/
const VERSION_RE = /\bv?\d+\.\d+(?:\.\d+)?\b/;
/**
* §6.0's carve-out. These stay in the floor by decision: the downside is
* asymmetric and the lines are cheap. Do not let a "the model knows this now"
* argument reach them.
*/
const POLICY_RE =
/\b(?:secret|secrets|credential|credentials|password|passphrase|api[\s-]?key|access[\s-]?token|keychain|\.env|production|prod|force[\s-]push|rm\s+-rf|destructive|hemmelighet|passord|untrusted|injection|prompt[\s-]injection|angrepsflate|attack[\s-]surface|exfiltrat\w*)\b/i;
/**
* An unresolved local entity: a mixed-case capitalized word appearing
* mid-sentence. In config prose that is almost always a product, service or
* tool name "push til deres egne Forgejo-remotes", "Bruk Explore for søk"
* i.e. exactly the local vocabulary that makes a line underivable, but with no
* literal syntax for the other markers to key on.
*
* This marker is the CONSERVATIVE DEFAULT, and it is deliberately blunt: the
* mechanism cannot tell "Forgejo" from an ordinary capitalized word without a
* dictionary, so it declines to decide and keeps the block. Any finer rule
* (lowercase-form-appears-elsewhere, curated entity lists) is a proxy for a
* dictionary that would be tuned against one machine's config and fail silently
* on the next one. Paying in recall is the direction brief §6.0 mandates.
*
* All-caps tokens are exempt: this config's emphasis convention is ALDRI /
* ALLTID / FØR / , and acronyms like AI and TDD are generic, not local.
*
* "Mid-sentence" is keyed on a preceding LOWERCASE letter (or comma) not on
* "anything that is not a full stop". The looser version cost 4 of 11 deletable
* groups in the dogfood run by firing on `**Bold labels:**` and on quoted
* sentence starts (`"Som AI kan jeg ikke…"`), both of which are sentence
* openings dressed in punctuation rather than local vocabulary.
*/
const ENTITY_RE = /[a-zæøå,;]\s+(?![A-ZÆØÅ]{2,}\b)[A-ZÆØÅ][a-zæøå][\wæøåÆØÅ-]*/;
/** The ordered veto table — exported so a finding can cite *why* it was floored. */
export const FLOOR_MARKERS = Object.freeze([
{ name: 'code-span', re: CODE_SPAN_RE, why: 'contains an inline code literal' },
{ name: 'url', re: URL_RE, why: 'names a specific host or URL' },
{ name: 'filename', re: FILENAME_RE, why: 'names a concrete file' },
{ name: 'path', re: PATH_RE, why: 'names a concrete path' },
{ name: 'version', re: VERSION_RE, why: 'pins a specific version' },
{ name: 'policy', re: POLICY_RE, why: 'is a policy invariant (floor by decision, §6.0)' },
{ name: 'unresolved-entity', re: ENTITY_RE, why: 'names a capitalized entity the mechanism cannot resolve' },
]);
/**
* The first floor marker present in `text`, or null if the text carries no
* underivable local fact.
* @param {string} text
* @returns {{name:string, why:string}|null}
*/
export function floorMarker(text) {
const s = String(text == null ? '' : text);
for (const m of FLOOR_MARKERS) {
if (m.re.test(s)) return { name: m.name, why: m.why };
}
return null;
}
/**
* True when the block must never be proposed for deletion.
* @param {string} text
* @returns {boolean}
*/
export function isLoadBearing(text) {
return floorMarker(text) !== null;
}

View file

@ -0,0 +1,160 @@
/**
* Hook additionalContext injection advisory (v5.10 B5).
*
* A hook that emits `hookSpecificOutput.additionalContext` has that content
* injected into Claude's context EVERY time the hook fires. Plain stdout on
* exit 0 does NOT enter context (it goes to the debug log only). So a hook that
* dumps large, unfiltered command output into additionalContext is a recurring
* per-turn token cost that compounds and is compaction-sensitive.
*
* This module is a STATIC heuristic over hook SCRIPT SOURCE. It is deliberately
* LOW PRECISION it cannot run the script or measure the real payload so it
* ships as an INFO advisory (weight 0, never severity-bearing), paired with a
* feature-gap "filter-before-Claude-reads" lever. The signal:
*
* the script references `additionalContext`
* AND captures output from a verbose-prone command (cat/find/git log/test/curl)
* AND applies no truncating/filtering tool anywhere (grep/head/jq/.slice).
*
* A script that pipes through a filter, or only captures cheap output ($(date)),
* is assumed bounded and is not flagged. Conversely a verbose capture with no
* filter is the un-grepped pattern worth surfacing.
*
* Mechanism verified 2026-06-23 against code.claude.com/docs:
* context-window.md "A PostToolUse hook reports back via
* hookSpecificOutput.additionalContext. That field enters Claude's context.
* Plain stdout on exit 0 does not." + tip: keep output concise; it enters
* context without truncation. The remediation lever is the documented
* filter-test-output.sh pattern (grep before Claude reads).
*
* The pure `assessHookAdditionalContext` takes already-read script text so it is
* fully unit-testable without IO. `assessHookContextForRepo` is the thin IO
* wrapper (walk hooks scripts assess) shared by feature-gap; the HKV scanner
* calls the pure function inline on scripts it already reads.
*/
import { readTextFile } from './file-discovery.mjs';
import { parseJson } from './yaml-parser.mjs';
import { stat } from 'node:fs/promises';
import { resolve, dirname } from 'node:path';
// The field that actually enters Claude's context (vs. plain stdout / debug log).
const ADDITIONAL_CONTEXT_RX = /additionalContext/;
// Commands whose UNfiltered output can be large. A capture invoking one of these
// with no filter anywhere is the low-precision "un-grepped output" signal.
// Shell substitution AND node child_process / file reads are both covered.
const VERBOSE_CAPTURE_RX =
/\b(?:cat|find|ls|git\s+(?:log|diff|status|show)|npm|yarn|pnpm|pytest|jest|go\s+test|cargo\s+test|curl|wget|env|printenv|dmesg|journalctl|execSync|spawnSync|readFileSync)\b/;
// Truncating / filtering tools that BOUND a payload before it reaches context.
// Their presence anywhere in the script suppresses the advisory (assumed bounded).
// Shell filters + the common node-side bounding operations.
const FILTER_RX =
/\b(?:grep|egrep|rg|head|tail|sed|awk|jq|cut|wc|uniq|sort)\b|\.(?:slice|substring|substr)\s*\(/;
/**
* Assess one hook script's source for unfiltered additionalContext injection.
*
* @param {{ scriptContent?: string }} [args]
* @returns {{
* buildsAdditionalContext: boolean,
* hasVerboseCapture: boolean,
* hasFilter: boolean,
* capturesUnfiltered: boolean,
* flagged: boolean,
* }}
*/
export function assessHookAdditionalContext({ scriptContent } = {}) {
const content = typeof scriptContent === 'string' ? scriptContent : '';
const buildsAdditionalContext = ADDITIONAL_CONTEXT_RX.test(content);
const hasVerboseCapture = VERBOSE_CAPTURE_RX.test(content);
const hasFilter = FILTER_RX.test(content);
const capturesUnfiltered = hasVerboseCapture && !hasFilter;
return {
buildsAdditionalContext,
hasVerboseCapture,
hasFilter,
capturesUnfiltered,
flagged: buildsAdditionalContext && capturesUnfiltered,
};
}
/**
* Extract a filesystem script path from a hook command string.
* Mirrors hook-validator's extractScriptPath (kept local so the shared lib has
* no upward dependency on a scanner). Handles ${CLAUDE_PLUGIN_ROOT}.
*/
function extractScriptPath(command, baseDir) {
const match = command.match(/(?:bash|node|sh)\s+(.+?)(?:\s|$)/);
if (!match) return null;
let scriptPath = match[1].trim();
scriptPath = scriptPath.replace(/\$\{CLAUDE_PLUGIN_ROOT\}/g, resolve(baseDir, '..'));
scriptPath = scriptPath.replace(/\$CLAUDE_PLUGIN_ROOT/g, resolve(baseDir, '..'));
if (scriptPath.includes('$')) return null;
return resolve(baseDir, scriptPath);
}
/** Yield every command-hook { event, command } from a hooks object. */
function* iterateCommandHooks(hooks) {
if (!hooks || typeof hooks !== 'object' || Array.isArray(hooks)) return;
for (const [event, handlers] of Object.entries(hooks)) {
if (!Array.isArray(handlers)) continue;
for (const group of handlers) {
const hookList = group && Array.isArray(group.hooks) ? group.hooks : [];
for (const hook of hookList) {
if (hook && hook.type === 'command' && typeof hook.command === 'string') {
yield { event, command: hook.command };
}
}
}
}
}
/**
* IO wrapper: walk discovered hooks (hooks.json + settings.json hooks), resolve
* each command hook's script, and return the ones flagged by the heuristic.
* Shared by feature-gap (the HKV scanner assesses inline on scripts it reads).
*
* @param {{ files: import('./file-discovery.mjs').ConfigFile[] }} discovery
* @returns {Promise<Array<{ event: string, scriptPath: string, file: string,
* assessment: ReturnType<typeof assessHookAdditionalContext> }>>}
*/
export async function assessHookContextForRepo(discovery) {
const flagged = [];
const files = (discovery && Array.isArray(discovery.files)) ? discovery.files : [];
const hooksObjects = [];
for (const file of files.filter((f) => f.type === 'hooks-json')) {
const content = await readTextFile(file.absPath);
const parsed = content ? parseJson(content) : null;
if (parsed) hooksObjects.push({ hooks: parsed.hooks || parsed, file });
}
for (const file of files.filter((f) => f.type === 'settings-json')) {
const content = await readTextFile(file.absPath);
const parsed = content ? parseJson(content) : null;
if (parsed && parsed.hooks && !Array.isArray(parsed.hooks)) {
hooksObjects.push({ hooks: parsed.hooks, file });
}
}
for (const { hooks, file } of hooksObjects) {
const baseDir = dirname(file.absPath);
for (const { event, command } of iterateCommandHooks(hooks)) {
const scriptPath = extractScriptPath(command, baseDir);
if (!scriptPath) continue;
try {
await stat(scriptPath);
} catch {
continue; // missing script — HKV reports that separately
}
const scriptContent = await readTextFile(scriptPath);
const assessment = assessHookAdditionalContext({ scriptContent });
if (assessment.flagged) {
flagged.push({ event, scriptPath, file: file.absPath, assessment });
}
}
}
return flagged;
}

View file

@ -32,6 +32,11 @@ export const TRANSLATIONS = {
description: 'Without `CLAUDE.md` at your project root, Claude has to work out your conventions from scratch every conversation. Project-specific guidance is the single highest-impact thing you can add.',
recommendation: 'Create a file called `CLAUDE.md` in your project root. Start with a one-paragraph project overview, common commands, and any quirks Claude should know about.',
},
'Nested CLAUDE.md is not re-injected after compaction': {
title: 'A `CLAUDE.md` in a subfolder can quietly drop out mid-session',
description: 'Only the main project `CLAUDE.md` is restored when Claude trims older history. A `CLAUDE.md` in a subfolder loads when you open a file there, but it does not come back after a trim until you open one again.',
recommendation: 'If its guidance must always apply, move that part into the main `CLAUDE.md`. Keep the subfolder file for things only needed when working in that folder.',
},
'CLAUDE.md is nearly empty': {
title: 'Your `CLAUDE.md` is mostly empty',
description: 'An empty instructions file gives Claude no project-specific context, so behavior falls back to defaults.',
@ -91,10 +96,10 @@ export const TRANSLATIONS = {
// ─────────────────────────────────────────────────────────────
SET: {
static: {
'Unknown settings key': {
title: 'A settings key isn\'t recognized',
description: 'A key in your settings file isn\'t one Claude Code understands. It will be ignored.',
recommendation: 'Check the key name for typos, or remove the key if it\'s no longer in use.',
'Possible typo in settings key': {
title: 'A settings key looks like a typo',
description: 'A key in your settings file isn\'t recognized, but it\'s very close to a real one — likely a typo. Claude Code forwards unrecognized keys unchanged rather than rejecting them, so a misspelled key silently has no effect.',
recommendation: 'Check the suggested key name. Fix the spelling, or keep the key if it\'s intentional (e.g. a newer key this audit doesn\'t know yet).',
},
'Deprecated settings key': {
title: 'A settings key is no longer supported',
@ -142,7 +147,27 @@ export const TRANSLATIONS = {
recommendation: 'Open the file and fix the JSON syntax shown in the details (often a missing comma or quote).',
},
},
patterns: [],
patterns: [
{
// Specific case first: well-formed autoMode placed in the wrong scope.
regex: /^autoMode in shared project settings/,
translation: {
title: 'Your auto-mode rules are in a file Claude Code ignores',
description: 'Claude Code doesn\'t read `autoMode` from shared project settings (`.claude/settings.json`), so a checked-in repo can\'t grant itself auto-approval rules. The block has no effect where it is.',
recommendation: 'Move `autoMode` to your user settings (`~/.claude/settings.json`), the local (gitignored) project settings (`.claude/settings.local.json`), or managed settings.',
},
},
{
// Catch-all for autoMode structure problems (not-an-object, unknown
// sub-key, sub-key not a string array).
regex: /^autoMode/,
translation: {
title: 'Your `autoMode` block is malformed',
description: '`autoMode` must be an object with `environment`, `allow`, `soft_deny`, and `hard_deny`, each a list of plain-text rules. Part of it doesn\'t match that shape, so those rules may not apply.',
recommendation: 'Fix the `autoMode` entry shown in the details — use only the four known keys, each set to a list of strings.',
},
},
],
_default: {
title: 'Your settings file has an issue',
description: 'A check on your settings file flagged something worth a look.',
@ -234,10 +259,15 @@ export const TRANSLATIONS = {
description: 'Without scoping, the rule loads on every conversation regardless of which files you\'re working with.',
recommendation: 'Add a scoping block at the top of the file to limit when the rule loads (see the details).',
},
'Rule uses deprecated "globs" field': {
title: 'A rule uses an old field name',
description: 'The field was renamed; the old name still works for now but may stop working in a future release.',
recommendation: 'Rename the field to the current equivalent shown in the details.',
'Large path-scoped rule is lost after compaction': {
title: 'A large scoped rule can quietly drop out mid-session',
description: 'Scoped rules load only when you open a matching file, and they fall out of context when Claude trims older history — they do not come back until you open a matching file again.',
recommendation: 'If part of it must always apply, move that part into the main `CLAUDE.md`, which is restored automatically.',
},
'Rule uses "globs" instead of documented "paths"': {
title: 'A rule uses an unrecognized scoping field',
description: 'Claude Code\'s docs use `paths:` to scope a rule; `globs:` is not the documented field, so the rule may not scope the way you intend.',
recommendation: 'Rename the field to `paths:` (see the details).',
},
'Rule file is not .md': {
title: 'A rule file uses an unexpected extension',
@ -273,16 +303,6 @@ export const TRANSLATIONS = {
description: 'The `type` field doesn\'t match one Claude Code knows how to start (typically `stdio`, `sse`, or `http`).',
recommendation: 'Change the `type` to one of the supported values shown in the details.',
},
'Invalid trust level': {
title: 'A connected service has an unrecognized trust setting',
description: 'Trust controls whether Claude can use the service\'s tools without asking.',
recommendation: 'Set the trust value to one of the accepted ones (see details).',
},
'Missing trust level': {
title: 'A connected service has no trust setting',
description: 'Without an explicit trust value, Claude has to ask before each tool use, which slows your work.',
recommendation: 'Add a trust value to the entry. The details show the accepted values.',
},
'Unknown MCP server field': {
title: 'A connected service has an unrecognized setting',
description: 'The setting isn\'t one Claude Code reads, so it will be ignored.',
@ -415,12 +435,12 @@ export const TRANSLATIONS = {
recommendation: 'Consider moving team-wide settings to project scope and keeping personal ones at user or local scope.',
},
'CLAUDE.md not modular': {
title: 'Your instructions file is one big block',
description: 'Splitting long instructions into smaller linked files makes them easier to maintain and easier on the loading time.',
title: 'Your instructions all live in one file',
description: 'Splitting your instructions into smaller linked files with `@import` or `.claude/rules/` keeps each part focused and easier to maintain.',
recommendation: 'Break out long sections into separate files and link them with `@import`.',
},
'No path-scoped rules': {
title: 'Your rules all load on every conversation',
title: 'You haven\'t set up path-scoped rules yet',
description: 'Path-scoped rules only load when you\'re working with files that match — keeps each conversation focused.',
recommendation: 'Add scoping to your rules so they only load for the files they apply to.',
},
@ -470,7 +490,7 @@ export const TRANSLATIONS = {
recommendation: 'Add fields like `model`, `tools`, or `description` to your skill files where useful.',
},
'No subagent isolation': {
title: 'Your subagents share Claude\'s main work folder',
title: 'You haven\'t set up subagent isolation yet',
description: 'Isolated subagents run in their own copy of the repo so they can\'t accidentally disturb your main work.',
recommendation: 'Add `isolation: worktree` to subagents that do destructive or experimental work.',
},
@ -549,6 +569,11 @@ export const TRANSLATIONS = {
description: 'Skill descriptions load on every turn whether you use the skill or not. Long descriptions add up.',
recommendation: 'Trim the description to one short sentence and move details into the skill body.',
},
'Stale plugin-cache versions (disk cleanup, zero live-context impact)': {
title: 'Old plugin versions are sitting on disk (safe to delete)',
description: 'Your plugin cache holds older versions that newer installs have replaced. They take up disk space but are never loaded into a conversation — so they cost zero tokens per turn. This is housekeeping, not a performance problem.',
recommendation: 'Delete the old version folders to reclaim disk. The details list exactly which ones; the active version of each plugin stays untouched.',
},
},
patterns: [
{
@ -669,11 +694,6 @@ export const TRANSLATIONS = {
description: 'The settings block tells Claude what tools and model the agent should use.',
recommendation: 'Add a settings block (delimited by `---`) at the top of the file.',
},
'Cross-plugin command name conflict': {
title: 'Two plugins both define a command with the same name',
description: 'When two plugins use the same command name, only one wins.',
recommendation: 'Rename the command in one of the plugins, or disable the one you don\'t need.',
},
'No plugins found': {
title: 'No plugins are installed in this location',
description: 'The location was checked but contains no plugins (or no plugins Claude Code recognizes).',
@ -701,6 +721,14 @@ export const TRANSLATIONS = {
},
},
patterns: [
{
regex: /^Plugin agent sets ".+", which Claude Code ignores$/,
translation: {
title: 'A plugin agent sets a field Claude Code ignores',
description: 'Plugin subagents ignore the `hooks`, `mcpServers`, and `permissionMode` settings — only agents in `.claude/agents/` honor them. The field shown here has no effect, and `permissionMode` can give a false sense of restriction.',
recommendation: 'Remove the field, or move the agent into `.claude/agents/`, where it takes effect.',
},
},
{
regex: /^Missing required field in plugin\.json/,
translation: {
@ -733,6 +761,38 @@ export const TRANSLATIONS = {
recommendation: 'Add the missing setting shown in the details.',
},
},
{
regex: /^Command name ".+" used by multiple plugins$/,
translation: {
title: 'Several plugins define a command with the same name',
description: 'Each plugin\'s commands are namespaced (like `/plugin:command`), so they all still work — but a shared command name makes error messages, search results, and the command listing ambiguous about which plugin you mean.',
recommendation: 'Rename the command in one of the plugins so each name points to a single plugin.',
},
},
{
regex: /^Plugin namespace collision:/,
translation: {
title: 'Two plugins share the same namespace, so one hides the other',
description: 'Claude Code names a plugin\'s commands, skills, and agents after the plugin (like `/name:command`). When two plugins declare the same name, they share one namespace and only one is reachable — the other\'s commands, skills, and agents silently disappear.',
recommendation: 'Give each plugin a distinct `name` in its `plugin.json`. The folder name doesn\'t matter — the `name` field is what forms the namespace.',
},
},
{
regex: /^plugin\.json ".+" path shadows the default /,
translation: {
title: 'A plugin folder is silently ignored because the manifest points elsewhere',
description: 'The plugin\'s `plugin.json` points a component type (like commands or agents) at a custom path. When it does, Claude Code stops scanning the default folder of the same name — so everything still sitting in that folder silently disappears.',
recommendation: 'Either delete the unused default folder, or keep it by listing it explicitly alongside the custom path. The details show which folder and field.',
},
},
{
regex: /^plugin\.json "skills" entry /,
translation: {
title: 'A plugin lists a skill path that doesn\'t point to a real skill folder',
description: 'The plugin\'s `plugin.json` lists a `skills` path that isn\'t a folder inside the plugin — it\'s missing, points at a file, sits outside the plugin, or isn\'t text. Claude Code loads no skill from it.',
recommendation: 'Point each `skills` entry at a real folder inside the plugin that contains a `SKILL.md`. The details show which entry and what\'s wrong.',
},
},
],
_default: {
title: 'A plugin has a configuration issue',
@ -740,4 +800,84 @@ export const TRANSLATIONS = {
recommendation: 'See the details for what needs to change.',
},
},
// ─────────────────────────────────────────────────────────────
// SKL — Skill-Listing Budget
// Category: Wasted tokens
// ─────────────────────────────────────────────────────────────
SKL: {
static: {
'Skill description exceeds the listing cap (Claude Code truncates it)': {
title: 'A skill description is too long and gets cut off',
description: 'Claude Code only shows the first part of each skill\'s description when it decides which skill to use. This one runs past the limit, so the end — often the trigger phrases — is silently dropped.',
recommendation: 'Shorten the description so the trigger words come first and it fits under the limit. For skills you do not use, turning them off frees up the listing entirely.',
},
'Aggregate skill descriptions may exceed the listing budget': {
title: 'All your skills together may be too much for the listing',
description: 'Claude Code keeps every active skill\'s description in one shared listing it reads to choose which skill to use, and that listing has a limited size. Added up, your skills\' descriptions run past that size on a smaller setup, so Claude Code may drop some of them — and stop seeing those skills. This is an estimate; a larger setup has more room.',
recommendation: 'Free up room: turn off bundled skills you do not use, collapse the heaviest ones so only their names show, or shorten the longest descriptions. The details show the measured total and the room available.',
},
'Skill body is large (loads on demand when the skill runs)': {
title: 'A skill\'s body is large (it loads only when that skill runs)',
description: 'This skill\'s instructions run longer than the rough guidance for a skill body. The body is not part of the always-loaded listing Claude reads every turn — it loads only when you invoke the skill, so it costs nothing until then. Once it loads, though, it stays in context for the rest of that session.',
recommendation: 'Move reference material into supporting files the skill opens only when needed, so the body stays lean. For a heavy skill you can also run its body in a separate context with `context: fork` in the skill\'s settings.',
},
},
patterns: [],
_default: {
title: 'A skill is using more of the listing budget than it should',
description: 'A check on how much room your skills take in Claude Code\'s skill listing flagged something worth a look.',
recommendation: 'See the details for which skill to trim or turn off.',
},
},
// ─────────────────────────────────────────────────────────────
// OST — Output-Style Validation
// Category: Configuration mistake
// ─────────────────────────────────────────────────────────────
OST: {
static: {
'Custom output style removes built-in coding instructions': {
title: 'A custom output style turns off Claude\'s coding know-how',
description: 'This style replaces Claude Code\'s built-in coding guidance with only your own text, so while it\'s active Claude forgets how to scope changes, comment, and verify work. The setting that keeps that guidance is off by default.',
recommendation: 'Add `keep-coding-instructions: true` to the top of the style file to keep that guidance. If you meant to drop it for a non-coding style, leave it as is.',
},
'Plugin output style overrides your selected output style': {
title: 'A plugin is forcing its own output style on you',
description: 'This plugin applies its own output style automatically whenever it\'s on, replacing the one you picked. If two plugins both do this, the first one to load wins.',
recommendation: 'If you didn\'t want this, turn off the plugin or remove the force setting from its style. Otherwise there\'s nothing to do — the plugin works this way on purpose.',
},
'Configured output style does not exist': {
title: 'Your chosen output style can\'t be found',
description: 'Your settings point to an output style that doesn\'t exist by that name, so Claude Code quietly uses the default instead. The style you wanted never takes effect.',
recommendation: 'Check the spelling against your styles (built-ins are Default, Explanatory, Learning, Proactive), add the missing style file, or remove the setting.',
},
},
patterns: [],
_default: {
title: 'Something about your output styles needs a look',
description: 'A check on your output styles flagged something worth reviewing.',
recommendation: 'See the details for which output style to adjust.',
},
},
// ─────────────────────────────────────────────────────────────
// OPT — Optimization Lens (mechanism-fit)
// Category: Missed opportunity
// ─────────────────────────────────────────────────────────────
OPT: {
static: {
'A multi-step procedure in CLAUDE.md belongs in a skill': {
title: 'A long checklist in CLAUDE.md could be a skill instead',
description: 'Your CLAUDE.md has a multi-step procedure that loads on every turn, costing tokens whether or not you\'re doing that task. Procedures fit better as a skill, whose steps load only when you actually run them.',
recommendation: 'Move the steps into a skill under `.claude/skills/`. Keep CLAUDE.md for facts Claude should always know, not step-by-step procedures.',
},
},
patterns: [],
_default: {
title: 'Your setup could fit Claude Code a little better',
description: 'A check found a setup that works but where a different mechanism would fit the job better.',
recommendation: 'See the details for the suggested change.',
},
},
};

View file

@ -37,9 +37,24 @@ const SCANNER_TO_CATEGORY = {
COL: 'Conflict',
TOK: 'Wasted tokens',
CPS: 'Wasted tokens',
SKL: 'Wasted tokens',
AGT: 'Wasted tokens',
DIS: 'Dead config',
GAP: 'Missed opportunity',
PLH: 'Configuration mistake',
OST: 'Configuration mistake',
OPT: 'Missed opportunity',
};
/**
* Per-finding `category` values that override the scanner-default impact label.
* Needed when one finding inside a scanner means something different from the
* scanner's usual bucket e.g. stale plugin-cache versions are emitted by TOK
* (normally "Wasted tokens") but load on ZERO turns, so their honest impact is
* "Dead config" (present on disk, never loaded), not wasted per-turn tokens.
*/
const CATEGORY_TO_IMPACT = {
'plugin-cache-hygiene': 'Dead config',
};
/**
@ -122,7 +137,8 @@ export function humanizeFinding(finding) {
}
const translation = lookupTranslation(finding.scanner, finding.title);
const category = SCANNER_TO_CATEGORY[finding.scanner] || 'Other';
const category =
CATEGORY_TO_IMPACT[finding.category] || SCANNER_TO_CATEGORY[finding.scanner] || 'Other';
const action = SEVERITY_TO_ACTION[finding.severity] || 'FYI';
const relevance = computeRelevanceContext(finding.file);

View file

@ -0,0 +1,94 @@
/**
* knowledge-refresh deterministic freshness core for the best-practices register.
*
* The "living" half of the v5.7 living knowledge base (Chunk 3). This module is the
* PURE, deterministic part of the hybrid `/config-audit knowledge-refresh` motor: given a
* register and an injected reference date, it classifies each entry as `fresh` or `stale`
* by the age of its `source.verified` stamp. It NEVER touches the network and NEVER writes
* candidate discovery (polling CC changelog + Anthropic blog) and the human-approved
* writes live in the command layer (Verifiseringsplikt: no unverified claim is auto-written).
*
* `referenceDate` is injected (not read from the clock here) so the function is fully
* deterministic and unit-testable; the CLI passes today's date. See
* docs/v5.7-optimization-lens-plan.md.
*/
/** Default re-verify cadence: a confirmed best-practice older than this needs a re-check. */
export const STALE_AFTER_DAYS_DEFAULT = 90;
const DATE_RE = /^\d{4}-\d{2}-\d{2}$/;
const DAY_MS = 86_400_000;
/**
* Normalize a reference date (Date or YYYY-MM-DD string) to a UTC-midnight {iso, ms}.
* Throws TypeError on anything else the reference date is required and must be valid.
*/
function normalizeReferenceDate(value) {
let iso;
if (value instanceof Date) {
if (Number.isNaN(value.getTime())) throw new TypeError('referenceDate is an invalid Date');
iso = value.toISOString().slice(0, 10);
} else if (typeof value === 'string' && DATE_RE.test(value)) {
iso = value;
} else {
throw new TypeError('referenceDate must be a Date or a YYYY-MM-DD string');
}
const ms = Date.parse(`${iso}T00:00:00Z`);
if (Number.isNaN(ms)) throw new TypeError(`referenceDate is not a real calendar date: ${iso}`);
return { iso, ms };
}
/** Parse an entry's `source.verified` to UTC-midnight ms, or null if missing/unparseable. */
function verifiedMs(entry) {
const v = entry && entry.source && entry.source.verified;
if (typeof v !== 'string' || !DATE_RE.test(v)) return null;
const ms = Date.parse(`${v}T00:00:00Z`);
return Number.isNaN(ms) ? null : ms;
}
/**
* Classify every register entry as fresh or stale by the age of its source.verified stamp.
*
* @param {{entries:object[]}} register
* @param {{ referenceDate: string|Date, staleAfterDays?: number }} opts
* @returns {{
* referenceDate: string,
* staleAfterDays: number,
* stale: Array<{id:string, verified:string|undefined, ageDays:number|null, url:string|undefined, claim:string|undefined}>,
* fresh: Array<{id:string, verified:string|undefined, ageDays:number}>,
* counts: { total:number, stale:number, fresh:number }
* }}
*/
export function assessFreshness(register, opts = {}) {
const ref = normalizeReferenceDate(opts.referenceDate);
const staleAfterDays =
typeof opts.staleAfterDays === 'number' ? opts.staleAfterDays : STALE_AFTER_DAYS_DEFAULT;
const entries = (register && Array.isArray(register.entries)) ? register.entries : [];
const stale = [];
const fresh = [];
for (const e of entries) {
const verified = e && e.source ? e.source.verified : undefined;
const vms = verifiedMs(e);
if (vms === null) {
// No re-checkable date → needs attention. Stale with ageDays null.
stale.push({ id: e && e.id, verified, ageDays: null, url: e && e.source && e.source.url, claim: e && e.claim });
continue;
}
const ageDays = Math.floor((ref.ms - vms) / DAY_MS);
if (ageDays > staleAfterDays) {
stale.push({ id: e.id, verified, ageDays, url: e.source && e.source.url, claim: e.claim });
} else {
fresh.push({ id: e.id, verified, ageDays });
}
}
return {
referenceDate: ref.iso,
staleAfterDays,
stale,
fresh,
counts: { total: entries.length, stale: stale.length, fresh: fresh.length },
};
}

View file

@ -0,0 +1,114 @@
/**
* lens-prefilter deterministic, recall-oriented candidate generator for the
* v5.7 optimization lens (CA-OPT) hybrid motor, Chunk 2b.
*
* The OPT *scanner* (Chunk 2a) handles the one mechanism-fit case it can decide
* deterministically with high precision (a long numbered procedure skill). The
* other three cases in the register are PROSE-JUDGMENT calls whether a line is
* really lifecycle automation, a path-specific constraint, or an absolute
* prohibition depends on reading intent, which a regex cannot settle. So the
* hybrid motor splits the work:
*
* pre-filter (this module, CHEAP, recall-oriented)
* surfaces candidate lines tagged with the register rule they might fit
* opus optimization-lens-agent (PRECISION gate)
* reads each candidate in context, keeps only genuine mechanism-fit
* opportunities, cites the register rule + source
*
* Therefore this pre-filter deliberately errs toward recall: a false candidate
* costs the agent a moment's judgement, a missed line is never recoverable. It
* does, however, avoid the two obvious noise sources fenced code blocks and
* (when the caller passes the parsed body) YAML frontmatter and it requires an
* imperative-looking line for the path-specific class so plain "see docs/x.md"
* references don't flood the candidate list.
*
* The detector names mirror the `lensCheck` fields of the register entries
* (knowledge/best-practices.json), so the agent can map each candidate straight
* back to its provenance.
*
* Zero external dependencies. Pure: input text candidate array.
*/
/**
* The three prose-judgment detectors, keyed to their register entries.
* `mechanism` is the better-fit mechanism the register recommends.
*/
export const LENS_DETECTORS = Object.freeze([
{ lensCheck: 'claude-md-lifecycle-phrasing', registerId: 'BP-MECH-001', mechanism: 'hook' },
{ lensCheck: 'unscoped-path-specific-instruction', registerId: 'BP-MECH-002', mechanism: 'rule' },
{ lensCheck: 'never-instruction', registerId: 'BP-MECH-004', mechanism: 'permission' },
]);
// Lifecycle automation phrased as an instruction: "after every commit", "before
// each push", "every time you …", "whenever you …", "always run". Recall-first.
const LIFECYCLE_RE =
/\b(?:after (?:every|each)|before (?:every|each)|on (?:every|each)|every time|each time|always run|whenever)\b/i;
// Absolute prohibition: a standalone "never" followed by an action word. Kept
// permissive (recall); the agent decides whether it is a real hard rule.
const NEVER_RE = /\bnever\s+[a-z]/i;
// A concrete path / glob / known-extension filename anywhere in the line.
const PATH_RE =
/(?:(?:\.{0,2}\/)?[\w.-]+\/[\w.*/-]+|\*\*?\/[\w.*-]+|\b[\w-]+\.(?:ts|tsx|js|jsx|mjs|cjs|py|md|json|ya?ml|toml|go|rs|java|rb|php|c|cpp|h|hpp|sh|sql|css|scss|html|env)\b)/;
// An imperative / modal verb that marks a line as an instruction rather than a
// bare cross-reference. Gates the path-specific class to cut "see foo/bar.md".
const INSTRUCTION_RE =
/\b(?:use|edit|run|always|must|should|put|place|write|add|modify|update|format|lint|test|name|store|keep|never|generate|build|deploy|commit)\b/i;
const getDetector = (lensCheck) => LENS_DETECTORS.find((d) => d.lensCheck === lensCheck);
function candidate(lensCheck, lineNo, lineText) {
const d = getDetector(lensCheck);
return {
lensCheck,
registerId: d.registerId,
mechanism: d.mechanism,
line: lineNo,
text: lineText.trim(),
};
}
/**
* Scan CLAUDE.md text for prose-judgment mechanism-fit candidates.
*
* Pass the file body (frontmatter stripped) for clean line numbers; the caller
* is then responsible for offsetting `line` by the body's start line. Raw text
* also works fenced code is skipped either way.
*
* @param {string} text
* @returns {Array<{lensCheck:string, registerId:string, mechanism:string, line:number, text:string}>}
*/
export function prefilterClaudeMd(text) {
const lines = String(text == null ? '' : text).split('\n');
const out = [];
let inFence = false;
for (let i = 0; i < lines.length; i++) {
const raw = lines[i];
const lineNo = i + 1;
// Toggle fenced code blocks (``` or ~~~). Fence lines themselves are skipped.
if (/^\s*(?:```|~~~)/.test(raw)) {
inFence = !inFence;
continue;
}
if (inFence) continue;
const trimmed = raw.trim();
if (trimmed === '') continue;
if (LIFECYCLE_RE.test(raw)) {
out.push(candidate('claude-md-lifecycle-phrasing', lineNo, raw));
}
if (NEVER_RE.test(raw)) {
out.push(candidate('never-instruction', lineNo, raw));
}
if (PATH_RE.test(raw) && INSTRUCTION_RE.test(raw)) {
out.push(candidate('unscoped-path-specific-instruction', lineNo, raw));
}
}
return out;
}

View file

@ -0,0 +1,206 @@
/**
* MCP tool-schema deferral assessment (v5.10 B4).
*
* By default Claude Code DEFERS MCP tool schemas: only tool *names* enter the
* always-loaded prefix (~120 tokens total) and full schemas load on demand via
* tool search. Certain conditions force ALL full schemas into the always-loaded
* prefix instead paid on every turn whether or not a tool is used.
*
* This module is a STATIC, CONFIG-FILE assessment. It triggers only on signals
* that live in config files a scanner can read deterministically:
* - settings.json `env.ENABLE_TOOL_SEARCH` = "false" (HIGH confidence)
* - settings.json `permissions.deny` contains "ToolSearch" (HIGH confidence)
* - settings.json `model` is a Haiku model (MEDIUM confidence)
* - a per-server `.mcp.json` `alwaysLoad: true` (HIGH confidence)
*
* It deliberately does NOT read process.env shell variables. Tool search is also
* disabled on Vertex AI, with a custom ANTHROPIC_BASE_URL (non-first-party host),
* or after a runtime `/model` switch to Haiku but those are launch/runtime
* state, not config files, so a static scan cannot see them without becoming
* machine-dependent. They are disclosed (DEFERRAL_DISCLOSURE), never triggered.
*
* Mechanism verified 2026-06-23 against code.claude.com/docs:
* context-window.md (MCP tools deferred, ~120 tok), mcp.md#configure-tool-search
* + #exempt-a-server-from-deferral, costs.md (reduce MCP server overhead).
*
* The pure `assessMcpDeferral` takes already-parsed inputs so it is fully
* unit-testable without file IO. `assessMcpDeferralForRepo` is the thin IO
* wrapper shared by token-hotspots and feature-gap.
*/
import { resolve } from 'node:path';
import { readTextFile } from './file-discovery.mjs';
import { parseJson } from './yaml-parser.mjs';
import { readActiveMcpServers } from './active-config-reader.mjs';
// Aggregate forced-upfront schema cost (tokens) → severity ladder. MCP token
// estimates are base 500 + ~200/tool (active-config-reader estimateTokens), so
// these anchor on a couple of small servers (medium) vs a large one / several
// (high). Heuristic, not measured — disclosed in every finding.
export const FORCED_SCHEMA_TOKENS_MEDIUM = 1500;
export const FORCED_SCHEMA_TOKENS_HIGH = 5000;
// Appended to every CA-TOK-006 finding: the launch/runtime conditions a static
// config scan cannot see, so the user knows to check them manually.
export const DEFERRAL_DISCLOSURE =
'static config-file check: tool search is ALSO disabled (all MCP schemas forced ' +
'upfront) on Vertex AI, with a custom ANTHROPIC_BASE_URL (non-first-party host), ' +
'or after a runtime /model switch to a Haiku model — none visible to a static ' +
'scan, so verify at launch. Tool-level "anthropic/alwaysLoad" set server-side is ' +
'likewise invisible. Per-server alwaysLoad requires Claude Code v2.1.121+.';
/**
* Does a permissions.deny list disable the ToolSearch tool? Matches a bare
* "ToolSearch" tool name (with or without an argument suffix), same shape Claude
* Code uses for built-in tool denies.
*/
function denyListDisablesToolSearch(deny) {
if (!Array.isArray(deny)) return false;
return deny.some((entry) => {
if (typeof entry !== 'string') return false;
const tool = entry.replace(/\(.*\)$/, '').trim();
return tool === 'ToolSearch';
});
}
function isHaikuModel(model) {
return typeof model === 'string' && /haiku/i.test(model);
}
/**
* Map severity for the forced-upfront aggregate. High-confidence reasons scale
* with the token cost; medium-confidence reasons (inferred from a configured
* model) are capped at medium so the finding never overstates certainty.
*
* @param {number} aggregateTokens
* @param {'high'|'medium'} confidence
* @returns {'high'|'medium'|'low'}
*/
export function severityForForcedSchemas(aggregateTokens, confidence) {
const tok = typeof aggregateTokens === 'number' ? aggregateTokens : 0;
if (confidence === 'high') {
if (tok >= FORCED_SCHEMA_TOKENS_HIGH) return 'high';
if (tok >= FORCED_SCHEMA_TOKENS_MEDIUM) return 'medium';
return 'low';
}
// medium confidence: cap at medium
if (tok >= FORCED_SCHEMA_TOKENS_HIGH) return 'medium';
return 'low';
}
/**
* Assess whether MCP tool schemas are forced into the always-loaded prefix.
*
* @param {object} args
* @param {{ env?: object, permissions?: { deny?: string[] }, model?: string }} [args.settings]
* merged settings.json view (env block, permissions, model).
* @param {Array<{name:string, source?:string, enabled?:boolean, toolCount?:number,
* estimatedTokens?:number, alwaysLoad?:boolean}>} [args.mcpServers]
* @returns {{
* toolSearchDisabled: boolean, reason: string|null,
* confidence: 'high'|'medium'|null, thresholdMode: boolean,
* alwaysLoadServers: object[], affectedServers: object[],
* aggregateTokens: number, forcedUpfront: boolean,
* }}
*/
export function assessMcpDeferral({ settings = {}, mcpServers = [] } = {}) {
const s = settings || {};
const tsRaw = s.env && typeof s.env === 'object' ? s.env.ENABLE_TOOL_SEARCH : undefined;
const ts = String(tsRaw ?? '').trim().toLowerCase();
const denyTS = denyListDisablesToolSearch(s.permissions?.deny);
const haiku = isHaikuModel(s.model);
let toolSearchDisabled = false;
let reason = null;
let confidence = null;
let thresholdMode = false;
if (ts === 'false') {
toolSearchDisabled = true;
reason = 'enable-tool-search-false';
confidence = 'high';
} else if (denyTS) {
toolSearchDisabled = true;
reason = 'deny-tool-search';
confidence = 'high';
} else if (haiku) {
// Haiku lacks tool_reference support, so tool search cannot run even when
// ENABLE_TOOL_SEARCH=true. Medium confidence: the configured model can be
// switched at runtime (/model), which a static scan cannot observe.
toolSearchDisabled = true;
reason = 'haiku-model';
confidence = 'medium';
} else if (ts === 'true') {
toolSearchDisabled = false;
} else if (ts.startsWith('auto')) {
// Threshold mode (auto / auto:N): schemas load upfront only when they fit a
// percentage of the context window — not a clear always-load. Info, not a
// forced-upfront trigger.
thresholdMode = true;
}
const active = (Array.isArray(mcpServers) ? mcpServers : []).filter(
(m) => m && m.enabled !== false,
);
const alwaysLoadServers = active.filter((m) => m.alwaysLoad === true);
const affectedServers = toolSearchDisabled ? active : alwaysLoadServers;
const aggregateTokens = affectedServers.reduce(
(sum, m) => sum + (typeof m.estimatedTokens === 'number' ? m.estimatedTokens : 0),
0,
);
return {
toolSearchDisabled,
reason,
confidence,
thresholdMode,
alwaysLoadServers,
affectedServers,
aggregateTokens,
forcedUpfront: affectedServers.length > 0,
};
}
/**
* Merge project + local settings.json for the deferral check. Scoped to the
* audited path (NOT the user cascade) so the result stays deterministic and free
* of ambient HOME leakage mirrors Pattern G's project-local scoping. Local
* overrides project; env/permissions shallow-merge, deny lists concatenate.
*/
async function readMergedProjectSettings(repoPath) {
const merged = { env: {}, permissions: {}, model: undefined };
for (const rel of ['.claude/settings.json', '.claude/settings.local.json']) {
const content = await readTextFile(resolve(repoPath, rel));
if (!content) continue;
const parsed = parseJson(content);
if (!parsed || typeof parsed !== 'object') continue;
if (parsed.env && typeof parsed.env === 'object') Object.assign(merged.env, parsed.env);
if (parsed.permissions && typeof parsed.permissions === 'object') {
const deny = [
...(Array.isArray(merged.permissions.deny) ? merged.permissions.deny : []),
...(Array.isArray(parsed.permissions.deny) ? parsed.permissions.deny : []),
];
merged.permissions = { ...merged.permissions, ...parsed.permissions, deny };
}
if (typeof parsed.model === 'string') merged.model = parsed.model;
}
return merged;
}
/**
* IO wrapper: assess MCP deferral for a repo path. Scopes MCP servers to active
* project-local `.mcp.json` (plugin / ~/.claude.json servers are the manifest's
* concern). Pass `mcpServers` (e.g. an already-loaded activeConfig.mcpServers) to
* avoid a re-read; otherwise they are read via readActiveMcpServers.
*
* @param {string} repoPath
* @param {{ mcpServers?: object[] }} [opts]
*/
export async function assessMcpDeferralForRepo(repoPath, opts = {}) {
const all = Array.isArray(opts.mcpServers)
? opts.mcpServers
: await readActiveMcpServers(repoPath);
const projectLocal = all.filter((m) => m && m.enabled && m.source === '.mcp.json');
const settings = await readMergedProjectSettings(repoPath);
return assessMcpDeferral({ settings, mcpServers: projectLocal });
}

View file

@ -0,0 +1,186 @@
/**
* Permission rule matching shared by the DIS scanner (dead-allow detection)
* and the CNF conflict-detector (cross-scope allow/deny conflicts).
*
* Claude Code permission rules are either bare (`Tool`) or param-qualified
* (`Tool(param)`): `Bash`, `Bash(npm:*)`, `Agent(model:opus)`,
* `WebFetch(domain:*)`. CC 2.1.178 made `Tool(param:value)` matching
* meaningful and 2.1.172 added `domain:` rules, so rule identity must be
* param-aware `Agent(model:opus)` and `Agent(model:sonnet)` are DISTINCT,
* not "the same tool".
*
* Two distinct predicates are exported because the scanners ask different
* questions:
* - DIS asks "is this allow entry dead?" does some deny fully COVER it
* (`dominates`). A bare allow survives a specific deny.
* - CNF asks "do these two cross-scope rules conflict?" do their match
* sets INTERSECT (`rulesIntersect`). A bare allow conflicts with a
* specific deny, because the denied case is a real contradiction.
*
* Zero external dependencies.
*/
/**
* Split a permission entry into its bare tool name and optional param.
* `Bash` { tool: 'Bash', param: null }
* `Bash(npm:*)` { tool: 'Bash', param: 'npm:*' }
* @param {string} entry
* @returns {{ tool: string|null, param: string|null }}
*/
export function parseRule(entry) {
if (typeof entry !== 'string') return { tool: null, param: null };
const idx = entry.indexOf('(');
if (idx === -1) return { tool: entry.trim() || null, param: null };
const tool = entry.slice(0, idx).trim();
let param = entry.slice(idx + 1).trim();
if (param.endsWith(')')) param = param.slice(0, -1).trim();
return { tool: tool || null, param };
}
/**
* Glob match a permission param against a concrete value. `*` matches any run
* of characters; everything else is literal. Anchored (full-string) match.
* @param {string} pattern
* @param {string} value
* @returns {boolean}
*/
export function paramMatches(pattern, value) {
if (typeof pattern !== 'string' || typeof value !== 'string') return false;
if (pattern === value) return true;
const rx = '^' + pattern
.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')
.replace(/\\\*/g, '.*') + '$';
return new RegExp(rx).test(value);
}
/**
* Does the deny entry fully cover the allow entry, making the allow dead config?
* Used by DIS. Bare deny and the equivalent `Tool(*)` deny-all glob covers
* everything (including a bare allow); a specific deny only covers the matching
* (or wildcard-subsumed) param and does NOT kill a bare allow.
*
* CC: "`Bash(*)` is equivalent to `Bash` ... As a deny rule, both forms remove
* the tool from Claude's context." (code.claude.com/docs/en/permissions)
* @param {string} denyEntry
* @param {string} allowEntry
* @returns {boolean}
*/
export function dominates(denyEntry, allowEntry) {
const d = parseRule(denyEntry);
const a = parseRule(allowEntry);
if (!d.tool || !a.tool || d.tool !== a.tool) return false;
if (d.param === null || d.param === '*') return true; // bare / Tool(*) deny covers all params
if (a.param === null) return false; // specific deny does not kill a bare allow
if (d.param === a.param) return true;
return paramMatches(d.param, a.param); // wildcard deny covers a matching literal allow
}
/**
* Do two permission rules' match sets intersect (genuine cross-scope conflict)?
* Used by CNF. A bare rule intersects any same-tool rule; param rules intersect
* when equal or when either wildcard-matches the other.
* @param {string} ruleA
* @param {string} ruleB
* @returns {boolean}
*/
export function rulesIntersect(ruleA, ruleB) {
const a = parseRule(ruleA);
const b = parseRule(ruleB);
if (!a.tool || !b.tool || a.tool !== b.tool) return false;
if (a.param === null || b.param === null) return true; // bare ∩ anything (same tool)
if (a.param === b.param) return true;
return paramMatches(a.param, b.param) || paramMatches(b.param, a.param);
}
/**
* Is this `permissions.allow` entry an UNANCHORED tool-name glob that Claude
* Code silently skips? CC accepts tool-name globs in an ALLOW rule only after a
* literal `mcp__<server>__` prefix (the server segment must be glob-free).
* Unanchored globs like `*`, `B*`, or `mcp__*` are skipped with a warning and
* auto-approve nothing dead config the author believes is granting access.
*
* Tool-name globs apply to the bare-name form only (no `(...)` specifier); a
* glob INSIDE a specifier such as `Bash(npm run *)` is normal and valid.
*
* CC: "An unanchored allow glob such as `"*"`, `"B*"`, or `"mcp__*"` is skipped
* with a warning and does not auto-approve anything."
* (code.claude.com/docs/en/permissions "Tool name wildcards")
*
* NOTE: deny/ask rules DO accept tool-name globs, so this predicate is for the
* allow list only.
* @param {string} entry
* @returns {boolean}
*/
export function isIneffectiveAllowGlob(entry) {
if (typeof entry !== 'string') return false;
if (entry.includes('(')) return false; // specifier form — glob lives inside the param
if (!entry.includes('*')) return false; // no glob — a normal match-all-tool allow
if (entry.startsWith('mcp__')) {
const rest = entry.slice(5); // after 'mcp__'
const sep = rest.indexOf('__');
// mcp__<server>__<rest> with a glob-free server segment is anchored & valid
if (sep > 0 && !rest.slice(0, sep).includes('*')) return false;
}
return true;
}
/**
* Tools whose canonicalizing input field collides with `Tool(param:value)`
* matching. CC ignores a rule whose param key is the tool's own field and
* emits a startup warning, because the rule would be bypassable (e.g. a
* compound command defeats `Bash(command:rm *)`).
*
* CC: "Fields that a tool already matches with its own canonicalizing rules are
* not matchable this way: `command` for Bash and PowerShell, `file_path` for
* Read, Edit, and Write, `path` for Grep and Glob, `notebook_path` for
* NotebookEdit, and `url` for WebFetch."
* (code.claude.com/docs/en/permissions "Match by input parameter")
*/
const FORBIDDEN_PARAMS = Object.freeze({
Bash: 'command',
PowerShell: 'command',
Read: 'file_path',
Edit: 'file_path',
Write: 'file_path',
Grep: 'path',
Glob: 'path',
NotebookEdit: 'notebook_path',
WebFetch: 'url',
});
/** Correct specifier syntax to suggest in place of the forbidden param form. */
const FORBIDDEN_PARAM_HINT = Object.freeze({
Bash: 'Bash(rm *)',
PowerShell: 'PowerShell(Remove-Item *)',
Read: 'Read(./path)',
Edit: 'Edit(/src/**)',
Write: 'Write(/src/**)',
Grep: 'a Read rule (covers Grep)',
Glob: 'a Read rule (covers Glob)',
NotebookEdit: 'Edit(/notebooks/**)',
WebFetch: 'WebFetch(domain:host)',
});
/**
* Is this entry a `Tool(param:value)` rule whose param KEY is the tool's own
* canonicalizing field? CC silently ignores these (any list) and emits a
* startup warning. Returns `{ tool, key, hint }` or `null`.
*
* Only the `param:value` form (a colon present) is forbidden `Bash(command)`
* is a literal command-prefix match and stays valid. The key must equal the
* tool's forbidden field, so `Bash(npm:*)`, `WebFetch(domain:x)`, and
* `Agent(model:opus)` are NOT flagged.
* @param {string} entry
* @returns {{ tool: string, key: string, hint: string }|null}
*/
export function forbiddenParamRule(entry) {
const { tool, param } = parseRule(entry);
if (!tool || param === null) return null;
const forbidden = FORBIDDEN_PARAMS[tool];
if (!forbidden) return null;
const colon = param.indexOf(':');
if (colon === -1) return null; // no `param:value` — literal specifier, valid
const key = param.slice(0, colon).trim();
if (key !== forbidden) return null;
return { tool, key, hint: FORBIDDEN_PARAM_HINT[tool] };
}

View file

@ -165,8 +165,12 @@ const SCANNER_AREA_MAP = {
GAP: 'Feature Coverage',
TOK: 'Token Efficiency',
CPS: 'Token Efficiency',
SKL: 'Token Efficiency',
AGT: 'Token Efficiency',
DIS: 'Settings',
COL: 'Plugin Hygiene',
OST: 'Settings',
OPT: 'CLAUDE.md',
};
/**

View file

@ -0,0 +1,199 @@
/**
* Skill-listing budget single source of truth.
*
* Claude Code shows the model a listing of every active skill's `description`
* so it can decide which skill to invoke. That listing is budgeted two ways:
* - per description: capped at 1,536 chars (CC 2.1.105, changelog L1502);
* anything past the cap is silently truncated.
* - in aggregate: the whole listing is allotted ~2% of the context window
* (CC 2.1.32, changelog L2860). We do NOT know the user's context window,
* so the aggregate budget anchors on a conservative 200k window (4,000 tok)
* and discloses the assumption.
*
* Two scanners consume this module so the budget is defined in exactly one place:
* - SKL (skill-listing-scanner) DIAGNOSES overflow (CA-SKL-001 per-description
* cap, CA-SKL-002 aggregate).
* - GAP (feature-gap-scanner) PRESCRIBES the remedy: when the listing is over
* budget and `disableBundledSkills` is un-pulled, it recommends that lever.
*
* Zero external dependencies.
*/
import { join } from 'node:path';
import { estimateTokens, enumeratePlugins, enumerateSkills } from './active-config-reader.mjs';
import { readTextFile } from './file-discovery.mjs';
import { parseFrontmatter, parseJson } from './yaml-parser.mjs';
import { CONTEXT_WINDOW_ANCHOR, LARGE_CONTEXT_WINDOW, withCommas } from './context-window.mjs';
// Verified per-description skill-listing cap (CC 2.1.105, changelog L1502).
// Descriptions longer than this are truncated in the listing the model sees.
export const DESCRIPTION_CAP = 1536;
// Aggregate listing budget (CC 2.1.32, changelog L2860): the skill listing the
// model reads is allotted ~2% of the context window. The context window is
// unknown, so we anchor on a conservative 200k window — the smallest common
// size, which fires earliest — and disclose the assumption in the evidence.
// The 200k/1M window constants live in context-window.mjs (single source of
// truth, shared with the CML CLAUDE.md char-budget check); re-exported here so
// existing importers of this module keep working.
export const BUDGET_FRACTION = 0.02;
export const AGGREGATE_BUDGET_TOKENS = Math.round(BUDGET_FRACTION * CONTEXT_WINDOW_ANCHOR); // 4000
export const LARGE_CONTEXT_BUDGET_TOKENS = Math.round(BUDGET_FRACTION * LARGE_CONTEXT_WINDOW); // 20000
export { CONTEXT_WINDOW_ANCHOR, LARGE_CONTEXT_WINDOW, withCommas };
// The honest framing required because the budget depends on a context window we
// cannot observe (jf. TOK CALIBRATION_NOTE). Appended to budget-overflow evidence.
export const BUDGET_CALIBRATION_NOTE =
'the budget scales with the context window - this anchors on a conservative 200k ' +
`window; at ${withCommas(LARGE_CONTEXT_WINDOW)} context the budget is ~${withCommas(LARGE_CONTEXT_BUDGET_TOKENS)} ` +
'tok and you are likely within it. this is an estimate, not measured telemetry';
// Skill-body size guidance (CA-SKL-003). A SKILL.md body over ~5,000 tokens
// (~500 lines / ~20k chars) should split reference content into supporting files
// (Claude Code skill-authoring guidance). Unlike the listing budget above, the
// body is an ON-DEMAND cost: it loads only when the skill is invoked, not every
// turn — so this is a LOW-severity efficiency signal, not an always-loaded bill.
export const BODY_TOKEN_THRESHOLD = 5000;
// Honest framing for the body-size finding: distinguishes on-demand from
// always-loaded cost and flags the figure as an estimate. Appended to evidence.
export const BODY_CALIBRATION_NOTE =
'this is the skill BODY (SKILL.md below the frontmatter), which loads ON DEMAND only when the ' +
'skill is invoked - NOT every turn like the always-loaded listing. estimate (chars/4), not measured telemetry';
/**
* @typedef {object} BudgetAssessment
* @property {number} scanned - number of descriptions assessed
* @property {number} aggregateChars - sum of each length capped at DESCRIPTION_CAP
* @property {number} aggregateTokens - estimateTokens(aggregateChars, 'markdown')
* @property {number} budgetTokens - AGGREGATE_BUDGET_TOKENS (the 200k-anchored budget)
* @property {boolean} overBudget - aggregateTokens strictly greater than budgetTokens
* @property {number} overBy - tokens over budget (0 when not over)
*/
/**
* Pure aggregate-budget assessment. Each description contributes only up to the
* cap (the tail past the cap is dropped from the listing, and CA-SKL-001 already
* flags it so the aggregate does not double-count it).
*
* @param {number[]} descLengths - one entry per active skill (description char count)
* @param {number} [budgetTokens=AGGREGATE_BUDGET_TOKENS] - the listing budget to
* measure against. Defaults to the 200k-anchored 4,000 tok; B8 passes a
* window-calibrated budget. Defaulting keeps existing callers byte-stable.
* @returns {BudgetAssessment}
*/
export function assessSkillListingBudget(descLengths, budgetTokens = AGGREGATE_BUDGET_TOKENS) {
let aggregateChars = 0;
for (const len of descLengths) {
const safe = (typeof len === 'number' && Number.isFinite(len) && len > 0) ? len : 0;
aggregateChars += Math.min(safe, DESCRIPTION_CAP);
}
const aggregateTokens = estimateTokens(aggregateChars, 'markdown');
const overBudget = aggregateTokens > budgetTokens;
return {
scanned: descLengths.length,
aggregateChars,
aggregateTokens,
budgetTokens,
overBudget,
overBy: overBudget ? aggregateTokens - budgetTokens : 0,
};
}
/**
* @typedef {object} ActiveSkillEntry
* @property {string} name
* @property {'user'|'plugin'} source
* @property {string|null} pluginName
* @property {string} path
* @property {number} descLength
* @property {number} bodyChars - SKILL.md body length below the frontmatter (on-demand cost)
* @property {number} bodyLines - body line count
* @property {number} bodyTokens - estimateTokens(bodyChars, 'markdown')
*/
/**
* Enumerate every active skill (user + plugin) and measure the listing budget.
* HOME-scoped: resolves ~/.claude via process.env.HOME (enumeratePlugins /
* enumerateSkills). Callers that run under test MUST override HOME (see the
* hermetic-home helper / runScannerWithHome pattern).
*
* @param {number} [budgetTokens=AGGREGATE_BUDGET_TOKENS] - listing budget for the
* aggregate assessment (B8 window-calibration); defaults keep callers byte-stable.
* @returns {Promise<{ skills: ActiveSkillEntry[], aggregate: BudgetAssessment }>}
*/
export async function measureActiveSkillListing(budgetTokens = AGGREGATE_BUDGET_TOKENS) {
const plugins = await enumeratePlugins();
const allSkills = await enumerateSkills(plugins);
const skills = [];
for (const skill of allSkills) {
if (!skill || typeof skill.path !== 'string') continue;
const content = await readTextFile(skill.path);
if (!content) continue;
const parsed = parseFrontmatter(content);
const fm = parsed?.frontmatter || null;
const desc = (fm && typeof fm.description === 'string') ? fm.description : '';
const body = (parsed && typeof parsed.body === 'string') ? parsed.body : '';
const bodyChars = body.length;
skills.push({
name: skill.name,
source: skill.source,
pluginName: skill.pluginName,
path: skill.path,
descLength: desc.length,
bodyChars,
bodyLines: bodyChars === 0 ? 0 : body.split('\n').length,
bodyTokens: estimateTokens(bodyChars, 'markdown'),
});
}
const aggregate = assessSkillListingBudget(skills.map((s) => s.descLength), budgetTokens);
return { skills, aggregate };
}
/**
* Read an env flag, treating null, "", "0", "false", "no", "off" as un-set.
* @param {string|undefined} v
* @returns {boolean}
*/
export function envFlag(v) {
if (v == null) return false;
const s = String(v).trim().toLowerCase();
return s !== '' && s !== '0' && s !== 'false' && s !== 'no' && s !== 'off';
}
/**
* Resolve whether the `disableBundledSkills` lever is effectively ON, reading the
* env var and the settings cascade directly (user ~/.claude, then project, then
* project-local).
*
* Reads the files directly rather than relying on config-discovery
* classification: when discovery walks ~/.claude from the .claude root, the
* user settings.json has a relPath of "settings.json" (no ".claude" segment)
* and is NOT tagged as settings-json so the dominant user-scope location for
* this global preference would otherwise be missed. HOME-scoped via
* process.env.HOME.
*
* @param {string} [projectPath] - project root, to also read project + local settings
* @returns {Promise<boolean>}
*/
export async function isBundledSkillsDisabled(projectPath) {
if (envFlag(process.env.CLAUDE_CODE_DISABLE_BUNDLED_SKILLS)) return true;
const home = process.env.HOME || process.env.USERPROFILE || '';
const candidates = [];
if (home) candidates.push(join(home, '.claude', 'settings.json'));
if (projectPath) {
candidates.push(join(projectPath, '.claude', 'settings.json'));
candidates.push(join(projectPath, '.claude', 'settings.local.json'));
}
for (const p of candidates) {
const content = await readTextFile(p);
if (!content) continue;
const parsed = parseJson(content);
if (parsed && parsed.disableBundledSkills === true) return true;
}
return false;
}

View file

@ -43,6 +43,41 @@ export function isSimilar(a, b, threshold = 0.8) {
return similarity >= threshold;
}
/**
* Levenshtein edit distance between two strings (insertions, deletions,
* substitutions; a transposition counts as 2). Used for typo detection on
* settings keys. Zero external dependencies, O(a*b) with two rolling rows.
* @param {string} a
* @param {string} b
* @returns {number}
*/
export function levenshtein(a, b) {
if (a === b) return 0;
const al = a.length;
const bl = b.length;
if (al === 0) return bl;
if (bl === 0) return al;
let prev = new Array(bl + 1);
let curr = new Array(bl + 1);
for (let j = 0; j <= bl; j++) prev[j] = j;
for (let i = 1; i <= al; i++) {
curr[0] = i;
const ac = a.charCodeAt(i - 1);
for (let j = 1; j <= bl; j++) {
const cost = ac === b.charCodeAt(j - 1) ? 0 : 1;
curr[j] = Math.min(
prev[j] + 1, // deletion
curr[j - 1] + 1, // insertion
prev[j - 1] + cost, // substitution
);
}
const tmp = prev;
prev = curr;
curr = tmp;
}
return prev[bl];
}
/**
* Extract all key-like patterns from a settings.json or similar config.
* @param {object} obj

View file

@ -0,0 +1,350 @@
/**
* subtraction-prefilter deterministic candidate generator for the v5.13
* subtraction lens (`/config-audit optimize --subtract`, BP-SUB-001).
*
* Every other command in this plugin asks an ADDITION question what could you
* add, what would fit a better mechanism, how expensive is what you have. This
* module asks the inverse: **what is no longer earning its always-loaded rent?**
*
* It is the mirror image of `lens-prefilter` in one important way. That module
* is recall-first, because a false candidate only costs the judge a moment's
* thought. Here a false candidate is a proposal to DELETE something, so the
* polarity flips: precision-first, and a hard deterministic floor
* (`floor-exclusion`) that the judge is not allowed to override.
*
* ## Granularity: leaf blocks
*
* The one design choice the hand-built fasit deliberately left open. A block is
* one markdown *leaf*: a list item including its wrapped continuation lines, or
* a paragraph. Headings, table rows and fenced code are structural, never
* candidates.
*
* Both halves of that choice are load-bearing, and the fasit tests both:
* - It must SPLIT. A numbered list whose steps 23 are local facts and whose
* steps 1 and 4 are filler is a mixed block; section granularity would have
* to keep or drop all four.
* - It must NOT split further. A bullet's load-bearing literal often sits on a
* wrapped continuation line ("…— `coord-send` er mekanismen"). A
* line-granular mechanism severs the first line from the fact that protects
* it and proposes a floor block for deletion the exact failure the gate
* exists to prevent.
*
* ## Two independent guarantees, not one
*
* A load-bearing block fails to become a candidate for either of two reasons,
* and both are needed:
* 1. it is *declarative* "Language: Norwegian for dialogue" states a fact
* about the human and corrects no behaviour, so no detector fires; or
* 2. `floor-exclusion` vetoes it for carrying an underivable local literal.
* Group 1 never reaches the veto at all, which is why the contract is asserted
* on the candidate list rather than on either mechanism alone.
*
* Norwegian and English are both first-class: the config this was designed
* against is Norwegian prose carrying English identifiers.
*
* Zero external dependencies. Pure: input text candidate array.
*/
import { floorMarker } from './floor-exclusion.mjs';
/**
* The subtraction detector, kept in its OWN table. `LENS_DETECTORS` drives the
* plain `optimize` payload's register block, and the subtraction axis must not
* fire on a plain run it asks a different question and the operator has to
* opt into it with `--subtract`.
*/
export const SUBTRACT_DETECTORS = Object.freeze([
{ lensCheck: 'compensatory-instruction', registerId: 'BP-SUB-001', mechanism: 'deletion' },
]);
/**
* Absolute / insistent phrasing. An instruction that has to shout is usually
* correcting behaviour rather than stating a fact.
*/
/**
* Word boundaries that understand æ/ø/å.
*
* JavaScript's `\b` is ASCII-only, so `/\bunngå\b/` never matches "unngå "
* the trailing "å" is not a word character, so there is no boundary after it.
* Every Norwegian keyword ending in æ/ø/å was silently dead until the dogfood
* run surfaced it. Do not reintroduce `\b` around this vocabulary.
*/
const LB = '(?<![\\wæøåÆØÅ])';
const RB = '(?![\\wæøåÆØÅ])';
const ABSOLUTE_RE = new RegExp(
LB +
'(?:never|always|avoid|don\'t|do not|must not|ensure|remember to|make sure|' +
'aldri|alltid|unngå|husk|sørg for|ikke)' +
RB,
'i',
);
/**
* Imperative verbs the grammatical signature of telling the model how to
* behave. Matched anywhere in the block, since Norwegian list prose puts them
* after a colon ("…oppgaver: forstå problemet, vurder alternativer").
*
* Word boundaries matter more than the list length: `\bdocument\b` must not
* match "documentation", or the declarative language-preference fact a floor
* block with no local literal to veto it would become a deletion candidate.
*/
const IMPERATIVE_RE = new RegExp(
LB +
'(?:' +
// English
'think|write|test|commit|use|read|check|ask|verify|stop|start|summarize|' +
'wait|match|change|fix|present|identify|keep|prefer|declare|refactor|' +
'document|explain|split|review|' +
// Norwegian
'tenk|skriv|test|commit|bruk|les|sjekk|spør|verifiser|dokumenter|stopp|' +
'start|oppsummer|vent|gjør|match|endre|fiks|presenter|identifiser|forstå|' +
'vurder|hold|siter|jobb|gjett|push|del|sett|forklar|utfør|følg' +
')' +
RB,
'i',
);
/**
* Minimum words for a block whose ONLY signal is an absolute marker. A bare
* "Haiku: aldri." is a declarative policy fact wearing the word "aldri", not an
* instruction about how to behave the same reason "Tone: direct and technical"
* never fires. Blocks carrying a real imperative verb are exempt from the floor,
* so "Test inkrementelt" still surfaces at two words.
*/
const ABSOLUTE_ONLY_MIN_WORDS = 6;
const HEADING_RE = /^\s*#{1,6}\s/;
const TABLE_RE = /^\s*\|/;
const FENCE_RE = /^\s*(?:```|~~~)/;
const LIST_ITEM_RE = /^\s*(?:[-*+]\s+|\d+[.)]\s+)/;
const ORDERED_ITEM_RE = /^\s*\d+[.)]\s+/;
const CONTINUATION_RE = /^\s+\S/;
/** A paragraph that introduces the list beneath it ("…tre lag med hver sin ene jobb:"). */
const STEM_RE = /:\s*$/;
/**
* Split markdown into leaf blocks.
*
* @param {string} text
* @returns {Array<{startLine:number, endLine:number, text:string, type:string}>}
* `type` is one of paragraph | list-item | heading | table | code.
*/
export function splitLeafBlocks(text) {
const lines = String(text == null ? '' : text).split('\n');
const blocks = [];
let current = null;
let inFence = false;
const flush = () => {
if (current) blocks.push(current);
current = null;
};
const indentOf = (line) => (line.match(/^\s*/) || [''])[0].length;
const open = (type, i, line) => {
current = { startLine: i + 1, endLine: i + 1, text: line, type, indent: indentOf(line) };
};
for (let i = 0; i < lines.length; i++) {
const line = lines[i];
if (FENCE_RE.test(line)) {
flush();
inFence = !inFence;
blocks.push({ startLine: i + 1, endLine: i + 1, text: line, type: 'code' });
continue;
}
if (inFence) {
blocks.push({ startLine: i + 1, endLine: i + 1, text: line, type: 'code' });
continue;
}
if (line.trim() === '') {
flush();
continue;
}
if (HEADING_RE.test(line)) {
flush();
blocks.push({ startLine: i + 1, endLine: i + 1, text: line, type: 'heading' });
continue;
}
if (TABLE_RE.test(line)) {
flush();
blocks.push({ startLine: i + 1, endLine: i + 1, text: line, type: 'table' });
continue;
}
if (LIST_ITEM_RE.test(line)) {
// Structural exception 1: a paragraph ending in ':' is this list's stem —
// it introduces the items rather than standing alone, so it merges with
// them. Deleting a stem without its list is meaningless, and a stem often
// carries no literal of its own to protect it.
//
// A list item can be a stem too ("5. …alltid eksplisitt:" over its
// sub-bullets) — but only when the following item is nested deeper, or
// sibling bullets would glue together.
const isStem =
current &&
STEM_RE.test(current.text) &&
(current.type === 'paragraph' || indentOf(line) > current.indent);
if (isStem) {
current.type = 'list-item';
if (current.ordered === undefined) current.ordered = ORDERED_ITEM_RE.test(line);
current.endLine = i + 1;
current.text += '\n' + line;
current.stemMerged = true;
current.stemIndent = indentOf(line);
continue;
}
if (current && current.stemMerged && current.endLine === i && indentOf(line) >= current.stemIndent) {
// Subsequent items of the same stemmed list join it too.
current.endLine = i + 1;
current.text += '\n' + line;
continue;
}
// Otherwise a new list item always ends the previous block, even mid-list.
flush();
open('list-item', i, line);
current.ordered = ORDERED_ITEM_RE.test(line);
continue;
}
if (current && CONTINUATION_RE.test(line)) {
// Indented wrap — belongs to the block it continues.
current.endLine = i + 1;
current.text += '\n' + line;
continue;
}
if (current && current.type === 'paragraph') {
current.endLine = i + 1;
current.text += '\n' + line;
continue;
}
flush();
open('paragraph', i, line);
}
flush();
return blocks.sort((a, b) => a.startLine - b.startLine);
}
/** Blocks that can carry a deletable instruction at all. */
const isProse = (block) => block.type === 'paragraph' || block.type === 'list-item';
const wordCount = (s) => s.trim().split(/\s+/).filter(Boolean).length;
/**
* Does the block instruct behaviour at all? A declarative fact does not, however
* absolute its wording: "Haiku: aldri." and "Tone: direct and technical" both
* state a decision rather than correcting how the model works.
*/
function correctsBehaviour(text) {
if (IMPERATIVE_RE.test(text)) return true;
return ABSOLUTE_RE.test(text) && wordCount(text) >= ABSOLUTE_ONLY_MIN_WORDS;
}
/**
* Structural exception 2: an ordered list is a CONTRACT. Numbered steps are a
* sequence whose items reference each other, so a floor marker on any step
* floors the whole run deleting step 2 of a five-step session protocol is not
* the same kind of act as deleting one bullet from a list of platitudes.
*
* Unordered lists deliberately do NOT inherit. B-32a and B-32b are opposite
* calls inside one bullet list, and "the container decides" is precisely the
* reasoning the fasit exists to refute.
*
* @returns {Set<number>} startLine of every block floored by inheritance
*/
function orderedContractFloor(blocks, floored) {
const inherited = new Set();
let run = [];
const closeRun = () => {
if (run.length > 1 && run.some((b) => floored.has(b.startLine))) {
for (const b of run) inherited.add(b.startLine);
}
run = [];
};
for (const block of blocks) {
const contiguous = run.length > 0 && block.startLine === run[run.length - 1].endLine + 1;
if (block.ordered && (run.length === 0 || contiguous)) {
run.push(block);
} else {
closeRun();
if (block.ordered) run.push(block);
}
}
closeRun();
return inherited;
}
/**
* Compensatory-phrasing candidates that survived floor-exclusion.
*
* @param {string} text
* @returns {Array<{lensCheck:string, registerId:string, mechanism:string,
* line:number, startLine:number, endLine:number, lineCount:number, text:string}>}
*/
export function subtractionCandidates(text) {
const detector = SUBTRACT_DETECTORS[0];
const out = [];
const blocks = splitLeafBlocks(text).filter(isProse);
// 1. The blocking floor veto, evaluated over the WHOLE leaf block — so a
// literal on a wrapped continuation line still protects its opening line.
const floored = new Set();
for (const block of blocks) {
if (floorMarker(block.text)) floored.add(block.startLine);
}
// 2. …then propagated across ordered-list contracts.
const inherited = orderedContractFloor(blocks, floored);
for (const block of blocks) {
// 3. Does it correct behaviour at all? A declarative local fact does not.
if (!correctsBehaviour(block.text)) continue;
if (floored.has(block.startLine) || inherited.has(block.startLine)) continue;
out.push({
lensCheck: detector.lensCheck,
registerId: detector.registerId,
mechanism: detector.mechanism,
line: block.startLine,
startLine: block.startLine,
endLine: block.endLine,
lineCount: block.endLine - block.startLine + 1,
text: block.text.trim(),
});
}
return out;
}
/**
* Diagnostics for the floor gate: every prose block that fired the detector but
* was vetoed, with the marker that saved it. Not user-facing this is how a
* later narrowing of the veto can be checked against the fasit.
*
* @param {string} text
* @returns {Array<{startLine:number, endLine:number, marker:string, why:string}>}
*/
export function floorExcluded(text) {
const blocks = splitLeafBlocks(text).filter(isProse);
const floored = new Set();
for (const block of blocks) {
if (floorMarker(block.text)) floored.add(block.startLine);
}
const inherited = orderedContractFloor(blocks, floored);
const out = [];
for (const block of blocks) {
if (!correctsBehaviour(block.text)) continue;
const marker = floorMarker(block.text);
if (marker) {
out.push({ startLine: block.startLine, endLine: block.endLine, marker: marker.name, why: marker.why });
} else if (inherited.has(block.startLine)) {
out.push({
startLine: block.startLine,
endLine: block.endLine,
marker: 'ordered-contract',
why: 'is a step of an ordered list whose sibling carries a local fact',
});
}
}
return out;
}

View file

@ -0,0 +1,39 @@
/**
* write-output the one place a scanner's `--output-file` payload is written.
*
* Every command in this plugin follows the same contract (`.claude/rules/ux-rules.md`):
* run the scanner with `--output-file <path> 2>/dev/null`, check the exit code, then Read
* the file. The path the command chooses is frequently one it has never created e.g.
* `commands/campaign.md` writes its report to
* `~/.claude/config-audit/sessions/campaign-report.json`, which on a fresh machine does not
* exist yet. That is precisely the FIRST run, the case campaign-cli otherwise handles
* gracefully by reporting `initialized: false`.
*
* Before this helper existed, all 13 payload writers called `writeFile` directly and threw
* ENOENT there. The exit code was 3, and the command's own exit-code table reads 3 as "the
* input is missing or corrupt" so the user was told the ledger might be corrupt and
* warned off the one action that would have fixed anything. `saveLedger` had always created
* its parent directory; the payload write simply never did. The asymmetry was accidental.
*
* Creating the parent is the honest behaviour: the caller asked for a file at a path, and
* nothing about a missing intermediate directory is an error the caller can learn from.
*/
import { writeFile, mkdir } from 'node:fs/promises';
import { dirname } from 'node:path';
/**
* Write a scanner payload, creating the parent directory if needed.
*
* Signature-compatible with `writeFile(path, contents, encoding)` so call sites are a pure
* rename the encoding argument is kept rather than defaulted away.
*
* @param {string} path - destination file
* @param {string} contents - serialized payload
* @param {string} [encoding='utf-8']
* @returns {Promise<void>}
*/
export async function writeOutputFile(path, contents, encoding = 'utf-8') {
await mkdir(dirname(path), { recursive: true });
await writeFile(path, contents, encoding);
}

View file

@ -34,17 +34,51 @@ export function parseSimpleYaml(yaml) {
let currentKey = null;
let multiLineValue = '';
let inMultiLine = false;
// Block-sequence state: an empty-valued key may be a YAML block sequence
// key:
// - item
// - item
// We defer the decision until the next line: a `- item` line starts an array;
// anything else (or end-of-input) leaves the empty-valued key as `null` —
// indistinguishable from a plain null value, preserving backwards compatibility.
let seqKey = null; // key awaiting/collecting block-sequence items
let seqItems = null; // null until the first `- ` item is seen
let seqPending = false; // an empty-valued key was seen; items may follow
// Commit the current block-sequence (or empty-valued key) to result.
const commitSeq = () => {
if (seqKey !== null) {
result[normalizeKey(seqKey)] = seqItems !== null ? seqItems : null;
}
seqKey = null;
seqItems = null;
seqPending = false;
};
for (const line of lines) {
// Skip comments and empty lines
// Skip comments and empty lines (these do NOT terminate a block sequence)
if (line.trim().startsWith('#') || line.trim() === '') {
if (inMultiLine) multiLineValue += '\n';
continue;
}
// Block-sequence item: only while a sequence is pending or active
const seqMatch = (seqPending || seqItems !== null) && !inMultiLine
? line.match(/^\s+-\s*(.*)$/)
: null;
if (seqMatch) {
if (seqItems === null) seqItems = [];
const itemVal = parseValue(seqMatch[1].trim());
if (itemVal !== null && itemVal !== '') seqItems.push(itemVal);
seqPending = false;
continue;
}
// Key-value pair
const kvMatch = line.match(/^(\w[\w-]*):\s*(.*)/);
if (kvMatch && !inMultiLine) {
// A new key terminates any pending/active block sequence.
commitSeq();
if (currentKey && multiLineValue) {
result[normalizeKey(currentKey)] = multiLineValue.trim();
}
@ -58,6 +92,15 @@ export function parseSimpleYaml(yaml) {
continue;
}
if (value === '') {
// Defer: this may be a block sequence (next line) or a null value.
seqKey = currentKey;
seqItems = null;
seqPending = true;
currentKey = null;
continue;
}
result[normalizeKey(currentKey)] = parseValue(value);
currentKey = null;
continue;
@ -79,6 +122,13 @@ export function parseSimpleYaml(yaml) {
currentKey = null;
}
}
continue;
}
// A non-item, non-key line while a sequence is pending/active terminates it
// (then the line is ignored, as before the fix).
if (seqPending || seqItems !== null) {
commitSeq();
}
}
@ -86,6 +136,8 @@ export function parseSimpleYaml(yaml) {
if (inMultiLine && currentKey) {
result[normalizeKey(currentKey)] = multiLineValue.trim();
}
// Flush a trailing block sequence / empty-valued key
commitSeq();
// Normalize arrays for known list fields
for (const field of ['allowed_tools', 'tools', 'paths', 'globs']) {

View file

@ -9,13 +9,29 @@
* {
* meta: { repoPath, generatedAt, durationMs },
* sources: [
* { kind: 'claude-md'|'plugin'|'skill'|'mcp-server'|'hook',
* name: string, source: string, estimated_tokens: number },
* { kind: 'claude-md'|'skill'|'rule'|'agent'|'output-style'|'mcp-server'|'hook',
* name: string, source: string, estimated_tokens: number,
* loadPattern: 'always'|'on-demand'|'external'|'unknown',
* survivesCompaction: 'yes'|'no'|'n/a',
* derivationConfidence: 'confirmed'|'inferred' },
* ...
* ],
* summary: {
* always: { tokens, count }, // enter context every turn before you type
* onDemand: { tokens, count }, // loaded on invoke / on file read
* external: { tokens, count }, // run outside the context window (hooks)
* unknown: { tokens, count },
* },
* total: <sum of sources.estimated_tokens>
* }
*
* v5.6 B load-pattern accounting. Sources are component-level: the coarse
* "plugin" roll-up was dropped because a plugin's contributions (skills, rules,
* agents, output styles, hooks, MCP) are each enumerated once on their own
* keeping the roll-up double-counted them and corrupted the always-loaded
* subtotal. Every record now carries the load pattern derived from the
* published Claude Code loading model (deriveLoadPattern).
*
* Usage:
* node manifest.mjs [path] [--json] [--output-file <path>]
*
@ -24,65 +40,190 @@
*/
import { resolve } from 'node:path';
import { writeFile, stat } from 'node:fs/promises';
import { readActiveConfig } from './lib/active-config-reader.mjs';
import { stat } from 'node:fs/promises';
import { writeOutputFile } from './lib/write-output.mjs';
import { readActiveConfig, deriveLoadPattern } from './lib/active-config-reader.mjs';
// CLAUDE.md cascade files are all discovered by walking UP from the repo, so
// each one is always-loaded; the scope only changes the derivation confidence.
const CLAUDE_MD_SCOPE_KIND = {
project: 'claude-md-root',
local: 'claude-md-root',
user: 'claude-md-user',
managed: 'claude-md-managed',
import: 'claude-md-import',
};
/** Spread the three load-pattern fields onto a source record. */
function withLoadPattern(record, lp) {
return {
...record,
loadPattern: lp.loadPattern,
survivesCompaction: lp.survivesCompaction,
derivationConfidence: lp.derivationConfidence,
};
}
const sourceLabel = (item, fallback) =>
item.pluginName ? `plugin:${item.pluginName}` : item.source || fallback;
/**
* Flatten an activeConfig snapshot into a single ranked array of sources.
* Flatten an activeConfig snapshot into a single ranked array of sources, each
* tagged with its load pattern, plus a load-pattern summary.
*/
export function buildManifest(activeConfig) {
const sources = [];
for (const f of activeConfig.claudeMd?.files || []) {
const tokens = estimateClaudeMdEntryTokens(f, activeConfig);
sources.push({
const kind = CLAUDE_MD_SCOPE_KIND[f.scope] || 'claude-md-root';
sources.push(withLoadPattern({
kind: 'claude-md',
name: f.path,
source: f.scope,
estimated_tokens: tokens,
});
}
for (const p of activeConfig.plugins || []) {
sources.push({
kind: 'plugin',
name: p.name,
source: p.path,
estimated_tokens: p.estimatedTokens || 0,
});
}, deriveLoadPattern(kind)));
}
// Skills: the measured tokens are the skill BODY (full file), paid on invoke.
// The always-loaded part (name+description listing) is small and tracked
// separately (skill-listing-budget / posture), so the body is tagged
// on-demand here rather than inflating the always-loaded subtotal.
for (const s of activeConfig.skills || []) {
sources.push({
sources.push(withLoadPattern({
kind: 'skill',
name: s.name,
source: s.pluginName ? `plugin:${s.pluginName}` : s.source || 'user',
source: sourceLabel(s, 'user'),
estimated_tokens: s.estimatedTokens || 0,
});
}, deriveLoadPattern('skill-body')));
}
// Rules / agents / output styles — the foundation enumeration already derived
// the load pattern (rules vary by `scoped`), so propagate it verbatim.
for (const r of activeConfig.rules || []) {
sources.push(withLoadPattern({
kind: 'rule',
name: r.name,
source: sourceLabel(r, 'project'),
estimated_tokens: r.estimatedTokens || 0,
}, r));
}
for (const a of activeConfig.agents || []) {
sources.push(withLoadPattern({
kind: 'agent',
name: a.name,
source: sourceLabel(a, 'project'),
estimated_tokens: a.estimatedTokens || 0,
}, a));
}
for (const o of activeConfig.outputStyles || []) {
sources.push(withLoadPattern({
kind: 'output-style',
name: o.name,
source: sourceLabel(o, 'project'),
estimated_tokens: o.estimatedTokens || 0,
}, o));
}
for (const m of activeConfig.mcpServers || []) {
if (m && m.enabled === false) continue;
sources.push({
sources.push(withLoadPattern({
kind: 'mcp-server',
name: m.name,
source: m.source || 'unknown',
estimated_tokens: m.estimatedTokens || 0,
});
}, deriveLoadPattern('mcp')));
}
for (const h of activeConfig.hooks || []) {
sources.push({
sources.push(withLoadPattern({
kind: 'hook',
name: `${h.event}${h.matcher ? `:${h.matcher}` : ''}`,
source: h.source || h.sourcePath || 'unknown',
estimated_tokens: h.estimatedTokens || 0,
});
}, deriveLoadPattern('hook')));
}
sources.sort((a, b) => b.estimated_tokens - a.estimated_tokens);
const total = sources.reduce((s, x) => s + (x.estimated_tokens || 0), 0);
return { sources, total };
const summary = summarizeByLoadPattern(sources);
return { sources, total, summary };
}
/**
* Bucket sources by load pattern into {tokens, count} subtotals. The `always`
* bucket is the headline: tokens that enter context every turn before the user
* types anything.
*/
export function summarizeByLoadPattern(sources) {
const mk = () => ({ tokens: 0, count: 0 });
const summary = { always: mk(), onDemand: mk(), external: mk(), unknown: mk() };
const BUCKET = { always: 'always', 'on-demand': 'onDemand', external: 'external' };
for (const s of sources) {
const key = BUCKET[s.loadPattern] || 'unknown';
summary[key].tokens += s.estimated_tokens || 0;
summary[key].count += 1;
}
return summary;
}
/**
* Source strings (the `source` field buildManifest stamps) that belong to the
* SHARED GLOBAL layer config paid once per machine and identical in every
* repo: the global ~/.claude CLAUDE.md (`user`) and managed enterprise policy
* (`managed`). Installed plugins are also shared but are matched by the
* `plugin:` prefix below, not by this set.
*
* Deliberately NOT here: `~/.claude.json:projects`. Although that file lives in
* HOME, `readClaudeJsonProjectSlice` returns the slice keyed to the SPECIFIC
* repo path those MCP servers are per-repo, load only in their own project,
* and differ across repos, so they are a delta (folding them into the
* once-counted shared layer would drop every repo's slice but the first). The
* only machine-global MCP is plugin-provided (caught by the `plugin:` prefix).
*/
const SHARED_GLOBAL_SOURCES = Object.freeze(new Set(['user', 'managed']));
/**
* Classify one manifest source as part of the once-counted shared global layer
* or a per-repo delta (v5.9 B2b). Anything not positively identified as global
* (project / local / .mcp.json / ~/.claude.json:projects / @import / unrecognized)
* falls to `delta`, so a source is never silently folded into the shared layer
* a wrong fold would HIDE machine-wide cost, whereas a wrong delta is at worst
* attributed visibly to a repo.
* @param {string} source
* @returns {'shared'|'delta'}
*/
export function classifyOwnership(source) {
if (typeof source === 'string') {
if (source.startsWith('plugin:')) return 'shared'; // installed plugins are machine-global
if (SHARED_GLOBAL_SOURCES.has(source)) return 'shared';
}
return 'delta';
}
/**
* Partition manifest sources by ownership for the machine-wide token roll-up,
* returning two load-pattern summaries in the exact shape `summarizeByLoadPattern`
* emits ({always,onDemand,external,unknown:{tokens,count}}), so the campaign
* ledger setters (`setSharedGlobal` / `setRepoTokens`) consume them verbatim.
*
* - `shared`: the global layer, identical across repos set ONCE on the ledger
* root so the roll-up counts it exactly once (the structural double-count guard).
* - `delta`: this repo's own project/local contribution beyond the shared layer.
*
* The split is total: every source lands in exactly one layer.
* @param {Array<{source:string, loadPattern:string, estimated_tokens:number}>} sources
* @returns {{shared:object, delta:object}}
*/
export function splitManifestByOwnership(sources) {
const shared = [];
const delta = [];
for (const s of sources || []) {
(classifyOwnership(s.source) === 'shared' ? shared : delta).push(s);
}
return { shared: summarizeByLoadPattern(shared), delta: summarizeByLoadPattern(delta) };
}
/**
@ -119,11 +260,13 @@ async function main() {
const s = await stat(absPath);
if (!s.isDirectory()) {
process.stderr.write(`Error: ${absPath} is not a directory\n`);
process.exit(3);
process.exitCode = 3;
return;
}
} catch {
process.stderr.write(`Error: path does not exist: ${absPath}\n`);
process.exit(3);
process.exitCode = 3;
return;
}
const start = Date.now();
@ -138,13 +281,14 @@ async function main() {
durationMs: Date.now() - start,
},
sources: manifest.sources,
summary: manifest.summary,
total: manifest.total,
};
const json = JSON.stringify(output, null, 2);
if (outputFile) {
await writeFile(outputFile, json, 'utf-8');
await writeOutputFile(outputFile, json, 'utf-8');
}
if (jsonMode || rawMode || !outputFile) {
@ -156,6 +300,6 @@ const isDirectRun = process.argv[1] && resolve(process.argv[1]) === resolve(new
if (isDirectRun) {
main().catch(err => {
process.stderr.write(`Fatal: ${err.message}\n`);
process.exit(3);
process.exitCode = 3;
});
}

View file

@ -1,6 +1,6 @@
/**
* MCP Scanner MCP Configuration Validator
* Validates .mcp.json files: server types, trust levels, env vars, unknown fields.
* Validates .mcp.json files: server types, env vars, unknown fields.
* Finding IDs: CA-MCP-NNN
*/
@ -13,12 +13,23 @@ import { truncate } from './lib/string-utils.mjs';
const SCANNER = 'MCP';
const VALID_SERVER_TYPES = new Set(['stdio', 'http', 'sse']);
const VALID_TRUST_LEVELS = new Set(['workspace', 'trusted', 'untrusted']);
// No `trust` field: MCP server approval is dialog/settings-based
// (enableAllProjectMcpServers / enabledMcpjsonServers / disabledMcpjsonServers),
// not a per-server .mcp.json field. Verified against code.claude.com/docs 2026-06-18.
const VALID_SERVER_FIELDS = new Set([
'type', 'command', 'args', 'env', 'url', 'headers', 'timeout', 'trust',
'type', 'command', 'args', 'env', 'url', 'headers', 'timeout',
// alwaysLoad: exempt a server from MCP tool-schema deferral (CC v2.1.121+).
// Verified against code.claude.com/docs/en/mcp.md#exempt-a-server-from-deferral 2026-06-23.
'alwaysLoad',
]);
const ENV_VAR_PATTERN = /\$\{([^}]+)\}/g;
// Match only bare ${IDENTIFIER} references. POSIX expansions like ${VAR%pattern}
// or ${VAR:-default} contain operators and are skipped — Claude Code resolves
// them at launch (CC 2.1.142), so they are not config-defined env references.
const ENV_VAR_PATTERN = /\$\{([A-Za-z_][A-Za-z0-9_]*)\}/g;
// Auto-injected by Claude Code at runtime — never require an env block (CC 2.1.139).
const AUTO_INJECTED_ENV_VARS = new Set(['CLAUDE_PROJECT_DIR']);
/**
* Scan all .mcp.json files discovered.
@ -86,28 +97,6 @@ export async function scan(targetPath, discovery) {
}));
}
// Check trust level
if (!config.trust) {
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: 'Missing trust level',
description: `${file.relPath}: Server "${name}" has no trust level configured.`,
file: file.absPath,
recommendation: 'Add "trust": "workspace"|"trusted"|"untrusted" to explicitly set the trust level.',
}));
} else if (!VALID_TRUST_LEVELS.has(config.trust)) {
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.high,
title: 'Invalid trust level',
description: `${file.relPath}: Server "${name}" has invalid trust level "${config.trust}".`,
file: file.absPath,
evidence: `trust: "${config.trust}"`,
recommendation: 'Use one of: workspace, trusted, untrusted.',
}));
}
// Check for env var references in args without env block
if (Array.isArray(config.args)) {
for (const arg of config.args) {
@ -116,6 +105,7 @@ export async function scan(targetPath, discovery) {
ENV_VAR_PATTERN.lastIndex = 0;
while ((match = ENV_VAR_PATTERN.exec(arg)) !== null) {
const varName = match[1];
if (AUTO_INJECTED_ENV_VARS.has(varName)) continue;
const hasEnvBlock = config.env && typeof config.env === 'object' && varName in config.env;
if (!hasEnvBlock) {
findings.push(finding({

View file

@ -0,0 +1,146 @@
/**
* OPT Scanner Optimization Lens / mechanism-fit (v5.7 Fase 1 Chunk 2a)
*
* The first detector of the "is the config OPTIMAL?" axis (vs. the existing
* "is it CORRECT?" scanners). It reads the machine-readable best-practices
* register (knowledge/best-practices.json) and flags config that works but uses
* a mechanism a better one would fit the deterministic half of the hybrid
* motor (the opus analyzer for prose-judgment cases is Chunk 2b).
*
* CA-OPT-001 A multi-step procedure in CLAUDE.md should be a SKILL (BP-MECH-003).
* CLAUDE.md is for facts Claude holds every turn; a procedure there
* costs always-loaded tokens whether or not you run it, and a skill's
* body loads only on invoke. Detection is deliberately CONSERVATIVE
* (a run of >= 6 consecutive numbered steps) to keep precision high
* the negative corpus in the tests proves null false-positives.
* Framed as a Missed opportunity (humanizer), severity LOW.
*
* Provenance: the recommendation + claim come from the register entry (only a
* `confirmed` entry is used user-facing Verifiseringsplikt); an inline default
* is the graceful fallback if the register is unavailable. Fixture-gated: the
* marketplace-medium CLAUDE.md has no numbered lists, so it emits nothing (SC-5
* byte-stable; the additive OPT scanner entry is stripped from frozen baselines).
*
* Zero external dependencies.
*/
import { readFile } from 'node:fs/promises';
import { finding, scannerResult } from './lib/output.mjs';
import { SEVERITY } from './lib/severity.mjs';
import { parseFrontmatter } from './lib/yaml-parser.mjs';
import { loadRegister, getEntry } from './lib/best-practices-register.mjs';
const SCANNER = 'OPT';
const STEP_THRESHOLD = 6;
const STEP_RE = /^\s*\d+\.\s+\S/;
const PROCEDURE_TITLE = 'A multi-step procedure in CLAUDE.md belongs in a skill';
// Graceful fallback if the register is missing/unreadable (the register is the
// source of truth; this keeps the scanner working without it).
const DEFAULT_MECH_003 = {
claim:
'A multi-step procedure in CLAUDE.md should be a skill — CLAUDE.md is for facts Claude ' +
'should hold all the time; procedures belong in skills.',
recommendation:
'Extract the procedure into .claude/skills/; its body then loads only on invoke instead ' +
'of every turn.',
};
/** Return the confirmed register entry for `id`, or null (→ caller uses default). */
function confirmedEntry(id) {
try {
const e = getEntry(loadRegister(), id);
return e && e.confidence === 'confirmed' ? e : null;
} catch {
return null;
}
}
/**
* Longest run of consecutive numbered-list items. Blank lines and indented
* continuation lines neither extend nor break a run; any other non-step line
* breaks it. Conservative by design (a wrapped, non-indented step line ends the
* run undercount, never overcount).
* @param {string} text
* @returns {{count:number, startIndex:number, firstStep:string}}
*/
function longestNumberedRun(text) {
const lines = String(text).split('\n');
let maxRun = 0;
let maxStartIdx = 0;
let maxFirstStep = '';
let run = 0;
let runStartIdx = 0;
let runFirstStep = '';
for (let i = 0; i < lines.length; i++) {
const line = lines[i];
if (STEP_RE.test(line)) {
if (run === 0) {
runStartIdx = i;
runFirstStep = line.trim();
}
run++;
if (run > maxRun) {
maxRun = run;
maxStartIdx = runStartIdx;
maxFirstStep = runFirstStep;
}
} else if (line.trim() === '' || /^\s+\S/.test(line)) {
// blank or indented continuation — part of the list, neither step nor break
} else {
run = 0;
}
}
return { count: maxRun, startIndex: maxStartIdx, firstStep: maxFirstStep };
}
export async function scan(targetPath, discovery) {
const start = Date.now();
const findings = [];
const claudeMdFiles = ((discovery && discovery.files) || []).filter((f) => f.type === 'claude-md');
const entry = confirmedEntry('BP-MECH-003');
const claim = (entry && entry.claim) || DEFAULT_MECH_003.claim;
const recommendation = (entry && entry.recommendation) || DEFAULT_MECH_003.recommendation;
let filesScanned = 0;
for (const file of claudeMdFiles) {
let content;
try {
content = await readFile(file.absPath, 'utf-8');
} catch {
continue;
}
filesScanned++;
const parsed = parseFrontmatter(content);
const body = parsed.body || content;
const bodyStartLine = parsed.bodyStartLine || 1;
const run = longestNumberedRun(body);
if (run.count >= STEP_THRESHOLD) {
findings.push(
finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: PROCEDURE_TITLE,
description: claim,
file: file.relPath || file.absPath,
line: bodyStartLine + run.startIndex,
evidence: run.firstStep,
recommendation,
category: 'mechanism-fit',
details: {
mechanism: 'skill',
steps: run.count,
register: entry ? entry.id : null,
source: entry ? entry.source.url : null,
confidence: entry ? entry.confidence : null,
},
}),
);
}
}
return scannerResult(SCANNER, 'ok', findings, filesScanned, Date.now() - start);
}

View file

@ -0,0 +1,235 @@
#!/usr/bin/env node
/**
* optimize-lens CLI feeds the v5.7 optimization lens (CA-OPT) `/config-audit
* optimize` command. It produces the two halves of the hybrid motor as one JSON
* payload:
*
* 1. `deterministic` the OPT scanner's high-precision findings (CA-OPT-001:
* a long numbered procedure in CLAUDE.md skill). Already part of the
* orchestrated audit; surfaced here so /optimize is a complete view.
* 2. `candidates` recall-oriented prose-judgment candidates from the
* lens-prefilter (lifecycle hook, unscoped path-specific rule, "never"
* permission), each stamped with the CONFIRMED register entry it might fit
* (claim / recommendation / source / severity). The opus
* optimization-lens-agent is the precision gate over these.
*
* Only CONFIRMED register entries are attached (Verifiseringsplikt); a candidate
* whose register rule is missing or unconfirmed is dropped, so the agent never
* sees an unverifiable recommendation.
*
* Usage:
* node optimize-lens-cli.mjs [path] [--output-file <path>] [--global]
*
* Exit codes: 0=ok, 3=unrecoverable error. Zero external dependencies.
*/
import { resolve, sep } from 'node:path';
import { readFile, stat } from 'node:fs/promises';
import { writeOutputFile } from './lib/write-output.mjs';
import { discoverConfigFiles } from './lib/file-discovery.mjs';
import { resetCounter } from './lib/output.mjs';
import { parseFrontmatter } from './lib/yaml-parser.mjs';
import { loadRegister, getEntry } from './lib/best-practices-register.mjs';
import { prefilterClaudeMd, LENS_DETECTORS } from './lib/lens-prefilter.mjs';
import { subtractionCandidates, SUBTRACT_DETECTORS } from './lib/subtraction-prefilter.mjs';
import { scan as optScan } from './optimization-lens-scanner.mjs';
// Files under `.claude/plugins/` are shipped by an installed plugin — vendored
// CLAUDE.md plus its bundled tests/fixtures and examples. They are not the user's
// authored config, so a mechanism-fit suggestion against them is not actionable
// (the user can't edit a file the plugin overwrites on update). Excluded from the
// lens regardless of active/stale version. (M-BUG-11; mirrors the M-BUG-2 rule
// that keeps plugin-bundled config out of the conflict detector.)
const PLUGIN_TREE_MARKER = `.claude${sep}plugins${sep}`;
const isPluginBundled = (file) => (file.absPath || '').includes(PLUGIN_TREE_MARKER);
/** Confirmed register entry for `id`, or null. */
function confirmedEntry(register, id) {
const e = getEntry(register, id);
return e && e.confidence === 'confirmed' ? e : null;
}
async function main() {
const args = process.argv.slice(2);
let targetPath = '.';
let outputFile = null;
let includeGlobal = false;
let subtract = false;
for (let i = 0; i < args.length; i++) {
if (args[i] === '--global') includeGlobal = true;
else if (args[i] === '--subtract') subtract = true;
else if (args[i] === '--output-file' && args[i + 1]) outputFile = args[++i];
else if (!args[i].startsWith('-')) targetPath = args[i];
}
const absPath = resolve(targetPath);
try {
const s = await stat(absPath);
if (!s.isDirectory()) {
process.stderr.write(`Error: ${absPath} is not a directory\n`);
process.exitCode = 3;
return;
}
} catch {
process.stderr.write(`Error: path does not exist: ${absPath}\n`);
process.exitCode = 3;
return;
}
// Load the register once; tolerate its absence (deterministic half still runs).
let register = null;
try {
register = loadRegister();
} catch {
register = null;
}
resetCounter();
const rawDiscovery = await discoverConfigFiles(absPath, { includeGlobal });
// Scope the lens to the user's authored config: drop plugin-bundled files for
// BOTH halves of the motor (the OPT scanner reads discovery.files directly).
const discovery = {
...rawDiscovery,
files: (rawDiscovery.files || []).filter((f) => !isPluginBundled(f)),
};
// ── Deterministic half: the OPT scanner (CA-OPT-001) ──
const opt = await optScan(absPath, discovery);
// ── Recall half: prose-judgment candidates from the pre-filter ──
const claudeMdFiles = (discovery.files || []).filter((f) => f.type === 'claude-md');
const candidates = [];
// Opt-in only: the subtraction axis asks a different question and must not
// fire on a plain `/config-audit optimize` run (brief §7 q3).
const subtractCands = [];
const subtractEntry = subtract && register ? confirmedEntry(register, 'BP-SUB-001') : null;
for (const file of claudeMdFiles) {
let content;
try {
content = await readFile(file.absPath, 'utf-8');
} catch {
continue;
}
const parsed = parseFrontmatter(content);
const body = parsed.body || content;
const bodyStartLine = parsed.bodyStartLine || 1;
if (subtractEntry) {
for (const cand of subtractionCandidates(body)) {
subtractCands.push({
file: file.absPath,
line: bodyStartLine - 1 + cand.startLine,
endLine: bodyStartLine - 1 + cand.endLine,
lineCount: cand.lineCount,
lensCheck: cand.lensCheck,
mechanism: cand.mechanism,
signalText: cand.text,
register: {
id: subtractEntry.id,
claim: subtractEntry.claim,
recommendation: subtractEntry.recommendation || null,
severity: subtractEntry.severity || 'low',
source: subtractEntry.source,
},
});
}
}
for (const cand of prefilterClaudeMd(body)) {
const entry = register ? confirmedEntry(register, cand.registerId) : null;
if (!entry) continue; // never surface an unverifiable recommendation
candidates.push({
// Absolute path: unique + readable. relPath collides across scopes
// (a repo-root `CLAUDE.md` and the user-global `~/.claude/CLAUDE.md`
// both relPath to `CLAUDE.md`), which would send the agent's Read() to
// the wrong file. (M-BUG-11)
file: file.absPath,
line: bodyStartLine - 1 + cand.line,
lensCheck: cand.lensCheck,
mechanism: cand.mechanism,
signalText: cand.text,
register: {
id: entry.id,
claim: entry.claim,
recommendation: entry.recommendation || null,
severity: entry.severity || 'low',
source: entry.source,
},
});
}
}
// The CONFIRMED prose-judgment entries, so the agent has full provenance even
// for a detector class that produced no candidates this run.
const registerEntries = register
? LENS_DETECTORS.map((d) => confirmedEntry(register, d.registerId))
.filter(Boolean)
.map((e) => ({
id: e.id,
lensCheck: e.lensCheck,
claim: e.claim,
recommendation: e.recommendation || null,
mechanism: e.mechanism || null,
severity: e.severity || 'low',
source: e.source,
}))
: [];
const payload = {
status: 'ok',
target: absPath,
deterministic: opt.findings || [],
candidates,
register: registerEntries,
counts: {
deterministic: (opt.findings || []).length,
candidates: candidates.length,
byLensCheck: candidates.reduce((acc, c) => {
acc[c.lensCheck] = (acc[c.lensCheck] || 0) + 1;
return acc;
}, {}),
},
};
// Additive ONLY under --subtract: a plain run's payload must stay byte-identical.
if (subtract) {
payload.subtract = {
enabled: true,
candidates: subtractCands,
register: subtractEntry
? [
{
id: subtractEntry.id,
lensCheck: subtractEntry.lensCheck,
claim: subtractEntry.claim,
recommendation: subtractEntry.recommendation || null,
mechanism: subtractEntry.mechanism || null,
severity: subtractEntry.severity || 'low',
source: subtractEntry.source,
},
]
: [],
detectors: SUBTRACT_DETECTORS.map((d) => ({ ...d })),
};
payload.counts.subtractCandidates = subtractCands.length;
}
const json = JSON.stringify(payload, null, 2);
if (outputFile) {
await writeOutputFile(outputFile, json, 'utf-8');
}
if (!outputFile) {
process.stdout.write(json + '\n');
}
}
const isDirectRun = process.argv[1] && resolve(process.argv[1]) === resolve(new URL(import.meta.url).pathname);
if (isDirectRun) {
main().catch((err) => {
process.stderr.write(`Fatal: ${err.message}\n`);
process.exitCode = 3;
});
}

View file

@ -0,0 +1,179 @@
/**
* OST Scanner Output-style validation (v5.6 C)
*
* Output styles are live (the standalone `/output-style` command was removed in
* v2.1.91; styles are now managed via `/config`). They are the most surprising
* steering surface because they rewrite the system prompt:
*
* CA-OST-001 A custom (user/project) output style that does NOT set
* `keep-coding-instructions: true` when active, Claude Code
* REMOVES its built-in software-engineering instructions (how to
* scope changes, write comments, verify work) and keeps only the
* style's text. `keep-coding-instructions` defaults to false, so
* this is the headline footgun. Severity medium.
*
* CA-OST-002 A PLUGIN output style with `force-for-plugin: true` Claude
* Code auto-applies it whenever the plugin is enabled, OVERRIDING
* the user's selected `outputStyle`. If several enabled plugins set
* it, the first loaded wins. Severity low (awareness). Note:
* `force-for-plugin` is plugin-styles-only per the docs, so this
* keys on `source === 'plugin'` a user/project style cannot
* trigger the override (it would simply be ignored).
*
* CA-OST-003 A settings `outputStyle` value that matches no built-in and no
* discovered custom style dead config: Claude Code falls back to
* the default style, so the configured behavior is silently not
* applied. Severity medium.
*
* Every claim traces to a CONFIRMED row of docs/v5.5-steering-model-plan.md
* (V9/V10/V11/V12), verified against code.claude.com/docs/en/output-styles and
* .../plugins-reference. The scanner is fixture-gated: with no output styles and
* no `outputStyle` setting it emits nothing (keeps the SC-5 snapshot byte-stable).
*
* Zero external dependencies.
*/
import { readFile } from 'node:fs/promises';
import { finding, scannerResult } from './lib/output.mjs';
import { SEVERITY } from './lib/severity.mjs';
import { readActiveConfig } from './lib/active-config-reader.mjs';
import { parseFrontmatter, parseJson } from './lib/yaml-parser.mjs';
const SCANNER = 'OST';
// Built-in output styles, verified against code.claude.com/docs/en/output-styles.
// Compared case-insensitively so OST-003 never false-flags a valid built-in.
const BUILTIN_STYLES = ['default', 'explanatory', 'learning', 'proactive'];
/**
* Read + parse the frontmatter of each enumerated output style once.
* @param {Array<object>} styles - readActiveConfig().outputStyles entries
*/
async function withFrontmatter(styles) {
const out = [];
for (const s of styles) {
let frontmatter = null;
try {
frontmatter = parseFrontmatter(await readFile(s.path, 'utf-8')).frontmatter;
} catch { /* unreadable → treat as no frontmatter */ }
out.push({ ...s, frontmatter });
}
return out;
}
/**
* Resolve the effective `outputStyle` setting from the cascade (user project
* local; later scope wins). Returns null when unset everywhere.
* @param {object} activeConfig
* @returns {Promise<{value:string, scope:string, path:string} | null>}
*/
async function resolveOutputStyleSetting(activeConfig) {
const cascade = (activeConfig.settings && activeConfig.settings.cascade) || [];
let resolved = null;
for (const entry of cascade) {
if (!entry.exists || !entry.path) continue;
let json = null;
try { json = parseJson(await readFile(entry.path, 'utf-8')); } catch { continue; }
if (json && typeof json.outputStyle === 'string' && json.outputStyle.trim()) {
resolved = { value: json.outputStyle.trim(), scope: entry.scope, path: entry.path };
}
}
return resolved;
}
/**
* Main scanner entry point.
* @param {string} targetPath - repo root to scan
* @param {object} _discovery - unused (OST reads the active config cascade itself)
*/
export async function scan(targetPath, _discovery) {
const start = Date.now();
const findings = [];
const activeConfig = await readActiveConfig(targetPath);
const styles = await withFrontmatter(activeConfig.outputStyles || []);
// CA-OST-001 — user/project custom style missing keep-coding-instructions:true.
for (const s of styles) {
if (s.source !== 'project' && s.source !== 'user') continue;
const kci = s.frontmatter ? s.frontmatter.keep_coding_instructions : undefined;
if (kci === true) continue;
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: 'Custom output style removes built-in coding instructions',
description:
`The ${s.source} output style "${s.name}" does not set ` +
'`keep-coding-instructions: true`. While this style is active, Claude Code ' +
'drops its built-in software-engineering instructions — how to scope changes, ' +
'write comments, and verify work — and keeps only this style\'s text. The ' +
'frontmatter flag defaults to false, so the strip is easy to miss.',
file: s.path,
evidence:
`output_style="${s.name}"; source=${s.source}; ` +
`keep-coding-instructions=${kci === undefined ? 'unset (default false)' : String(kci)}`,
recommendation:
'To keep Claude Code\'s software-engineering behavior while applying this style, ' +
'add `keep-coding-instructions: true` to the frontmatter. If the strip is ' +
'intentional (a non-coding persona), no change is needed.',
category: 'output-styles',
}));
}
// CA-OST-002 — plugin output style with force-for-plugin:true (overrides user choice).
for (const s of styles) {
if (s.source !== 'plugin') continue;
const ffp = s.frontmatter ? s.frontmatter.force_for_plugin : undefined;
if (ffp !== true) continue;
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: 'Plugin output style overrides your selected output style',
description:
`The plugin "${s.pluginName}" ships an output style "${s.name}" with ` +
'`force-for-plugin: true`, so Claude Code applies it automatically whenever the ' +
'plugin is enabled — overriding whatever `outputStyle` you selected. When more ' +
'than one enabled plugin does this, the first one loaded wins.',
file: s.path,
evidence: `output_style="${s.name}"; source=plugin:${s.pluginName}; force-for-plugin=true`,
recommendation:
'If you did not expect this style, disable the plugin or remove ' +
'`force-for-plugin: true` from its output style. This is awareness only — the ' +
'plugin is behaving as designed.',
category: 'output-styles',
}));
}
// CA-OST-003 — settings outputStyle resolving to a non-existent style (dead config).
const resolved = await resolveOutputStyleSetting(activeConfig);
if (resolved) {
const known = new Set([
...BUILTIN_STYLES,
...styles.map(s => String(s.name).toLowerCase()),
]);
if (!known.has(resolved.value.toLowerCase())) {
const customNames = styles.map(s => s.name);
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: 'Configured output style does not exist',
description:
`Your ${resolved.scope} settings set \`outputStyle: "${resolved.value}"\`, but no ` +
'built-in or discovered custom style has that name. Claude Code falls back to the ' +
'default style, so the output behavior you configured is silently never applied.',
file: resolved.path,
evidence:
`outputStyle="${resolved.value}"; scope=${resolved.scope}; ` +
`builtins=[${BUILTIN_STYLES.join(', ')}]; ` +
`known_custom=[${customNames.join(', ')}]`,
recommendation:
'Fix the value to match an existing style (built-ins: Default, Explanatory, ' +
'Learning, Proactive), create the missing style under `.claude/output-styles/`, ' +
'or remove the `outputStyle` setting.',
category: 'output-styles',
}));
}
}
return scannerResult(SCANNER, 'ok', findings, styles.length, Date.now() - start);
}

View file

@ -9,7 +9,8 @@
*/
import { readdir, stat, readFile } from 'node:fs/promises';
import { join, basename, resolve } from 'node:path';
import { writeOutputFile } from './lib/write-output.mjs';
import { join, basename, resolve, sep } from 'node:path';
import { finding, scannerResult, resetCounter } from './lib/output.mjs';
import { SEVERITY } from './lib/severity.mjs';
import { parseFrontmatter } from './lib/yaml-parser.mjs';
@ -19,20 +20,127 @@ const SCANNER = 'PLH';
const REQUIRED_PLUGIN_JSON_FIELDS = ['name', 'description', 'version'];
const RECOMMENDED_CLAUDE_MD_SECTIONS = ['commands', 'agents', 'hooks'];
// Keys as they appear after yaml-parser normalizeKey (hyphens → underscores)
// A CLAUDE.md need only document the component types the plugin actually ships. Mirrors the
// optional-frontmatter rule: do not demand docs for commands/agents/hooks that do not exist.
async function pluginShipsComponent(pluginDir, section) {
if (section === 'hooks') {
try { await readFile(join(pluginDir, 'hooks', 'hooks.json'), 'utf-8'); return true; }
catch { return false; }
}
try {
const entries = await readdir(join(pluginDir, section));
return entries.some(f => f.endsWith('.md'));
} catch { return false; }
}
// Keys as they appear after yaml-parser normalizeKey (hyphens → underscores).
// Field requirements are pinned to the primary docs, NOT to "every field a plugin could set":
// - Commands/skills (code.claude.com/docs slash-commands): "All fields are optional. Only
// `description` is recommended." `name` defaults to the directory name; `model` and
// `allowed-tools` are optional. So only `description` is flagged.
// - Subagents (code.claude.com/docs sub-agents): "Only `name` and `description` are required."
// `model` (defaults to `inherit`) and `tools` (inherits all) are optional.
const REQUIRED_COMMAND_FRONTMATTER = [
{ key: 'name', display: 'name' },
{ key: 'description', display: 'description' },
{ key: 'model', display: 'model' },
{ key: 'allowed_tools', display: 'allowed-tools' },
];
const REQUIRED_AGENT_FRONTMATTER = [
{ key: 'name', display: 'name' },
{ key: 'description', display: 'description' },
{ key: 'model', display: 'model' },
{ key: 'tools', display: 'tools' },
];
// Plugin subagents silently ignore these frontmatter keys — they are honored
// ONLY for user/project agents in .claude/agents/ (code.claude.com/docs
// sub-agents, "ignored for plugin subagents"). Setting them in a plugin agent
// is dead config; permissionMode is MEDIUM because it implies a restriction
// that Claude Code does not actually apply (false sense of security).
const PLUGIN_AGENT_IGNORED_FIELDS = [
{ key: 'permissionMode', severity: SEVERITY.medium },
{ key: 'hooks', severity: SEVERITY.low },
{ key: 'mcpServers', severity: SEVERITY.low },
];
// Component-path keys that REPLACE the default folder (per code.claude.com/docs
// plugins-reference#path-behavior-rules). When such a key is set, Claude Code
// stops scanning the default folder; if that folder still exists, its contents
// are silently ignored (dead config). CC v2.1.140+ flags this in /doctor and
// `claude plugin list`. Excluded by design: `skills` (ADDS to the default —
// both load, never a shadow), and `hooks`/`mcpServers`/`lspServers` (own merge
// rules, not a folder shadow). Experimental themes/monitors are omitted: the
// docs warn their manifest schema may change between releases.
const SHADOWING_PATH_FIELDS = [
{ key: 'commands', defaultDir: 'commands' },
{ key: 'agents', defaultDir: 'agents' },
{ key: 'outputStyles', defaultDir: 'output-styles' },
];
/** Normalize a manifest path: strip a leading "./" and trailing slashes. */
function normalizeManifestPath(p) {
return String(p).replace(/^\.\//, '').replace(/\/+$/, '');
}
/**
* True when a custom manifest path addresses the default folder (equals it or
* points inside it) Claude Code shows no warning in that case because the
* folder is referenced explicitly (e.g. "commands": ["./commands/deploy.md"]).
*/
function addressesDefaultDir(customPath, defaultDir) {
const norm = normalizeManifestPath(customPath);
return norm === defaultDir || norm.startsWith(defaultDir + '/');
}
/** True when `p` exists and is a directory. */
async function dirExists(p) {
try {
return (await stat(p)).isDirectory();
} catch {
return false;
}
}
/** Stat `p`, or null when it does not exist (distinguishes missing from file/dir). */
async function statOrNull(p) {
try {
return await stat(p);
} catch {
return null;
}
}
/**
* True when a `skills` entry resolves outside the plugin root. Installed plugins
* cannot reference files outside their own directory (docs: path-traversal
* limitations), so "../shared" or an absolute path will not load.
*/
function skillsEntryEscapesRoot(pluginDir, entry) {
const resolved = resolve(pluginDir, entry.replace(/^\.\//, ''));
return resolved !== pluginDir && !resolved.startsWith(pluginDir + sep);
}
// Per-problem prose for a malformed `skills` entry. Each `title` starts with
// `plugin.json "skills" entry` so the family is greppable.
const SKILLS_ENTRY_MESSAGES = {
'non-string': {
title: e => `plugin.json "skills" entry is not a string: ${JSON.stringify(e)}`,
description: 'Each "skills" entry must be a relative path string (starting with "./") to a skill directory.',
recommendation: 'Replace the non-string entry with a path like "./my-skill/", or remove it.',
},
'escapes-root': {
title: e => `plugin.json "skills" entry escapes the plugin root: ${e}`,
description: 'Installed plugins cannot reference files outside their own directory, so a skills path that traverses outside the plugin root (e.g. "../shared") will not load.',
recommendation: 'Point the entry at a directory inside the plugin, or vendor the skill into the plugin.',
},
'not-found': {
title: e => `plugin.json "skills" entry does not exist: ${e}`,
description: 'The "skills" entry points at a path that does not exist in the plugin, so no skill loads from it.',
recommendation: 'Create the directory, fix the path, or remove the entry.',
},
'not-a-directory': {
title: e => `plugin.json "skills" entry is a file, not a directory: ${e}`,
description: 'A "skills" entry must be a directory containing a SKILL.md (or <name>/SKILL.md), not a file.',
recommendation: 'Point the entry at the skill directory (the folder that contains SKILL.md), not the file.',
},
};
/**
* Discover plugins under a path.
* Looks for .claude-plugin/plugin.json pattern.
@ -99,6 +207,9 @@ async function scanSinglePlugin(pluginDir) {
const pluginName = basename(pluginDir);
let commandCount = 0;
let agentCount = 0;
// Declared namespace from plugin.json `name` (the prefix for /name:command,
// name:skill, agent "name"). Folder basename is NOT the namespace.
let declaredName = null;
// 1. Validate plugin.json
const pluginJsonPath = join(pluginDir, '.claude-plugin', 'plugin.json');
@ -119,6 +230,9 @@ async function scanSinglePlugin(pluginDir) {
}
if (parsed) {
if (typeof parsed.name === 'string' && parsed.name.trim()) {
declaredName = parsed.name.trim();
}
for (const field of REQUIRED_PLUGIN_JSON_FIELDS) {
if (!parsed[field]) {
findings.push(finding({
@ -131,6 +245,69 @@ async function scanSinglePlugin(pluginDir) {
}));
}
}
// Shadow check: a manifest component-path key that REPLACES a default
// folder which still exists → that folder is silently ignored (dead config).
for (const { key, defaultDir } of SHADOWING_PATH_FIELDS) {
const value = parsed[key];
if (value === undefined || value === null) continue;
const customPaths = (Array.isArray(value) ? value : [value]).filter(p => typeof p === 'string');
if (customPaths.length === 0) continue;
// If any custom path addresses the default folder, CC keeps scanning it → no shadow.
if (customPaths.some(p => addressesDefaultDir(p, defaultDir))) continue;
if (!(await dirExists(join(pluginDir, defaultDir)))) continue;
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: `plugin.json "${key}" path shadows the default ${defaultDir}/ folder`,
description:
`Plugin "${pluginName}" sets "${key}" in plugin.json to ${customPaths.map(p => `"${p}"`).join(', ')}, ` +
`which replaces the default ${defaultDir}/ folder. That folder still exists but Claude Code no longer ` +
`scans it, so its contents are silently ignored (dead config). Claude Code flags this in /doctor and ` +
'`claude plugin list` (v2.1.140+).',
file: pluginJsonPath,
evidence: `${key}=${JSON.stringify(value)}; ignored folder=${defaultDir}/`,
recommendation:
`Either remove the unused ${defaultDir}/ folder, or keep it by listing it explicitly in "${key}" ` +
`(e.g. "${key}": ["./${defaultDir}/", ...]).`,
category: 'plugin-hygiene',
details: { field: key, ignoredDir: defaultDir, customPaths },
}));
}
// skills:-array validation: each entry must resolve to an existing
// directory inside the plugin root. Mirrors `claude plugin validate`.
// skills is string|array (a single string is one entry). Unlike the
// shadow check, skills ADDS to the default skills/ scan, so a custom path
// here is never a shadow — it just has to be a real directory.
if (parsed.skills !== undefined && parsed.skills !== null) {
const entries = Array.isArray(parsed.skills) ? parsed.skills : [parsed.skills];
for (const entry of entries) {
let problem = null;
if (typeof entry !== 'string') {
problem = 'non-string';
} else if (skillsEntryEscapesRoot(pluginDir, entry)) {
problem = 'escapes-root';
} else {
const st = await statOrNull(resolve(pluginDir, entry.replace(/^\.\//, '')));
if (!st) problem = 'not-found';
else if (!st.isDirectory()) problem = 'not-a-directory';
}
if (!problem) continue;
const m = SKILLS_ENTRY_MESSAGES[problem];
findings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: m.title(entry),
description: `Plugin "${pluginName}": ${m.description}`,
file: pluginJsonPath,
evidence: `skills entry=${JSON.stringify(entry)}; problem=${problem}`,
recommendation: m.recommendation,
category: 'plugin-hygiene',
details: { field: 'skills', entry, problem },
}));
}
}
}
} catch {
findings.push(finding({
@ -150,6 +327,9 @@ async function scanSinglePlugin(pluginDir) {
const lower = content.toLowerCase();
for (const section of RECOMMENDED_CLAUDE_MD_SECTIONS) {
// Only require a section for a component the plugin actually ships (mirrors the
// optional-frontmatter rule — no docs demanded for absent commands/agents/hooks).
if (!(await pluginShipsComponent(pluginDir, section))) continue;
// Look for markdown table header or section header
const hasSection = lower.includes(`## ${section}`) ||
lower.includes(`| ${section}`) ||
@ -195,7 +375,7 @@ async function scanSinglePlugin(pluginDir) {
title: 'Command missing frontmatter',
description: `Command "${file}" in plugin "${pluginName}" has no frontmatter`,
file: filePath,
recommendation: 'Add YAML frontmatter with name, description, model',
recommendation: 'Add YAML frontmatter with a description (other command fields are optional)',
}));
continue;
}
@ -234,7 +414,7 @@ async function scanSinglePlugin(pluginDir) {
title: 'Agent missing frontmatter',
description: `Agent "${file}" in plugin "${pluginName}" has no frontmatter`,
file: filePath,
recommendation: 'Add YAML frontmatter with name, description, model, tools',
recommendation: 'Add YAML frontmatter with name and description (model and tools are optional)',
}));
continue;
}
@ -251,6 +431,23 @@ async function scanSinglePlugin(pluginDir) {
}));
}
}
// Plugin subagents ignore hooks/mcpServers/permissionMode frontmatter (V15)
// — dead config (permissionMode = medium: false sense of restriction).
for (const { key, severity } of PLUGIN_AGENT_IGNORED_FIELDS) {
if (frontmatter[key] !== undefined) {
findings.push(finding({
scanner: SCANNER,
severity,
title: `Plugin agent sets "${key}", which Claude Code ignores`,
description: `Agent "${file}" in plugin "${pluginName}" sets "${key}" in frontmatter, but Claude Code ignores ${key} for plugin subagents — ${key === 'permissionMode' ? 'the agent runs with default permissions, not the restricted mode this implies' : 'this configuration has no effect'}.`,
file: filePath,
evidence: `${key}: ${JSON.stringify(frontmatter[key])}`,
recommendation: `Remove "${key}" from the agent frontmatter, or ship the agent as a user/project agent in .claude/agents/, where ${key} is honored.`,
autoFixable: false,
}));
}
}
}
} catch { /* no agents dir */ }
@ -294,7 +491,13 @@ async function scanSinglePlugin(pluginDir) {
const pluginMetaDir = join(pluginDir, '.claude-plugin');
try {
const entries = await readdir(pluginMetaDir);
const known = new Set(['plugin.json']);
// `marketplace.json` belongs here: it is the documented, required location
// for a marketplace catalog (code.claude.com/docs plugin-marketplaces —
// "Create `.claude-plugin/marketplace.json` in your repository root"), and a
// marketplace entry with `"source": "./"` makes the repo root its own
// plugin. Such a repo legitimately carries both files, so flagging the
// catalog as an unknown file was a false positive.
const known = new Set(['plugin.json', 'marketplace.json']);
for (const entry of entries) {
if (!known.has(entry)) {
findings.push(finding({
@ -309,30 +512,65 @@ async function scanSinglePlugin(pluginDir) {
}
} catch { /* skip */ }
return { name: pluginName, findings, commandCount, agentCount };
return { name: pluginName, declaredName, findings, commandCount, agentCount };
}
/**
* Per-plugin score and grade. Single source for both the terminal report and
* the --output-file payload the grade formula used to live only inside
* `formatPluginHealthReport`, which nothing called.
* @param {number} issueCount
* @returns {{ score: number, grade: string }}
*/
export function pluginGrade(issueCount) {
const score = Math.max(0, 100 - issueCount * 10);
const grade = score >= 90 ? 'A' : score >= 75 ? 'B' : score >= 60 ? 'C' : score >= 40 ? 'D' : 'F';
return { score, grade };
}
/**
* Scan one or more plugins and return aggregated results.
*
* The envelope is frozen at the v5.0.0 shape (byte-stable `--raw`/`--json`), so
* per-plugin rows and the cross-plugin/per-plugin split are NOT in it. Callers
* that need those the `--output-file` payload, and therefore
* `/config-audit plugin-health` use `scanDetailed`.
*
* @param {string} targetPath - Plugin dir or marketplace root
* @returns {Promise<object>} Scanner result
*/
export async function scan(targetPath) {
return (await scanDetailed(targetPath)).result;
}
/**
* Scan, and also return what `scan()`'s frozen envelope cannot carry: one row
* per plugin (name, declared namespace, component counts, grade) and the
* cross-plugin findings as a distinct set.
*
* @param {string} targetPath - Plugin dir or marketplace root
* @returns {Promise<{ result: object, plugins: object[], crossPluginFindings: object[] }>}
*/
export async function scanDetailed(targetPath) {
const start = Date.now();
resetCounter();
const pluginDirs = await discoverPlugins(resolve(targetPath));
if (pluginDirs.length === 0) {
return scannerResult(SCANNER, 'ok', [
finding({
scanner: SCANNER,
severity: SEVERITY.info,
title: 'No plugins found',
description: `No Claude Code plugins found under ${targetPath}`,
recommendation: 'Ensure plugins have .claude-plugin/plugin.json',
}),
], 0, Date.now() - start);
return {
result: scannerResult(SCANNER, 'ok', [
finding({
scanner: SCANNER,
severity: SEVERITY.info,
title: 'No plugins found',
description: `No Claude Code plugins found under ${targetPath}`,
recommendation: 'Ensure plugins have .claude-plugin/plugin.json',
}),
], 0, Date.now() - start),
plugins: [],
crossPluginFindings: [],
};
}
const allFindings = [];
@ -344,10 +582,22 @@ export async function scan(targetPath) {
allFindings.push(...result.findings);
}
// Cross-plugin checks: command name conflicts
const commandNames = new Map(); // name → plugin
// Everything pushed from here on is a cross-plugin finding — the boundary the
// payload uses to split them out (they are flattened into `findings` in the
// frozen envelope, where `category: 'plugin-hygiene'` cannot tell them apart
// from the per-plugin shadow/skills findings that share it).
const crossPluginStart = allFindings.length;
// Cross-plugin checks: command-name ambiguity across DIFFERENT plugin namespaces.
// Commands are namespaced by the plugin's declared name (/name:command), so a
// shared command name across DIFFERENT plugins is ambiguity — not a hard
// conflict — mirroring COL's plugin-vs-plugin skill check (low). When two
// plugins share the SAME declared namespace, the namespace-collision finding
// below already covers it, so this check keys on the namespace and fires only
// when a command name spans 2+ DISTINCT namespaces.
const commandsByNamespace = new Map(); // cmdName → Map<namespace, { path }>
for (let idx = 0; idx < pluginResults.length; idx++) {
const pr = pluginResults[idx];
const namespace = pluginResults[idx].declaredName || basename(pluginDirs[idx]);
const commandsDir = join(pluginDirs[idx], 'commands');
try {
const entries = await readdir(commandsDir);
@ -357,24 +607,92 @@ export async function scan(targetPath) {
const { frontmatter } = parseFrontmatter(content);
if (frontmatter && frontmatter.name) {
const cmdName = frontmatter.name;
if (commandNames.has(cmdName)) {
allFindings.push(finding({
scanner: SCANNER,
severity: SEVERITY.high,
title: 'Cross-plugin command name conflict',
description: `Command "${cmdName}" exists in both "${commandNames.get(cmdName)}" and "${pr.name}"`,
file: filePath,
recommendation: `Rename one of the conflicting commands to avoid ambiguity`,
}));
} else {
commandNames.set(cmdName, pr.name);
}
if (!commandsByNamespace.has(cmdName)) commandsByNamespace.set(cmdName, new Map());
const nsMap = commandsByNamespace.get(cmdName);
if (!nsMap.has(namespace)) nsMap.set(namespace, { path: filePath });
}
}
} catch { /* no commands dir */ }
}
for (const [cmdName, nsMap] of commandsByNamespace) {
if (nsMap.size < 2) continue; // single namespace → no cross-plugin ambiguity
const entries = [...nsMap.entries()].map(([namespace, v]) => ({ namespace, path: v.path }));
const namespaceList = entries.map(e => e.namespace).join(', ');
allFindings.push(finding({
scanner: SCANNER,
severity: SEVERITY.low,
title: `Command name "${cmdName}" used by multiple plugins`,
description:
`${entries.length} plugins (${namespaceList}) expose a command named "${cmdName}". ` +
'Even when invocation is namespaced via /plugin:command, shared names create ambiguity ' +
'in error messages, search results, and the command listing.',
file: entries[0].path,
evidence: `name="${cmdName}"; plugins=${entries.map(e => e.namespace).join(',')}`,
recommendation:
'Coordinate command naming across plugins, or rename one to clarify intent. The shared ' +
'name forces every reader to disambiguate by plugin.',
category: 'plugin-hygiene',
details: {
namespaces: entries.map(e => ({ source: `plugin:${e.namespace}`, name: cmdName, path: e.path })),
},
}));
}
return scannerResult(SCANNER, 'ok', allFindings, pluginDirs.length, Date.now() - start);
// Cross-plugin checks: plugin namespace (declared name) collisions.
// Claude Code namespaces every plugin component by the plugin's declared
// `name` (/name:command, name:skill, agent "name"). Two plugins that declare
// the SAME name collapse into one namespace; the resolution between two
// installed plugins is undocumented, so one plugin's components are silently
// shadowed. Name-less plugins are flagged elsewhere and never grouped here.
const byDeclaredName = new Map(); // declaredName → string[] of plugin dirs
for (let idx = 0; idx < pluginResults.length; idx++) {
const declaredName = pluginResults[idx].declaredName;
if (!declaredName) continue;
if (!byDeclaredName.has(declaredName)) byDeclaredName.set(declaredName, []);
byDeclaredName.get(declaredName).push(pluginDirs[idx]);
}
for (const [declaredName, dirs] of byDeclaredName) {
if (dirs.length < 2) continue;
allFindings.push(finding({
scanner: SCANNER,
severity: SEVERITY.medium,
title: `Plugin namespace collision: "${declaredName}"`,
description:
`${dirs.length} plugins declare the same name "${declaredName}" in plugin.json. ` +
`Claude Code namespaces every plugin component by that name ` +
`(/${declaredName}:command, ${declaredName}:skill, agent "${declaredName}"), so the ` +
'namespaces collapse into one. Resolution between two installed plugins of the same ' +
"name is undocumented — one plugin's commands, skills, and agents are silently shadowed " +
'and become unreachable.',
file: join(dirs[0], '.claude-plugin', 'plugin.json'),
evidence: `name="${declaredName}"; plugins=${dirs.map(d => basename(d)).join(',')}`,
recommendation:
'Rename one plugin\'s "name" in plugin.json so each plugin owns a distinct namespace. ' +
'The folder name does not matter — the declared "name" field is the namespace.',
category: 'plugin-hygiene',
details: {
namespaces: dirs.map(d => ({
source: `plugin:${basename(d)}`,
name: declaredName,
path: join(d, '.claude-plugin', 'plugin.json'),
})),
},
}));
}
return {
result: scannerResult(SCANNER, 'ok', allFindings, pluginDirs.length, Date.now() - start),
plugins: pluginResults.map((p, idx) => ({
name: p.name,
declaredName: p.declaredName,
path: pluginDirs[idx],
commandCount: p.commandCount,
agentCount: p.agentCount,
findingCount: p.findings.length,
...pluginGrade(p.findings.length),
})),
crossPluginFindings: allFindings.slice(crossPluginStart),
};
}
/**
@ -391,9 +709,7 @@ export function formatPluginHealthReport(pluginResults, crossPluginFindings) {
lines.push('');
for (const p of pluginResults) {
const issueCount = p.findings.length;
const score = Math.max(0, 100 - issueCount * 10);
const grade = score >= 90 ? 'A' : score >= 75 ? 'B' : score >= 60 ? 'C' : score >= 40 ? 'D' : 'F';
const { score, grade } = pluginGrade(p.findings.length);
const padding = '.'.repeat(Math.max(1, 25 - p.name.length));
lines.push(` ${p.name} ${padding} ${grade} (${score}) ${p.commandCount} commands, ${p.agentCount} agents`);
}
@ -417,19 +733,42 @@ export function formatPluginHealthReport(pluginResults, crossPluginFindings) {
}
// --- CLI entry point ---
const BOOL_FLAGS = ['--json', '--raw'];
const VALUE_FLAGS = ['--output-file'];
async function main() {
const args = process.argv.slice(2);
let targetPath = '.';
let jsonMode = false;
let rawMode = false;
let outputFile = null;
// M-BUG-21, third arm: this loop used to end in
// `else if (!args[i].startsWith('-')) targetPath = args[i]` with no
// unknown-flag branch. An unrecognised flag was dropped silently and its
// VALUE became the scan target, so `--output-file /tmp/x.json` scanned
// /tmp/x.json. Unlike drift-cli, the result LOOKS fine: a non-existent path
// discovers no plugins, so the scanner reported "No plugins found" (info) and
// exit 0 — a green answer to a question nobody asked. Now it fails loudly.
for (let i = 0; i < args.length; i++) {
if (args[i] === '--json') {
jsonMode = true;
} else if (args[i] === '--raw') {
rawMode = true;
} else if (!args[i].startsWith('-')) {
targetPath = args[i];
const arg = args[i];
if (BOOL_FLAGS.includes(arg)) {
if (arg === '--json') jsonMode = true;
else if (arg === '--raw') rawMode = true;
} else if (VALUE_FLAGS.includes(arg)) {
const value = args[i + 1];
if (value === undefined || value.startsWith('-')) {
throw new Error(`Option ${arg} requires a value.`);
}
outputFile = value;
i++;
} else if (arg.startsWith('-')) {
throw new Error(
`Unknown option: ${arg}\n` +
`Valid options: ${[...BOOL_FLAGS, ...VALUE_FLAGS].join(' ')}`
);
} else {
targetPath = arg;
}
}
@ -437,7 +776,7 @@ async function main() {
process.stderr.write(humanizedProgress ? `Plugin Health v2.1.0\n` : `Plugin Health Scanner v2.1.0\n`);
process.stderr.write(`Target: ${resolve(targetPath)}\n\n`);
const result = await scan(targetPath);
const { result, plugins, crossPluginFindings } = await scanDetailed(targetPath);
if (jsonMode || rawMode) {
// --json and --raw both write the v5.0.0-shape result (byte-identical).
@ -450,6 +789,24 @@ async function main() {
for (const f of findings) {
process.stderr.write(` [${f.severity}] ${f.title}\n`);
}
// ux-rules rule 2: the command runs with `2>/dev/null`, so anything it must
// ACT on has to ride in the --output-file payload. Everything above this
// point is stderr, i.e. invisible to `/config-audit plugin-health`.
if (outputFile) {
const crossIds = new Set(crossPluginFindings.map(f => f.id));
for (const f of findings) {
if (crossIds.has(f.id)) f.crossPlugin = true;
}
const payload = {
...result,
findings,
plugins,
cross_plugin_findings: findings.filter(f => crossIds.has(f.id)),
};
await writeOutputFile(outputFile, JSON.stringify(payload, null, 2), 'utf-8');
process.stderr.write(`\nResults written to ${outputFile}\n`);
}
}
}
@ -457,6 +814,6 @@ const isDirectRun = process.argv[1] && resolve(process.argv[1]) === resolve(new
if (isDirectRun) {
main().catch(err => {
process.stderr.write(`Fatal: ${err.message}\n`);
process.exit(3);
process.exitCode = 3;
});
}

Some files were not shown because too many files have changed in this diff Show more