config-audit/commands/knowledge-refresh.md
Kjell Tore Guttormsen caea8aca23 fix(commands): stop answering questions the caller did not ask
Dogfooding `campaign` + `knowledge-refresh` against a throwaway ledger. Seven
defects, all found by running the commands as written and measuring, not by
reading them.

The headline pair only existed together. `knowledge-refresh` built
`STALE_AFTER="--stale-after 30"` and expanded it unquoted, trusting the shell to
split it in two. bash does; zsh — the macOS default, and what the Bash tool runs
here — does not. The CLI got one argv entry, matched no flag, and because it had
no unknown-flag branch, silently kept the 90-day default and reported "✓ All 14
register entries were re-verified within the last 90 days": a true-sounding
sentence about a threshold the user had just overridden. Fixing either half alone
leaves a silent wrong answer or a loud one; both are fixed, and a guard now
rejects any template that packs a flag and its value into one variable.

`knowledge-refresh` also read one register and wrote another: step 6 named an
unanchored `knowledge/best-practices.json` while the CLI reads
`${CLAUDE_PLUGIN_ROOT}/…`, which for an installed plugin is the cache. The
validation gate then ran the cached test against the cached register — green no
matter what was written. The two copies were byte-identical that day, which is
exactly why it was invisible.

`campaign` vouched for repos it could not read. `add /finnes/ikke` returned
`added` + exit 0; `refresh-tokens` then put the phantom in `swept[]` with a
0-token delta and left `skipped[]` empty, so the machine-wide bill claimed
coverage of three repos on a machine with two. Paths stay tracked — an unmounted
volume is a legitimate absence — but are reported as `addedUnverified`, and the
command names them.

Two class sweeps, both measured rather than assumed. `posture` was the single
scanner (1 of 14) whose fatal catch exited 1, which ux-rules defines as a normal
WARNING grade — a crash indistinguishable from a result. And all 13 payload
writers failed on a `--output-file` whose parent did not exist, which on a fresh
machine turned `campaign`'s first run into "the ledger may be corrupt"; they now
share `scanners/lib/write-output.mjs`.

Predicted breadth was too wide for the first time in five sessions: 6 of 8 CLIs
predicted to lack unknown-flag rejection, 4 measured. `drift` and `fix` already
reject them, via a construct the grep did not recognise — a grep matches an
implementation, the invariant is a behaviour. The sweep was rewritten to run each
CLI with a bogus flag and read the exit code.

Suite 1453 → 1469/0. Frozen snapshots untouched. `optimize-lens-cli` and
`token-hotspots-cli` share the unknown-flag defect and are deferred to the v5.14
argument-handling chunk with their positional-swallow arm; the count is recorded
in the guard rather than rounded down to zero.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012NHWjN8EnoxSqRvMTLK2NE
2026-08-01 21:26:39 +02:00

7.8 KiB

name description argument-hint allowed-tools model
config-audit:knowledge-refresh Keep the best-practices register fresh — flag stale entries, poll sources for new/changed practices, with human-approved writes only [--stale-after N] [--no-candidates] Read, Write, Edit, Bash, WebSearch, WebFetch opus

Config-Audit: Knowledge Refresh

The "living" part of the living knowledge base. The optimization lens (/config-audit optimize) is only as good as the best-practices register it reads — and Claude Code moves fast. This command keeps knowledge/best-practices.json current in two ways:

  • Stale check (deterministic): every CONFIRMED entry carries a source.verified date. An entry older than the threshold (default 90 days) is flagged for re-verification — its source may have changed since.
  • Candidate poll (web): scan the CC changelog + the Anthropic "Steering Claude Code" docs/blog for new best-practices the register doesn't yet hold, or changed guidance that contradicts an existing entry.

The Iron rule (Verifiseringsplikt): nothing is ever auto-written. Every change — a bumped verified date, an updated claim, a brand-new entry — is presented to the user and applied only on explicit approval, and only after the live source has actually been re-read. No unverified claim enters the register.

Implementation

Step 1: Parse arguments

From $ARGUMENTS:

  • --stale-after N → override the staleness threshold (integer days; default 90).
  • --no-candidates → run the deterministic stale check only; skip the web poll.

Tell the user what's happening:

## Knowledge Refresh

Checking the best-practices register for stale entries (sources that may need
re-verification) and polling for new Claude Code practices...

Step 2: Run the stale-check CLI

# Pass the threshold as its OWN quoted argument. Building "--stale-after 30" into
# one variable and expanding it unquoted only works if the shell word-splits —
# bash does, zsh (the macOS default) does not, and there the flag silently
# reverted to the 90-day default while the command reported success.
TODAY=$(date +%F)
STALE_AFTER_DAYS=$(echo "$ARGUMENTS" | sed -nE 's/.*--stale-after[ =]+([0-9]+).*/\1/p')
if [ -n "$STALE_AFTER_DAYS" ]; then
  node ${CLAUDE_PLUGIN_ROOT}/scanners/knowledge-refresh-cli.mjs \
    --reference-date "$TODAY" --stale-after "$STALE_AFTER_DAYS" \
    --output-file ~/.claude/config-audit/sessions/knowledge-refresh.json 2>/dev/null; echo $?
else
  node ${CLAUDE_PLUGIN_ROOT}/scanners/knowledge-refresh-cli.mjs \
    --reference-date "$TODAY" \
    --output-file ~/.claude/config-audit/sessions/knowledge-refresh.json 2>/dev/null; echo $?
fi

Exit code 0 = all fresh, 1 = some stale (advisory, normal), 3 = real error → "The refresh check couldn't run — the register file may be missing or invalid."

Step 3: Read the payload + present stale entries

Read ~/.claude/config-audit/sessions/knowledge-refresh.json with the Read tool. It has counts {total, stale, fresh}, a stale[] array (each: id, verified, ageDays, url, claim), referenceDate, and staleAfterDays.

Present the stale entries as a markdown table (per the UX rules — never show the raw JSON):

Entry Claim (short) Verified Age (days) Source

If counts.stale === 0, say so plainly: "✓ All N register entries were re-verified within the last {staleAfterDays} days." Then continue to the candidate poll (unless --no-candidates).

Step 4: Candidate + source-change poll (web — skip if --no-candidates)

Tell the user this takes a moment ("Polling the changelog + Anthropic docs, ~20-40s...").

  1. Re-verify each stale entry. WebFetch the entry's source.url and check whether the claim it backs is still accurate. Three outcomes:
    • Still holds → propose bumping source.verified to today (no claim change).
    • Changed → propose an updated claim/recommendation quoting the new source text.
    • Cannot verify (page gone, paywalled, contradicts) → propose nothing; flag it "needs manual review" (Verifiseringsplikt: never bump a date you couldn't confirm).
  2. Look for new practices. WebSearch the CC changelog and the "Steering Claude Code" blog/docs for steering/config guidance not already represented by an entry's lensCheck. For each genuine new practice, draft a candidate entry (next free BP-<TOPIC>-NNN id, confidence: "confirmed" only if a primary source confirms it — otherwise mark inferred and do not present it as user-facing).

Step 5: Present everything for approval — write nothing yet

Group the proposals and ask the user to approve per item:

  • Re-verify (date bump): "{id} — source re-read, claim still holds → bump verified to {today}?"
  • Update (claim drift): show the old vs. new claim + the quoted source line.
  • New candidate: show the drafted entry (id, claim, mechanism, recommendation, source).
  • Needs manual review: list, with why it couldn't be auto-verified. (No write offered.)

Be explicit: "I will not change any file until you approve specific items."

Step 6: Apply approved writes (only the approved ones)

Where the register lives — say this before writing. The register is part of the plugin, and the stale check above read it from ${CLAUDE_PLUGIN_ROOT}/knowledge/best-practices.json (the registerPath field in the payload names the exact file). For a marketplace install that is the plugin cache, so an edit there is discarded by the next plugin upgrade — the durable home for an approved change is the plugin's own checkout. Tell the user which of the two they are about to write to, using the registerPath they can see, before asking for approval.

For each approved item:

  1. Edit ${CLAUDE_PLUGIN_ROOT}/knowledge/best-practices.json — bump source.verified, update the claim/recommendation, or append the new entry. Keep the file's 2-space JSON formatting. Use the anchored path, never a bare knowledge/… — a relative path resolves against the user's current repo, which is not the file the CLI read.
  2. If a ${CLAUDE_PLUGIN_ROOT}/knowledge/*.md mirror states the same fact, update it too so the human-readable mirror doesn't drift from the register.
  3. Validate before declaring done — re-run the register schema check and confirm zero errors:
    node --test ${CLAUDE_PLUGIN_ROOT}/tests/lib/best-practices-register.test.mjs 2>&1 | tail -5
    
    This test loads the register through the same anchored path, so it validates the file you just edited — that only holds while step 1 uses the anchored path too. If validation fails, revert that edit and report it — never leave the register invalid.

Report exactly what changed (ids + fields), and what was deferred to manual review.

Step 7: Next steps

  • /config-audit optimize — the lens now reads the refreshed register; re-run it to pick up any new or changed mechanism-fit rules.
  • Re-run /config-audit knowledge-refresh --no-candidates anytime for a quick staleness scan without the web poll.
  • Commit the register change (knowledge/best-practices.json + any .md mirror) with a chore(knowledge): message so the provenance bump is in git history.

Notes

  • Deterministic core, web-driven shell. The stale classification is byte-stable and unit-tested (tests/lib/knowledge-refresh.test.mjs, tests/scanners/knowledge-refresh-cli.test.mjs); the candidate poll + writes are web/judgment-driven and deliberately not byte-stable (mirrors /config-audit optimize).
  • Read-only CLI. knowledge-refresh-cli.mjs never writes the register; all writes happen here, in the command, after approval.
  • The -cli suffix keeps it out of the scan-orchestrator, so the scanner count and the snapshot suite are unaffected.