config-audit/commands/knowledge-refresh.md
Kjell Tore Guttormsen caea8aca23 fix(commands): stop answering questions the caller did not ask
Dogfooding `campaign` + `knowledge-refresh` against a throwaway ledger. Seven
defects, all found by running the commands as written and measuring, not by
reading them.

The headline pair only existed together. `knowledge-refresh` built
`STALE_AFTER="--stale-after 30"` and expanded it unquoted, trusting the shell to
split it in two. bash does; zsh — the macOS default, and what the Bash tool runs
here — does not. The CLI got one argv entry, matched no flag, and because it had
no unknown-flag branch, silently kept the 90-day default and reported "✓ All 14
register entries were re-verified within the last 90 days": a true-sounding
sentence about a threshold the user had just overridden. Fixing either half alone
leaves a silent wrong answer or a loud one; both are fixed, and a guard now
rejects any template that packs a flag and its value into one variable.

`knowledge-refresh` also read one register and wrote another: step 6 named an
unanchored `knowledge/best-practices.json` while the CLI reads
`${CLAUDE_PLUGIN_ROOT}/…`, which for an installed plugin is the cache. The
validation gate then ran the cached test against the cached register — green no
matter what was written. The two copies were byte-identical that day, which is
exactly why it was invisible.

`campaign` vouched for repos it could not read. `add /finnes/ikke` returned
`added` + exit 0; `refresh-tokens` then put the phantom in `swept[]` with a
0-token delta and left `skipped[]` empty, so the machine-wide bill claimed
coverage of three repos on a machine with two. Paths stay tracked — an unmounted
volume is a legitimate absence — but are reported as `addedUnverified`, and the
command names them.

Two class sweeps, both measured rather than assumed. `posture` was the single
scanner (1 of 14) whose fatal catch exited 1, which ux-rules defines as a normal
WARNING grade — a crash indistinguishable from a result. And all 13 payload
writers failed on a `--output-file` whose parent did not exist, which on a fresh
machine turned `campaign`'s first run into "the ledger may be corrupt"; they now
share `scanners/lib/write-output.mjs`.

Predicted breadth was too wide for the first time in five sessions: 6 of 8 CLIs
predicted to lack unknown-flag rejection, 4 measured. `drift` and `fix` already
reject them, via a construct the grep did not recognise — a grep matches an
implementation, the invariant is a behaviour. The sweep was rewritten to run each
CLI with a bogus flag and read the exit code.

Suite 1453 → 1469/0. Frozen snapshots untouched. `optimize-lens-cli` and
`token-hotspots-cli` share the unknown-flag defect and are deferred to the v5.14
argument-handling chunk with their positional-swallow arm; the count is recorded
in the guard rather than rounded down to zero.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012NHWjN8EnoxSqRvMTLK2NE
2026-08-01 21:26:39 +02:00

152 lines
7.8 KiB
Markdown

---
name: config-audit:knowledge-refresh
description: Keep the best-practices register fresh — flag stale entries, poll sources for new/changed practices, with human-approved writes only
argument-hint: "[--stale-after N] [--no-candidates]"
allowed-tools: Read, Write, Edit, Bash, WebSearch, WebFetch
model: opus
---
# Config-Audit: Knowledge Refresh
The "living" part of the living knowledge base. The optimization lens (`/config-audit
optimize`) is only as good as the best-practices register it reads — and Claude Code moves
fast. This command keeps `knowledge/best-practices.json` current in two ways:
- **Stale check (deterministic):** every CONFIRMED entry carries a `source.verified` date.
An entry older than the threshold (default **90 days**) is flagged for re-verification —
its source may have changed since.
- **Candidate poll (web):** scan the CC changelog + the Anthropic "Steering Claude Code"
docs/blog for *new* best-practices the register doesn't yet hold, or *changed* guidance
that contradicts an existing entry.
**The Iron rule (Verifiseringsplikt): nothing is ever auto-written.** Every change — a
bumped `verified` date, an updated claim, a brand-new entry — is presented to the user and
applied only on explicit approval, and only after the live source has actually been
re-read. No unverified claim enters the register.
## Implementation
### Step 1: Parse arguments
From `$ARGUMENTS`:
- `--stale-after N` → override the staleness threshold (integer days; default 90).
- `--no-candidates` → run the deterministic stale check only; skip the web poll.
Tell the user what's happening:
```
## Knowledge Refresh
Checking the best-practices register for stale entries (sources that may need
re-verification) and polling for new Claude Code practices...
```
### Step 2: Run the stale-check CLI
```bash
# Pass the threshold as its OWN quoted argument. Building "--stale-after 30" into
# one variable and expanding it unquoted only works if the shell word-splits —
# bash does, zsh (the macOS default) does not, and there the flag silently
# reverted to the 90-day default while the command reported success.
TODAY=$(date +%F)
STALE_AFTER_DAYS=$(echo "$ARGUMENTS" | sed -nE 's/.*--stale-after[ =]+([0-9]+).*/\1/p')
if [ -n "$STALE_AFTER_DAYS" ]; then
node ${CLAUDE_PLUGIN_ROOT}/scanners/knowledge-refresh-cli.mjs \
--reference-date "$TODAY" --stale-after "$STALE_AFTER_DAYS" \
--output-file ~/.claude/config-audit/sessions/knowledge-refresh.json 2>/dev/null; echo $?
else
node ${CLAUDE_PLUGIN_ROOT}/scanners/knowledge-refresh-cli.mjs \
--reference-date "$TODAY" \
--output-file ~/.claude/config-audit/sessions/knowledge-refresh.json 2>/dev/null; echo $?
fi
```
Exit code **0** = all fresh, **1** = some stale (advisory, normal), **3** = real error →
"The refresh check couldn't run — the register file may be missing or invalid."
### Step 3: Read the payload + present stale entries
Read `~/.claude/config-audit/sessions/knowledge-refresh.json` with the Read tool. It has
`counts {total, stale, fresh}`, a `stale[]` array (each: `id, verified, ageDays, url,
claim`), `referenceDate`, and `staleAfterDays`.
Present the stale entries as a markdown table (per the UX rules — never show the raw JSON):
| Entry | Claim (short) | Verified | Age (days) | Source |
|-------|---------------|----------|------------|--------|
If `counts.stale === 0`, say so plainly: "✓ All N register entries were re-verified within
the last {staleAfterDays} days." Then continue to the candidate poll (unless `--no-candidates`).
### Step 4: Candidate + source-change poll (web — skip if `--no-candidates`)
Tell the user this takes a moment ("Polling the changelog + Anthropic docs, ~20-40s...").
1. **Re-verify each stale entry.** `WebFetch` the entry's `source.url` and check whether the
claim it backs is **still accurate**. Three outcomes:
- *Still holds* → propose bumping `source.verified` to today (no claim change).
- *Changed* → propose an updated `claim`/`recommendation` quoting the new source text.
- *Cannot verify* (page gone, paywalled, contradicts) → propose **nothing**; flag it
"needs manual review" (Verifiseringsplikt: never bump a date you couldn't confirm).
2. **Look for new practices.** `WebSearch` the CC changelog and the "Steering Claude Code"
blog/docs for steering/config guidance not already represented by an entry's `lensCheck`.
For each genuine new practice, draft a **candidate** entry (next free `BP-<TOPIC>-NNN` id,
`confidence: "confirmed"` only if a primary source confirms it — otherwise mark `inferred`
and do **not** present it as user-facing).
### Step 5: Present everything for approval — write nothing yet
Group the proposals and ask the user to approve per item:
- **Re-verify (date bump):** "{id} — source re-read, claim still holds → bump verified to {today}?"
- **Update (claim drift):** show the old vs. new claim + the quoted source line.
- **New candidate:** show the drafted entry (id, claim, mechanism, recommendation, source).
- **Needs manual review:** list, with why it couldn't be auto-verified. (No write offered.)
Be explicit: **"I will not change any file until you approve specific items."**
### Step 6: Apply approved writes (only the approved ones)
**Where the register lives — say this before writing.** The register is part of the plugin,
and the stale check above read it from `${CLAUDE_PLUGIN_ROOT}/knowledge/best-practices.json`
(the `registerPath` field in the payload names the exact file). For a marketplace install that
is the plugin cache, so **an edit there is discarded by the next plugin upgrade** — the durable
home for an approved change is the plugin's own checkout. Tell the user which of the two they
are about to write to, using the `registerPath` they can see, before asking for approval.
For each approved item:
1. Edit `${CLAUDE_PLUGIN_ROOT}/knowledge/best-practices.json` — bump `source.verified`, update
the `claim`/`recommendation`, or append the new entry. Keep the file's 2-space JSON
formatting. Use the anchored path, never a bare `knowledge/…` — a relative path resolves
against the user's current repo, which is not the file the CLI read.
2. If a `${CLAUDE_PLUGIN_ROOT}/knowledge/*.md` mirror states the same fact, update it too so
the human-readable mirror doesn't drift from the register.
3. **Validate before declaring done** — re-run the register schema check and confirm zero errors:
```bash
node --test ${CLAUDE_PLUGIN_ROOT}/tests/lib/best-practices-register.test.mjs 2>&1 | tail -5
```
This test loads the register through the same anchored path, so it validates the file you
just edited — that only holds while step 1 uses the anchored path too. If validation fails,
revert that edit and report it — never leave the register invalid.
Report exactly what changed (ids + fields), and what was deferred to manual review.
### Step 7: Next steps
- `/config-audit optimize` — the lens now reads the refreshed register; re-run it to pick up
any new or changed mechanism-fit rules.
- Re-run `/config-audit knowledge-refresh --no-candidates` anytime for a quick staleness scan
without the web poll.
- Commit the register change (`knowledge/best-practices.json` + any `.md` mirror) with a
`chore(knowledge):` message so the provenance bump is in git history.
## Notes
- **Deterministic core, web-driven shell.** The stale classification is byte-stable and
unit-tested (`tests/lib/knowledge-refresh.test.mjs`, `tests/scanners/knowledge-refresh-cli.test.mjs`);
the candidate poll + writes are web/judgment-driven and deliberately **not** byte-stable
(mirrors `/config-audit optimize`).
- **Read-only CLI.** `knowledge-refresh-cli.mjs` never writes the register; all writes happen
here, in the command, after approval.
- The `-cli` suffix keeps it out of the scan-orchestrator, so the scanner count and the
snapshot suite are unaffected.