fix(commands): stop answering questions the caller did not ask
Dogfooding `campaign` + `knowledge-refresh` against a throwaway ledger. Seven
defects, all found by running the commands as written and measuring, not by
reading them.
The headline pair only existed together. `knowledge-refresh` built
`STALE_AFTER="--stale-after 30"` and expanded it unquoted, trusting the shell to
split it in two. bash does; zsh — the macOS default, and what the Bash tool runs
here — does not. The CLI got one argv entry, matched no flag, and because it had
no unknown-flag branch, silently kept the 90-day default and reported "✓ All 14
register entries were re-verified within the last 90 days": a true-sounding
sentence about a threshold the user had just overridden. Fixing either half alone
leaves a silent wrong answer or a loud one; both are fixed, and a guard now
rejects any template that packs a flag and its value into one variable.
`knowledge-refresh` also read one register and wrote another: step 6 named an
unanchored `knowledge/best-practices.json` while the CLI reads
`${CLAUDE_PLUGIN_ROOT}/…`, which for an installed plugin is the cache. The
validation gate then ran the cached test against the cached register — green no
matter what was written. The two copies were byte-identical that day, which is
exactly why it was invisible.
`campaign` vouched for repos it could not read. `add /finnes/ikke` returned
`added` + exit 0; `refresh-tokens` then put the phantom in `swept[]` with a
0-token delta and left `skipped[]` empty, so the machine-wide bill claimed
coverage of three repos on a machine with two. Paths stay tracked — an unmounted
volume is a legitimate absence — but are reported as `addedUnverified`, and the
command names them.
Two class sweeps, both measured rather than assumed. `posture` was the single
scanner (1 of 14) whose fatal catch exited 1, which ux-rules defines as a normal
WARNING grade — a crash indistinguishable from a result. And all 13 payload
writers failed on a `--output-file` whose parent did not exist, which on a fresh
machine turned `campaign`'s first run into "the ledger may be corrupt"; they now
share `scanners/lib/write-output.mjs`.
Predicted breadth was too wide for the first time in five sessions: 6 of 8 CLIs
predicted to lack unknown-flag rejection, 4 measured. `drift` and `fix` already
reject them, via a construct the grep did not recognise — a grep matches an
implementation, the invariant is a behaviour. The sweep was rewritten to run each
CLI with a bogus flag and read the exit code.
Suite 1453 → 1469/0. Frozen snapshots untouched. `optimize-lens-cli` and
`token-hotspots-cli` share the unknown-flag defect and are deferred to the v5.14
argument-handling chunk with their positional-swallow arm; the count is recorded
in the guard rather than rounded down to zero.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012NHWjN8EnoxSqRvMTLK2NE
This commit is contained in:
parent
acd1cf1248
commit
caea8aca23
23 changed files with 742 additions and 48 deletions
91
tests/commands/command-flag-value-portability.test.mjs
Normal file
91
tests/commands/command-flag-value-portability.test.mjs
Normal file
|
|
@ -0,0 +1,91 @@
|
|||
/**
|
||||
* Session #51 — command-template flag-value portability.
|
||||
*
|
||||
* Dogfooding `knowledge-refresh` surfaced a defect that only exists at the
|
||||
* seam between the command template and the shell that runs it:
|
||||
*
|
||||
* STALE_AFTER="--stale-after 30"
|
||||
* node …-cli.mjs --reference-date "$TODAY" $STALE_AFTER --output-file …
|
||||
*
|
||||
* The unquoted `$STALE_AFTER` is meant to split into TWO argv entries. Under
|
||||
* **bash** it does. Under **zsh** — the macOS default since Catalina, and the
|
||||
* shell the Bash tool actually runs on this machine — unquoted parameter
|
||||
* expansions are NOT word-split, so the CLI receives ONE argv entry with the
|
||||
* literal text `--stale-after 30`, matches no known flag, and (because the CLI
|
||||
* silently ignored unknown flags — see cli-unknown-flag-rejection.test.mjs)
|
||||
* falls back to the 90-day default while reporting success. Measured:
|
||||
*
|
||||
* $ STALE_AFTER="--stale-after 30"; set -- $STALE_AFTER; echo $#
|
||||
* 1 # zsh (bash prints 2)
|
||||
* → payload staleAfterDays: 90, exit 0, "✓ All 14 entries fresh"
|
||||
*
|
||||
* The user-facing knob was silently dead. Note the asymmetry that makes this
|
||||
* survivable elsewhere: an EMPTY unquoted expansion yields ZERO argv entries in
|
||||
* both shells, so the `FLAG=""` idiom used by ~25 other sites is portable. Only
|
||||
* a variable that can hold a flag AND its value is affected.
|
||||
*
|
||||
* The invariant asserted here is therefore about VALUE-carrying flags, not
|
||||
* about quoting in general: a command template must never depend on the shell
|
||||
* splitting one variable into a flag plus its argument. Pass the value through
|
||||
* its own quoted variable instead.
|
||||
*/
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { readFile, readdir } from 'node:fs/promises';
|
||||
import { resolve, dirname } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
|
||||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||
const COMMANDS_DIR = resolve(__dirname, '..', '..', 'commands');
|
||||
|
||||
async function commandFiles() {
|
||||
const entries = await readdir(COMMANDS_DIR);
|
||||
return entries.filter((e) => e.endsWith('.md')).sort();
|
||||
}
|
||||
|
||||
/**
|
||||
* Assignments whose right-hand side contains a flag followed by a value —
|
||||
* i.e. the value only reaches argv if the shell word-splits. Matches both
|
||||
* `X="--flag value"` and `X="--flag $(cmd)"`.
|
||||
*/
|
||||
const MULTIWORD_FLAG_ASSIGN = /^\s*([A-Z_][A-Z0-9_]*)=(["'])(--[a-z0-9-]+)[ \t]+\S.*\2\s*$/;
|
||||
|
||||
test('no command template builds a flag AND its value into one shell variable', async () => {
|
||||
const offenders = [];
|
||||
|
||||
for (const file of await commandFiles()) {
|
||||
const content = await readFile(resolve(COMMANDS_DIR, file), 'utf-8');
|
||||
content.split('\n').forEach((line, i) => {
|
||||
const m = line.match(MULTIWORD_FLAG_ASSIGN);
|
||||
if (m) offenders.push(`${file}:${i + 1} ${m[1]}=${m[2]}${m[3]} …${m[2]}`);
|
||||
});
|
||||
}
|
||||
|
||||
assert.deepEqual(
|
||||
offenders,
|
||||
[],
|
||||
'A variable holding "--flag value" only reaches argv correctly if the shell\n' +
|
||||
'word-splits an unquoted expansion. zsh does not. Pass the value in its own\n' +
|
||||
'quoted variable instead:\n' +
|
||||
' N=$(… extract …); [ -n "$N" ] && node cli.mjs --flag "$N"\n' +
|
||||
'Offending assignments:\n ' + offenders.join('\n '),
|
||||
);
|
||||
});
|
||||
|
||||
test('knowledge-refresh.md passes --stale-after with a quoted value', async () => {
|
||||
const content = await readFile(resolve(COMMANDS_DIR, 'knowledge-refresh.md'), 'utf-8');
|
||||
|
||||
assert.ok(
|
||||
!/\$STALE_AFTER\b(?!")/.test(content.replace(/"\$STALE_AFTER"/g, '')),
|
||||
'knowledge-refresh.md still expands a flag-carrying variable unquoted; under zsh the\n' +
|
||||
'threshold silently reverts to the 90-day default while the command reports success.',
|
||||
);
|
||||
|
||||
assert.match(
|
||||
content,
|
||||
/--stale-after "\$[A-Z_]+"/,
|
||||
'knowledge-refresh.md must pass the extracted threshold as its own quoted argument\n' +
|
||||
'(`--stale-after "$STALE_AFTER_DAYS"`), so no word-splitting is required.',
|
||||
);
|
||||
});
|
||||
|
|
@ -215,3 +215,38 @@ test('state.yaml: phase commands name all four fields the rule requires', async
|
|||
}
|
||||
assert.deepEqual(violations, [], `Incomplete state.yaml contracts:\n${violations.join('\n')}`);
|
||||
});
|
||||
|
||||
/**
|
||||
* Session #51 — the instruction that CAUSED the $TODAY defect must not outlive its fix.
|
||||
*
|
||||
* #49 fixed campaign.md by re-deriving `TODAY=$(date +%F)` inside each of the six write
|
||||
* blocks. But step 1 still carried the original prose — "Set a shared date stamp for any
|
||||
* write: `TODAY=$(date +%F)`" — which is not in a fence, cannot set anything, and directly
|
||||
* contradicts the six comments added below it. A template that argues with itself is not a
|
||||
* contract, and the next edit is the one that believes the wrong half.
|
||||
*
|
||||
* The guard above asserts fences; this one asserts that no PROSE line instructs the reader
|
||||
* to establish shell state for later blocks.
|
||||
*/
|
||||
test('no command template instructs shell state to be set outside a fence', async () => {
|
||||
const offenders = [];
|
||||
|
||||
for (const file of await commandFiles()) {
|
||||
const content = await readFile(resolve(COMMANDS_DIR, file), 'utf-8');
|
||||
const { lines, blockIndexOf } = parseFences(content);
|
||||
lines.forEach((line, i) => {
|
||||
if (blockIndexOf(i) !== -1) return; // inside a fence: that is where state belongs
|
||||
if (/`[A-Z_][A-Z0-9_]*=\$\(/.test(line) || /^\s*[A-Z_][A-Z0-9_]*=\$\(/.test(line)) {
|
||||
offenders.push(`${file}:${i + 1} ${line.trim()}`);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
assert.deepEqual(
|
||||
offenders,
|
||||
[],
|
||||
'Prose that tells the reader to set a shell variable implies it survives to a later\n' +
|
||||
'block. It does not — every fence is its own process. Delete the instruction; the\n' +
|
||||
'blocks that need the value derive it themselves:\n ' + offenders.join('\n '),
|
||||
);
|
||||
});
|
||||
|
|
|
|||
91
tests/commands/knowledge-refresh-write-target.test.mjs
Normal file
91
tests/commands/knowledge-refresh-write-target.test.mjs
Normal file
|
|
@ -0,0 +1,91 @@
|
|||
/**
|
||||
* Session #51 — knowledge-refresh must write the register it actually read.
|
||||
*
|
||||
* The command reads the register through `knowledge-refresh-cli.mjs`, whose `REGISTER_PATH`
|
||||
* is anchored to the CLI module itself — i.e. `${CLAUDE_PLUGIN_ROOT}/knowledge/
|
||||
* best-practices.json`. For any installed plugin that is the marketplace cache; measured
|
||||
* live during this chunk:
|
||||
*
|
||||
* registerPath: …/.claude/plugins/cache/ktg-plugin-marketplace/config-audit/5.13.0/
|
||||
* knowledge/best-practices.json
|
||||
*
|
||||
* Step 6 then said: Edit `knowledge/best-practices.json` — an UNANCHORED relative path,
|
||||
* which resolves against whatever repo the user happens to be sitting in. Three consequences,
|
||||
* none of them visible at the time:
|
||||
*
|
||||
* 1. For a normal user, that path does not exist in their repo at all, so the approved
|
||||
* change either fails or drops a stray file into their project.
|
||||
* 2. In the plugin's own checkout it resolves to the working tree, so the command reads
|
||||
* one file and writes a different one — the refresh appears to do nothing, because the
|
||||
* CLI keeps reporting the cached copy's dates.
|
||||
* 3. Step 6.3's validation gate runs the plugin's own test file, which loads the register
|
||||
* via the same anchored `REGISTER_PATH`. So the gate validates the copy that was NOT
|
||||
* edited and passes no matter what was written — a guard that cannot see the file it
|
||||
* guards.
|
||||
*
|
||||
* The two register copies were byte-identical on the day this was found, which is exactly
|
||||
* why the defect was invisible to a casual dogfood run.
|
||||
*/
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { readFile } from 'node:fs/promises';
|
||||
import { resolve, dirname } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
|
||||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||
const FILE = resolve(__dirname, '..', '..', 'commands', 'knowledge-refresh.md');
|
||||
|
||||
/**
|
||||
* Scoped to lines that tell the model to MODIFY the register. Descriptive prose ("this
|
||||
* command keeps knowledge/best-practices.json current") and the closing `git commit`
|
||||
* suggestion name the file without acting on it, and a git path is correctly repo-relative.
|
||||
* The defect was specifically an unanchored path in an imperative write step.
|
||||
*/
|
||||
const WRITE_VERB = /\b(edit|write|append|overwrite|save)\b/i;
|
||||
|
||||
function unanchoredWriteTargets(content) {
|
||||
const out = [];
|
||||
content.split('\n').forEach((line, i) => {
|
||||
if (!WRITE_VERB.test(line)) return;
|
||||
for (const m of line.matchAll(/(\S*)knowledge\/best-practices\.json/g)) {
|
||||
if (!m[1].endsWith('${CLAUDE_PLUGIN_ROOT}/')) out.push(`${i + 1}: ${line.trim()}`);
|
||||
}
|
||||
});
|
||||
return out;
|
||||
}
|
||||
|
||||
test('the guard still catches the original defect', () => {
|
||||
const original = '1. Edit `knowledge/best-practices.json` — bump `source.verified`, update the `claim`/';
|
||||
assert.equal(
|
||||
unanchoredWriteTargets(original).length,
|
||||
1,
|
||||
'A narrowed guard must still fail on the text it was written for.',
|
||||
);
|
||||
});
|
||||
|
||||
test('every register path in knowledge-refresh.md is anchored to the plugin root', async () => {
|
||||
const content = await readFile(FILE, 'utf-8');
|
||||
const unanchored = unanchoredWriteTargets(content);
|
||||
|
||||
assert.deepEqual(
|
||||
unanchored,
|
||||
[],
|
||||
'A bare `knowledge/best-practices.json` resolves against the user\'s current repo, not\n' +
|
||||
'against the register the CLI actually read. Anchor every reference to\n' +
|
||||
'${CLAUDE_PLUGIN_ROOT}/ so the file that is read, written, and validated is one file:\n ' +
|
||||
unanchored.join('\n '),
|
||||
);
|
||||
});
|
||||
|
||||
test('the write step says where the register really lives', async () => {
|
||||
const content = await readFile(FILE, 'utf-8');
|
||||
assert.match(
|
||||
content,
|
||||
/marketplace|plugin cache|replaced on upgrade|plugin's own checkout/i,
|
||||
'Step 6 must state that the register is part of the installed plugin, so an approved\n' +
|
||||
'edit to a marketplace-installed copy is discarded by the next plugin upgrade and the\n' +
|
||||
'durable change belongs in the plugin\'s own checkout. Silently editing a cache\n' +
|
||||
'directory is the kind of write that looks successful and evaporates.',
|
||||
);
|
||||
});
|
||||
|
|
@ -85,15 +85,21 @@ describe('campaign-write-cli — add', () => {
|
|||
it('auto-initializes when no ledger exists, then adds repos (exit 0)', async () => {
|
||||
const dir = newDir();
|
||||
const file = join(dir, 'ledger.json');
|
||||
// Real directories: since #51 `add` reports a path it cannot read under
|
||||
// `addedUnverified` instead of vouching for it. This test is about auto-init and
|
||||
// tracking, so its fixtures must be repos that actually exist.
|
||||
const a = newDir();
|
||||
const b = newDir();
|
||||
assert.ok(!existsSync(file));
|
||||
const { status, stdout } = runWrite([
|
||||
'add', '/r/a', '/r/b', '--ledger-file', file, '--reference-date', NOW,
|
||||
'add', a, b, '--ledger-file', file, '--reference-date', NOW,
|
||||
]);
|
||||
assert.equal(status, 0);
|
||||
const out = JSON.parse(stdout);
|
||||
assert.equal(out.action, 'add');
|
||||
assert.equal(out.autoInitialized, true);
|
||||
assert.deepEqual(out.added.sort(), [resolve('/r/a'), resolve('/r/b')].sort());
|
||||
assert.deepEqual(out.added.sort(), [resolve(a), resolve(b)].sort());
|
||||
assert.deepEqual(out.addedUnverified, []);
|
||||
|
||||
const { ledger, validation } = await readLedger(file);
|
||||
assert.ok(validation.valid, validation.errors.join('; '));
|
||||
|
|
@ -104,16 +110,19 @@ describe('campaign-write-cli — add', () => {
|
|||
it('appends to an existing ledger and is idempotent on re-add (added vs skipped)', async () => {
|
||||
const dir = newDir();
|
||||
const file = join(dir, 'ledger.json');
|
||||
runWrite(['add', '/r/a', '/r/b', '--ledger-file', file, '--reference-date', NOW]);
|
||||
const a = newDir();
|
||||
const b = newDir();
|
||||
const c = newDir();
|
||||
runWrite(['add', a, b, '--ledger-file', file, '--reference-date', NOW]);
|
||||
|
||||
const { status, stdout } = runWrite([
|
||||
'add', '/r/a', '/r/c', '--ledger-file', file, '--reference-date', NOW,
|
||||
'add', a, c, '--ledger-file', file, '--reference-date', NOW,
|
||||
]);
|
||||
assert.equal(status, 0);
|
||||
const out = JSON.parse(stdout);
|
||||
assert.equal(out.autoInitialized, false);
|
||||
assert.deepEqual(out.added, [resolve('/r/c')]);
|
||||
assert.deepEqual(out.skipped, [resolve('/r/a')]);
|
||||
assert.deepEqual(out.added, [resolve(c)]);
|
||||
assert.deepEqual(out.skipped, [resolve(a)]);
|
||||
|
||||
const { ledger } = await readLedger(file);
|
||||
assert.equal(ledger.repos.length, 3);
|
||||
|
|
@ -316,3 +325,69 @@ describe('campaign-write-cli — determinism + --output-file', () => {
|
|||
assert.equal(status, 3);
|
||||
});
|
||||
});
|
||||
|
||||
// ── Session #51: honest coverage — the ledger must not vouch for repos it cannot read ──
|
||||
//
|
||||
// Dogfooding the campaign lifecycle put a path that does not exist into the ledger and
|
||||
// swept it. Both halves lied, quietly:
|
||||
//
|
||||
// add → added: ["/finnes/absolutt/ikke/noe-repo"], exit 0, no warning.
|
||||
// sweep → swept: [… , "/finnes/absolutt/ikke/noe-repo"], skipped: [], byRepo entry with 0
|
||||
// tokens, reposWithTokens: 3 for a machine with 2 real repos.
|
||||
//
|
||||
// The root cause is that `readActiveConfig` resolves a path and every sub-reader tolerates
|
||||
// ENOENT, so a missing repo yields an EMPTY config rather than an error — and the sweep's
|
||||
// try/catch only routes THROWN errors to `skipped`. The command's own honesty clause ("If
|
||||
// anything was skipped, name those repos plainly so the user knows the bill omits them")
|
||||
// could therefore never fire. A token bill that silently counts phantom repos as 0 is worse
|
||||
// than one that refuses to answer: it looks complete.
|
||||
|
||||
describe('campaign-write-cli — honest coverage for unreadable repos', () => {
|
||||
it('add flags a path that does not exist instead of vouching for it', async () => {
|
||||
const dir = newDir();
|
||||
const file = join(dir, 'ledger.json');
|
||||
const real = newDir();
|
||||
const phantom = join(dir, 'no-such-repo');
|
||||
|
||||
const { status, stdout } = runWrite([
|
||||
'add', real, phantom, '--ledger-file', file, '--reference-date', NOW,
|
||||
]);
|
||||
assert.equal(status, 0, 'a missing path is reported, not rejected — it may be unmounted');
|
||||
|
||||
const out = JSON.parse(stdout);
|
||||
assert.deepEqual(out.added, [resolve(real)], 'only the readable repo is vouched for');
|
||||
assert.deepEqual(
|
||||
out.addedUnverified,
|
||||
[resolve(phantom)],
|
||||
'a path that does not exist must be reported separately so the command can say so',
|
||||
);
|
||||
|
||||
// Still tracked — the campaign is a work register, and an unmounted volume is a
|
||||
// legitimate reason for a path to be absent today and present tomorrow.
|
||||
const { ledger } = await readLedger(file);
|
||||
assert.equal(ledger.repos.length, 2);
|
||||
});
|
||||
|
||||
it('refresh-tokens skips an unreadable repo instead of counting it as 0 tokens', async () => {
|
||||
const dir = newDir();
|
||||
const file = join(dir, 'ledger.json');
|
||||
const phantom = join(dir, 'no-such-repo');
|
||||
|
||||
runWrite(['add', phantom, '--ledger-file', file, '--reference-date', NOW]);
|
||||
const { status, stdout } = runWrite([
|
||||
'refresh-tokens', '--ledger-file', file, '--reference-date', NOW,
|
||||
]);
|
||||
assert.equal(status, 0);
|
||||
|
||||
const out = JSON.parse(stdout);
|
||||
assert.deepEqual(out.swept, [], 'a repo that cannot be read was never actually swept');
|
||||
assert.equal(out.skipped.length, 1, 'it belongs in skipped, with a reason the user can act on');
|
||||
assert.equal(out.skipped[0].path, resolve(phantom));
|
||||
assert.match(out.skipped[0].reason, /not readable|does not exist|ENOENT/i);
|
||||
assert.equal(
|
||||
out.rollUp.tokens.reposWithTokens,
|
||||
0,
|
||||
'the machine-wide bill must not claim coverage of a repo it could not read',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
|
|
|||
112
tests/scanners/cli-unknown-flag-rejection.test.mjs
Normal file
112
tests/scanners/cli-unknown-flag-rejection.test.mjs
Normal file
|
|
@ -0,0 +1,112 @@
|
|||
/**
|
||||
* Session #51 — CLIs must reject an unknown flag, never ignore it.
|
||||
*
|
||||
* Third arm of the argument-handling class first measured in #44/#47/#50. The
|
||||
* earlier arms were about a flag's VALUE being swallowed as the scan target
|
||||
* (`else if (!args[i].startsWith('-')) targetPath = args[i]`). This arm is
|
||||
* quieter and worse: several CLIs have no `else` branch at all, so an
|
||||
* unrecognised flag falls out of the parse loop leaving no trace — exit 0, a
|
||||
* full payload, and an answer to a question the caller did not ask.
|
||||
*
|
||||
* Measured cost, live, in the same session that wrote this test: the
|
||||
* `knowledge-refresh` command's only user-facing knob (`--stale-after N`)
|
||||
* reached the CLI as one malformed argv entry (see
|
||||
* command-flag-value-portability.test.mjs). Because the CLI ignored it, the
|
||||
* command reported "✓ All 14 register entries were re-verified within the last
|
||||
* 90 days" — a true-sounding sentence about a threshold the user had just
|
||||
* overridden. Had the CLI failed loudly, the shell bug would have been a
|
||||
* one-line exit-3 message instead of a silent wrong answer.
|
||||
*
|
||||
* Scope of this guard: the CLIs whose commands were dogfooded in this chunk.
|
||||
* `optimize-lens-cli.mjs` and `token-hotspots-cli.mjs` share the defect but
|
||||
* also carry the still-open positional-swallow arm; both are fixed together in
|
||||
* the v5.14 arg-handling chunk, where every call site's flags can be audited at
|
||||
* once. They are listed in KNOWN_OPEN so the number stays visible rather than
|
||||
* being quietly rounded down to zero.
|
||||
*/
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { spawn } from 'node:child_process';
|
||||
import { readFile } from 'node:fs/promises';
|
||||
import { resolve, dirname } from 'node:path';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
|
||||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||
const SCANNERS_DIR = resolve(__dirname, '..', '..', 'scanners');
|
||||
|
||||
/** CLIs this guard holds to the invariant, with the argv they need to get past required-arg checks. */
|
||||
const GUARDED = [
|
||||
{ cli: 'campaign-cli.mjs', argv: [] },
|
||||
{ cli: 'knowledge-refresh-cli.mjs', argv: [] },
|
||||
{ cli: 'campaign-write-cli.mjs', argv: ['init'] },
|
||||
{ cli: 'campaign-export-cli.mjs', argv: ['--repo', '.'] },
|
||||
];
|
||||
|
||||
/** Same defect, deferred to the v5.14 arg-handling chunk together with their positional-swallow arm. */
|
||||
const KNOWN_OPEN = ['optimize-lens-cli.mjs', 'token-hotspots-cli.mjs'];
|
||||
|
||||
function run(cli, argv) {
|
||||
return new Promise((res) => {
|
||||
const child = spawn(process.execPath, [resolve(SCANNERS_DIR, cli), ...argv], {
|
||||
cwd: resolve(__dirname, '..', '..'),
|
||||
});
|
||||
let stderr = '';
|
||||
child.stderr.on('data', (d) => { stderr += d; });
|
||||
child.stdout.on('data', () => {});
|
||||
child.on('close', (code) => res({ code, stderr }));
|
||||
});
|
||||
}
|
||||
|
||||
for (const { cli, argv } of GUARDED) {
|
||||
test(`${cli} rejects an unknown flag with exit 3`, async () => {
|
||||
const { code, stderr } = await run(cli, [...argv, '--zzz-not-a-real-flag']);
|
||||
|
||||
assert.equal(
|
||||
code,
|
||||
3,
|
||||
`${cli} accepted an unknown flag (exit ${code}). A flag the CLI does not understand\n` +
|
||||
'must fail loudly — silently ignoring it turns a caller-side bug into a confident\n' +
|
||||
'wrong answer. stderr was: ' + JSON.stringify(stderr),
|
||||
);
|
||||
assert.match(
|
||||
stderr,
|
||||
/--zzz-not-a-real-flag/,
|
||||
`${cli} must name the offending flag so the caller can find it.`,
|
||||
);
|
||||
});
|
||||
}
|
||||
|
||||
test('the deferred CLIs are still deferred, and still counted', () => {
|
||||
assert.equal(
|
||||
KNOWN_OPEN.length,
|
||||
2,
|
||||
'When the v5.14 arg-handling chunk closes optimize-lens-cli and token-hotspots-cli,\n' +
|
||||
'move them from KNOWN_OPEN into GUARDED rather than deleting them — the count is\n' +
|
||||
'the record of how wide the class was.',
|
||||
);
|
||||
});
|
||||
|
||||
/**
|
||||
* The caller arm. #45/#46/#47 all taught the same lesson: fixing a CLI does not fix the
|
||||
* command that reads its payload. `addedUnverified` and the `skipped` reasons only reach the
|
||||
* user if the command template is told to report them — otherwise the CLI is honest into a
|
||||
* void, and the phantom repo is just as invisible as before.
|
||||
*/
|
||||
const CAMPAIGN_MD = resolve(__dirname, '..', '..', 'commands', 'campaign.md');
|
||||
|
||||
test('campaign.md reports the fields the write-CLI added for honest coverage', async () => {
|
||||
const content = await readFile(CAMPAIGN_MD, 'utf-8');
|
||||
|
||||
assert.match(
|
||||
content,
|
||||
/addedUnverified/,
|
||||
'campaign.md must report `addedUnverified` after an add — a tracked path the CLI could\n' +
|
||||
'not read is exactly the row that silently pollutes the backlog and the token bill.',
|
||||
);
|
||||
assert.match(
|
||||
content,
|
||||
/skipped/,
|
||||
'campaign.md must report `skipped[]` after a token sweep so the bill\'s coverage is honest.',
|
||||
);
|
||||
});
|
||||
127
tests/scanners/output-file-robustness.test.mjs
Normal file
127
tests/scanners/output-file-robustness.test.mjs
Normal file
|
|
@ -0,0 +1,127 @@
|
|||
/**
|
||||
* Session #51 — `--output-file` must be usable, and a crash must never look normal.
|
||||
*
|
||||
* Two defects found while sweeping the campaign/knowledge-refresh chunk, both about the
|
||||
* seam every command depends on: the command runs a scanner with `--output-file <path>
|
||||
* 2>/dev/null`, checks the exit code, and Reads the file (ux-rules 2-4).
|
||||
*
|
||||
* 1. NO SCANNER CREATES THE PARENT DIRECTORY. `saveLedger` does; the payload write does
|
||||
* not. `commands/campaign.md` step 2 writes to
|
||||
* `~/.claude/config-audit/sessions/campaign-report.json` — on a machine where that
|
||||
* directory does not exist yet (the FIRST run, exactly the case campaign-cli otherwise
|
||||
* handles gracefully with `initialized:false`) the write throws ENOENT, and the command's
|
||||
* own exit-code table then tells the user "the campaign ledger couldn't be read — it may
|
||||
* be corrupt. Stop; do not attempt a write over a corrupt ledger." The ledger is not
|
||||
* corrupt; it does not exist. The user is steered away from the one action that helps.
|
||||
*
|
||||
* 2. `posture.mjs` REPORTED A CRASH AS A PASSING GRADE. Its top-level catch set
|
||||
* `process.exitCode = 1`, and every command in this plugin is instructed that "codes 0,
|
||||
* 1, 2 are normal (PASS/WARNING/FAIL). Only 3 is a real error." So a fatal error was
|
||||
* indistinguishable from a WARNING grade — measured live: posture exited 1, wrote no
|
||||
* file, and the command would have gone on to Read a file that was never created.
|
||||
* Every other scanner used 3; posture was the single outlier (1 of 14, measured).
|
||||
*/
|
||||
|
||||
import { test } from 'node:test';
|
||||
import { strict as assert } from 'node:assert';
|
||||
import { spawn } from 'node:child_process';
|
||||
import { readFile, mkdtemp, mkdir, writeFile, readdir } from 'node:fs/promises';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { resolve, dirname, join } from 'node:path';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { fileURLToPath } from 'node:url';
|
||||
|
||||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||
const ROOT = resolve(__dirname, '..', '..');
|
||||
const SCANNERS_DIR = resolve(ROOT, 'scanners');
|
||||
|
||||
/**
|
||||
* Every scanner that writes a `--output-file` payload, with argv that reaches the write.
|
||||
* Determined empirically (each one was confirmed to produce the file when the parent
|
||||
* directory already exists); `self-audit.mjs` is absent because it has no such flag.
|
||||
*/
|
||||
const WRITERS = [
|
||||
{ cli: 'campaign-cli.mjs', argv: () => ['--ledger-file', LEDGER] },
|
||||
{ cli: 'campaign-write-cli.mjs', argv: (d) => ['init', '--ledger-file', join(d, 'l.json')] },
|
||||
{ cli: 'campaign-export-cli.mjs', argv: () => ['--repo', REPO, '--ledger-file', LEDGER] },
|
||||
{ cli: 'knowledge-refresh-cli.mjs', argv: () => [] },
|
||||
{ cli: 'optimize-lens-cli.mjs', argv: () => [ROOT] },
|
||||
{ cli: 'token-hotspots-cli.mjs', argv: () => [ROOT] },
|
||||
{ cli: 'drift-cli.mjs', argv: () => [ROOT] },
|
||||
{ cli: 'fix-cli.mjs', argv: () => [ROOT] },
|
||||
{ cli: 'manifest.mjs', argv: () => [ROOT] },
|
||||
{ cli: 'posture.mjs', argv: () => [ROOT] },
|
||||
{ cli: 'whats-active.mjs', argv: () => [ROOT] },
|
||||
{ cli: 'plugin-health-scanner.mjs', argv: () => [ROOT] },
|
||||
{ cli: 'scan-orchestrator.mjs', argv: () => [ROOT] },
|
||||
];
|
||||
|
||||
let LEDGER;
|
||||
let REPO;
|
||||
|
||||
function run(cli, argv) {
|
||||
return new Promise((res) => {
|
||||
const child = spawn(process.execPath, [resolve(SCANNERS_DIR, cli), ...argv], { cwd: ROOT });
|
||||
child.stdout.on('data', () => {});
|
||||
child.stderr.on('data', () => {});
|
||||
child.on('close', (code) => res(code));
|
||||
});
|
||||
}
|
||||
|
||||
test('every --output-file writer creates its parent directory', async (t) => {
|
||||
const base = await mkdtemp(join(tmpdir(), 'ca-outfile-'));
|
||||
// A tracked repo + ledger so the campaign CLIs get past their own gates and reach the write.
|
||||
REPO = join(base, 'repo');
|
||||
await mkdir(REPO, { recursive: true });
|
||||
LEDGER = join(base, 'ledger.json');
|
||||
await writeFile(
|
||||
LEDGER,
|
||||
JSON.stringify({
|
||||
schemaVersion: 1, createdDate: '2026-06-22', updatedDate: '2026-06-22',
|
||||
repos: [{ path: REPO, name: 'repo', status: 'pending', sessionId: null, findingsBySeverity: null, tokens: null, updatedDate: '2026-06-22' }],
|
||||
}, null, 2),
|
||||
'utf-8',
|
||||
);
|
||||
|
||||
const failures = [];
|
||||
for (const { cli, argv } of WRITERS) {
|
||||
const dir = join(base, `work-${cli}`);
|
||||
await mkdir(dir, { recursive: true });
|
||||
// The parent of the output file deliberately does not exist.
|
||||
const out = join(dir, 'not', 'created', 'yet', 'payload.json');
|
||||
const code = await run(cli, [...argv(dir), '--output-file', out]);
|
||||
if (!existsSync(out)) failures.push(`${cli} (exit ${code})`);
|
||||
}
|
||||
|
||||
assert.deepEqual(
|
||||
failures,
|
||||
[],
|
||||
'These scanners crash instead of creating the output path. A command that writes its\n' +
|
||||
'payload under ~/.claude/config-audit/sessions/ on a fresh machine gets ENOENT and\n' +
|
||||
'reports it as a corrupt/unreadable input. Add mkdir(dirname(outputFile),\n' +
|
||||
'{recursive:true}) before the write:\n ' + failures.join('\n '),
|
||||
);
|
||||
});
|
||||
|
||||
test('a fatal error exits 3 — never a code the commands treat as a normal grade', async () => {
|
||||
const entries = (await readdir(SCANNERS_DIR)).filter((f) => f.endsWith('.mjs')).sort();
|
||||
const offenders = [];
|
||||
|
||||
for (const file of entries) {
|
||||
const src = await readFile(resolve(SCANNERS_DIR, file), 'utf-8');
|
||||
const idx = src.indexOf('main().catch');
|
||||
if (idx === -1) continue;
|
||||
const block = src.slice(idx, idx + 400);
|
||||
const m = block.match(/process\.exitCode\s*=\s*(\d+)/);
|
||||
if (m && m[1] !== '3') offenders.push(`${file}: exitCode = ${m[1]}`);
|
||||
}
|
||||
|
||||
assert.deepEqual(
|
||||
offenders,
|
||||
[],
|
||||
'ux-rules tells every command that exit 0/1/2 are normal results (PASS/WARNING/FAIL)\n' +
|
||||
'and only 3 is a real error. A fatal catch that sets anything else makes a crash\n' +
|
||||
'indistinguishable from a grade, and the command goes on to Read a file that was\n' +
|
||||
'never written:\n ' + offenders.join('\n '),
|
||||
);
|
||||
});
|
||||
Loading…
Add table
Add a link
Reference in a new issue