Q1 was gate-in-prose, Q2 was argv-in-prose; the sweep's answer to 'what is the third' is data-contract-in-prose: templates hand-build files (backup manifest, session state) that engines and hooks later parse. Rated list of 9, topped by two recovery-path findings: the rollback engine has no CLI entry (restore runs as model prose with a pre-rendered 'checksum verified' line), and implement's hand-built manifest format is pinned only by a hand-written fixture — the seam that already produced one success-shaped no-op. Every number in the doc is command-produced (Fable session, no advisor). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BeVSUuNqSCDUgTvfLKv4fK
15 KiB
Q_AUDIT — the prose-invariant sweep (v6 quality plan §3)
Session #68, 2026-08-12, Fable 5/xhigh (no advisor — every number below is
command-produced; the commands are in the appendix). Q1 closed the write gate
in code, Q2 closed the argv contract in code. This sweep asked what ELSE is
enforced only by prose across the three surfaces the plan names: 21 command
templates, 7 agent prompts, 4 .claude/rules/ files (4,950 lines), against
the 17 guard files that already exist under tests/commands/ + tests/agents/.
Output is a rated list, not code. Nothing here was fixed in this session.
1. The taxonomy — what counts as a prose invariant
The open decision this session owned. A statement counts when all three hold:
- Something else relies on it — another component (code, hook, agent, downstream command) behaves correctly only while the statement is true.
- Its violation is silent — nothing fails loudly when it stops holding.
- No code or test detects the violation.
What deliberately does NOT count (the negative list matters — Q1's own lesson is that a gate firing on legitimate writes gets switched off):
- Judgment rubrics given to agents (analyzer's 100-point CLAUDE.md rubric, planner's risk formula) — prose is the medium of a judgment task, not a bug.
- Narration/UX rules (ux-rules.md) — degraded output, self-correcting.
- Output budgets ("MUST NOT exceed 300 lines") — worst case is a long report.
- Harness facts ("the Write tool requires a prior Read") — enforced upstream.
Marker-grep is a non-detector here: only 18 MUST/NEVER/ALWAYS-class markers exist across all 28 command+agent files. The invariants are procedural steps whose omission is silent, not shouted rules. They are found by reading, which is why this was a session, not a script.
2. The class finding — the THIRD instance
Q1 was gate-in-prose (policy paraphrased in five templates). Q2 was argv-in-prose (caller contract unchecked, 54 pairs). The third instance this sweep asked for is:
Class 3 — data-contract-in-prose: prose WRITES what code READS.
Q2's mirror image. Templates instruct the model to hand-build files —
manifest.yaml, state.yaml, scope.yaml, register edits — that engines,
hooks, and later commands then parse. The writer side is a prose schema; the
reader side either trusts it or has quietly learned one measured variant of it.
The class has already bitten once: parseManifest grew its second format branch
after implement-produced backups made restoreBackup "a success-shaped no-op"
(comment in scanners/lib/backup.mjs:172).
Two further classes surfaced (the sweep found things it wasn't looking for, again): Class 4 — contracts an agent cannot honor (instructions colliding with harness behavior or the agent's own declared tools), and Class 5 — knowledge tables duplicated between prose and code.
3. The rated list
Rated by what breaks if the invariant silently stops holding. R1/R2 outrank everything because they sit on the recovery path — the surface that runs exactly when the user is already in trouble, and the least-exercised one.
R1 — the restore flow is model-executed prose; the code engine has no CLI entry
Where: commands/rollback.md §Implementation; scanners/rollback-engine.mjs.
Measured: 16 scanner CLIs carry a process.argv entry — rollback-engine.mjs
is not one of them. The template's "Implementation" shows an ESM import block a
command template cannot execute, then offers ad-hoc cp as the runnable
alternative. The engine's restoreBackup() verifies checksums before AND after
each write and returns createdNotRemoved — none of it reachable from the
command without node -e. The success template pre-renders "(checksum verified)" — a claim the runnable path never establishes.
What breaks: the recovery path for every other write the plugin makes. A
half-restore or a stale-backup restore lands on user config at the worst
possible moment, reported as verified.
Also unguardable as-is: command-cli-contract.test.mjs probes CLIs; with no
CLI here, the whole Q2 guard class is structurally blind to this command.
Guard shape: a thin rollback-cli.mjs over listBackups/restoreBackup/
deleteBackup; template calls it like every other command; the contract test
then covers it for free. The "(checksum verified)" line becomes payload-driven.
R2 — implement's backup manifest is hand-built prose; the parser knows one frozen sample of it
Where: commands/implement.md Step 3 (mkdir/cp + hand-written
manifest.yaml with sha256 lines); scanners/lib/backup.mjs parseManifest;
tests/scanners/rollback-paths.test.mjs:157-186.
Measured: the fix pipeline's backups go through code (fix-engine.mjs:10
imports createBackup); the implement pipeline's backup is template prose —
two copies of the backup policy, one per pipeline. parseManifest's
implement-format branch exists because the seam already failed silently once.
The test fixture pinning that format is hand-written, not derived from the
template's own example block — the exact #63 defect shape ("a hand-typed call is
a path no user takes").
What breaks: implement.md's example drifts (a key renamed, quoting added) →
parseManifest finds 0 files → every implement backup is unrestorable while
rollback reports success. Second occurrence of a failure that already happened.
Guard shape: minimum — derive the parser fixture from implement.md's own
fenced YAML (the Q2 extractor trick pointed at an example block). Real fix —
implement's Step 3 calls the same createBackup code path fix already uses
(via the R1 CLI), and the prose format dies entirely.
R3 — optimize + feature-gap still instruct agents to write reports the harness blocks
Where: commands/optimize.md:121-128, commands/feature-gap.md:123;
agents optimization-lens-agent.md §Output, feature-gap-agent.md §Output.
Measured: both templates tell the agent to write *-report.md to the
session dir, then Read that file. The harness note measured on analyze
(M-BUG-18) blocks agent writes of exactly the report/findings file class;
analyze was converted to orchestrator-writes, these two arms were left open
(tracked in STATE as the open M-BUG-18 class — this rating is its severity call).
What breaks: the Read step fails or the model improvises a rescue; the
command's documented artifact (optimization-lens-report.md) may never exist.
User-visible flow breakage, no data corruption.
Guard shape: the analyze pattern, already proven: agent returns inline,
command persists; analyze-report-persistence.test.mjs is the template to copy.
R4 — verifier-agent contradicts itself, and its "Read-Only Guarantee" is unenforced
Where: agents/verifier-agent.md — §Output Format says "Append to:
implementation-log.md"; §Read-Only Guarantee says "only uses Read, Glob, Grep /
never modifies any files"; frontmatter tools: lists no write tool.
Measured: implement.md Step 5 was already fixed to return-inline and
append orchestrator-side — so the agent's system prompt and the spawn prompt now
give OPPOSITE instructions to the same agent. Harness enforcement of tools:
is measured absent (memory: verifier/implement-log writes went through live).
What breaks: which instruction wins is nondeterministic; if the agent
improvises a write to satisfy its own §Output Format, a full-file Write on the
shared log clobbers parallel implementer entries — the precise defect
implement-log-append.test.mjs exists to prevent, entering through the file
that test does not read.
Guard shape: rewrite verifier-agent's Output section to return-inline (one
file), and extend implement-log-append/agent-prompt-shape to assert no agent
prompt instructs appending to the shared log. Cheap.
R5 — session state (state.yaml, scope.yaml) is a model-written machine contract with no schema anywhere
Where: every phase template ("Write scope.yaml and state.yaml", "append —
never replace — completed_phases"), .claude/rules/state-management.md,
readers in hooks/scripts/session-start.mjs + stop-session-reminder.mjs +
every session-aware command.
Measured: the phase vocabulary (discover…verify) appears as a shared
constant in zero code files — it lives only in prose copies (status.md's
table, state-management.md, each template). Hooks parse with a line-grep
(parseYamlValue) and print whatever they find. The existing guard
(command-shell-state-shape: "phase commands name all four fields") checks the
template text, not the written file.
What breaks: resume-by-recency picks wrong sessions, status misnarrates,
session-start reminders go quiet — degradation, not corruption, but it erodes
exactly the "can resume if interrupted" promise the rule exists for.
Guard shape: either a state-write CLI (heavy) or a defensive reader: a lib
that validates phase tokens + required fields and flags malformed state, used
by hooks and dogfooded in a test. The plugin flags drift in everyone else's
config; its own session state deserves the same reader.
R6 — both "Required Frontmatter" contracts are unguarded (currently compliant)
Where: .claude/rules/agent-development.md, .claude/rules/command-development.md.
Measured: no test outside fixtures matches allowed-tools;
agent-prompt-shape asserts only name: on 3 of 7 agents. Measured today:
7/7 agent frontmatters match the CLAUDE.md table; duplicate colors: 0. So —
compliant, unwatched. Every MUST in those two rules is enforced by nothing.
What breaks: a new agent/command ships with missing allowed-tools or a
duplicate color; nothing fails; the rules files become fiction one file at a
time (the exemption-table lesson from Q1: what nothing declares, everyone
re-answers by reading).
Guard shape: near-free shape test walking agents/*.md + commands/*.md
asserting the two rules' required keys, name conventions, color uniqueness.
Note the irony budget: plugin-health-scanner already audits other plugins'
frontmatter — pointing it at its own repo in a test is the dogfood version.
R7 — secret detection exists only as agent prose, in a domain STATE has parked elsewhere
Where: agents/scanner-agent.md §Secret Detection Patterns (xoxb/sk-/ghp_
regexes); agents/verifier-agent.md Check 7 ("Secrets Scan ✓").
Measured: xoxb/ghp_ appear in zero files under scanners/;
mcp-config-validator.mjs contains the string "secret" zero times. The
deterministic pipeline has no secret scanning at all; the agent path claims it
in prose, and the verifier's report template renders "Secrets Scan ✓ Pass" as a
table row regardless. STATE parks secrets as the llm-security plugin's domain.
What breaks: a user reads "Secrets Scan ✓" as an executed check. The lie is
in the reporting, not in a missing feature — the feature is deliberately owned
elsewhere.
Guard shape: this is a removal candidate, not a gate (measurement can
decline the feature): strip the prose secret patterns + the verifier's Check 7,
say "secrets: out of scope, see llm-security" where the row used to be. If the
capability is ever wanted deterministically, it starts life as a scanner with a
finding code, not as agent prose.
R8 — knowledge tables duplicated between prose and code
Where/measured: managed-path table — three copies (scanner-agent prose +
file-discovery.mjs + active-config-reader.mjs). Precedence — analyzer prose
("global beats managed (user preference)") vs conflict-detector.mjs:138
("local > project > user"): different vocabularies for the same claim, no link.
Optimization-lens agent's mechanism table restates register entries
(BP-MECH-001/002/004, BP-SUB-001) that live as data in
knowledge/best-practices.json.
What breaks: slow divergence — an agent narrates precedence or hierarchy the
deterministic scanners no longer implement. Confusing, not corrupting.
Guard shape: two-copies rule applies but measure first (#67): the two code
copies may legitimately differ; the prose copies should cite the code as owner
("hierarchy per file-discovery.mjs") rather than restate values.
R9 — knowledge-refresh applies approved writes by model edit, validated only afterwards
Where: commands/knowledge-refresh.md Step 6 (model Edit of
best-practices.json, then run the schema test and revert on failure).
Measured: already tracked open in STATE ("knowledge-refresh skrive-CLI");
path anchoring is guarded (knowledge-refresh-write-target.test.mjs), the write
itself is not. The post-hoc validation step is real mitigation — this ranks
last because the failure is loud (a failing schema test in the same flow).
Guard shape: the campaign pattern, which this command's own sibling already
implements: every mutation a subcommand of a write-CLI. campaign.md is the
in-repo proof that "human-approved writes" and "CLI-executed writes" compose.
4. What this changes in the plan
- Q3 (severity axis) should carry R1/R2 as its first two rows — recovery-path items never live in a backlog paragraph (plan §2 property 2).
- Q4 (v6.0.0 release) is NOT blocked by this list (the release gate is Q1 + green suite); but R1+R2 are the strongest candidates for the first post-v6 chunk, as one chunk: a rollback/backup CLI closes both, and converts R2's guard from "pin the prose format" to "delete the prose format".
- R4 + R6 are lunch-sized; they can ride along with any adjacent session the way M-BUG fixes have.
- R7 is an operator decision (removing a claimed capability): propose, don't do.
Appendix — measurement log
Every number above, and the command that produced it (run at 30c78ae):
| # | Number | Command |
|---|---|---|
| 1 | 21 / 7 / 4 files | ls commands/*.md | wc -l etc. |
| 2 | 4,950 lines | wc -l commands/*.md agents/*.md .claude/rules/*.md |
| 3 | 17 guard files | find tests/commands tests/agents -name '*.test.mjs' | wc -l |
| 4 | 18 markers | grep -cE '\b(MUST|NEVER|ALWAYS|DO NOT|…)\b' commands/*.md agents/*.md |
| 5 | 16 CLIs, rollback-engine absent | grep -ln "process.argv" scanners/*.mjs |
| 6 | parseManifest dual-format + its cause | sed -n '140,220p' scanners/lib/backup.mjs (comment at :172) |
| 7 | hand-written fixture | grep -n "sha256" tests/scanners/rollback-paths.test.mjs (:159-186) |
| 8 | fix uses code backup | grep -n "backup" scanners/fix-engine.mjs (:10) |
| 9 | 0 secret patterns in code | grep -rln "xoxb|ghp_" scanners/ (empty); grep -n "secret" scanners/mcp-config-validator.mjs (empty) |
| 10 | 3 copies managed-path table | grep -rln "Library/Application Support" scanners/ (2) + scanner-agent prose (1) |
| 11 | 0 phase-vocabulary constants | grep -rn "'discover'" scanners/lib/*.mjs hooks/scripts/*.mjs (empty) |
| 12 | 3/7 agents in shape guard | AGENT_FILES array read in tests/agents/agent-prompt-shape.test.mjs |
| 13 | 0 frontmatter guards | grep -rln "allowed-tools" tests/ (fixtures + yaml-parser only) |
| 14 | 0 duplicate colors | grep -h "^color:" agents/*.md | sort | uniq -d | wc -l |
| 15 | verifier self-contradiction | Read agents/verifier-agent.md (:143 append vs :242 read-only) |
Not covered by this sweep (deliberate): knowledge/*.md content freshness
(knowledge-refresh's domain), README/plugin.json surface (repo-standard's
domain), the deterministic scanners themselves (Q1/Q2 territory, already coded).