docs(plan): the prose-invariant sweep finds the class's third instance (Q_AUDIT)

Q1 was gate-in-prose, Q2 was argv-in-prose; the sweep's answer to 'what is
the third' is data-contract-in-prose: templates hand-build files (backup
manifest, session state) that engines and hooks later parse. Rated list of
9, topped by two recovery-path findings: the rollback engine has no CLI
entry (restore runs as model prose with a pre-rendered 'checksum verified'
line), and implement's hand-built manifest format is pinned only by a
hand-written fixture — the seam that already produced one success-shaped
no-op. Every number in the doc is command-produced (Fable session, no
advisor).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BeVSUuNqSCDUgTvfLKv4fK
This commit is contained in:
Kjell Tore Guttormsen 2026-08-12 22:14:48 +02:00
commit dc800560b7

View file

@ -0,0 +1,242 @@
# Q_AUDIT — the prose-invariant sweep (v6 quality plan §3)
**Session #68, 2026-08-12, Fable 5/xhigh (no advisor — every number below is
command-produced; the commands are in the appendix).** Q1 closed the write gate
in code, Q2 closed the argv contract in code. This sweep asked what ELSE is
enforced only by prose across the three surfaces the plan names: **21** command
templates, **7** agent prompts, **4** `.claude/rules/` files (4,950 lines), against
the **17** guard files that already exist under `tests/commands/` + `tests/agents/`.
Output is a rated list, not code. Nothing here was fixed in this session.
## 1. The taxonomy — what counts as a prose invariant
The open decision this session owned. A statement counts when all three hold:
1. **Something else relies on it** — another component (code, hook, agent,
downstream command) behaves correctly only while the statement is true.
2. **Its violation is silent** — nothing fails loudly when it stops holding.
3. **No code or test detects the violation.**
What deliberately does NOT count (the negative list matters — Q1's own lesson is
that a gate firing on legitimate writes gets switched off):
- **Judgment rubrics** given to agents (analyzer's 100-point CLAUDE.md rubric,
planner's risk formula) — prose is the *medium* of a judgment task, not a bug.
- **Narration/UX rules** (ux-rules.md) — degraded output, self-correcting.
- **Output budgets** ("MUST NOT exceed 300 lines") — worst case is a long report.
- **Harness facts** ("the Write tool requires a prior Read") — enforced upstream.
Marker-grep is a non-detector here: only **18** MUST/NEVER/ALWAYS-class markers
exist across all 28 command+agent files. The invariants are procedural steps
whose omission is silent, not shouted rules. They are found by reading, which is
why this was a session, not a script.
## 2. The class finding — the THIRD instance
Q1 was *gate-in-prose* (policy paraphrased in five templates). Q2 was
*argv-in-prose* (caller contract unchecked, 54 pairs). The third instance this
sweep asked for is:
**Class 3 — data-contract-in-prose: prose WRITES what code READS.**
Q2's mirror image. Templates instruct the model to hand-build files —
`manifest.yaml`, `state.yaml`, `scope.yaml`, register edits — that engines,
hooks, and later commands then parse. The writer side is a prose schema; the
reader side either trusts it or has quietly learned one measured variant of it.
The class has already bitten once: `parseManifest` grew its second format branch
*after* implement-produced backups made `restoreBackup` "a success-shaped no-op"
(comment in `scanners/lib/backup.mjs:172`).
Two further classes surfaced (the sweep found things it wasn't looking for,
again): **Class 4 — contracts an agent cannot honor** (instructions colliding
with harness behavior or the agent's own declared tools), and **Class 5 —
knowledge tables duplicated between prose and code**.
## 3. The rated list
Rated by what breaks if the invariant silently stops holding. R1/R2 outrank
everything because they sit on the *recovery* path — the surface that runs
exactly when the user is already in trouble, and the least-exercised one.
### R1 — the restore flow is model-executed prose; the code engine has no CLI entry
**Where:** `commands/rollback.md` §Implementation; `scanners/rollback-engine.mjs`.
**Measured:** 16 scanner CLIs carry a `process.argv` entry — `rollback-engine.mjs`
is not one of them. The template's "Implementation" shows an ESM `import` block a
command template cannot execute, then offers ad-hoc `cp` as the runnable
alternative. The engine's `restoreBackup()` verifies checksums before AND after
each write and returns `createdNotRemoved` — none of it reachable from the
command without `node -e`. The success template pre-renders "`(checksum
verified)`" — a claim the runnable path never establishes.
**What breaks:** the recovery path for every other write the plugin makes. A
half-restore or a stale-backup restore lands on user config at the worst
possible moment, reported as verified.
**Also unguardable as-is:** `command-cli-contract.test.mjs` probes CLIs; with no
CLI here, the whole Q2 guard class is structurally blind to this command.
**Guard shape:** a thin `rollback-cli.mjs` over `listBackups`/`restoreBackup`/
`deleteBackup`; template calls it like every other command; the contract test
then covers it for free. The "(checksum verified)" line becomes payload-driven.
### R2 — implement's backup manifest is hand-built prose; the parser knows one frozen sample of it
**Where:** `commands/implement.md` Step 3 (mkdir/cp + hand-written
`manifest.yaml` with sha256 lines); `scanners/lib/backup.mjs` `parseManifest`;
`tests/scanners/rollback-paths.test.mjs:157-186`.
**Measured:** the fix pipeline's backups go through code (`fix-engine.mjs:10`
imports `createBackup`); the implement pipeline's backup is template prose —
two copies of the backup policy, one per pipeline. `parseManifest`'s
implement-format branch exists because the seam already failed silently once.
The test fixture pinning that format is **hand-written**, not derived from the
template's own example block — the exact #63 defect shape ("a hand-typed call is
a path no user takes").
**What breaks:** implement.md's example drifts (a key renamed, quoting added) →
`parseManifest` finds 0 files → every implement backup is unrestorable while
`rollback` reports success. Second occurrence of a failure that already happened.
**Guard shape:** minimum — derive the parser fixture from `implement.md`'s own
fenced YAML (the Q2 extractor trick pointed at an example block). Real fix —
implement's Step 3 calls the same `createBackup` code path fix already uses
(via the R1 CLI), and the prose format dies entirely.
### R3 — optimize + feature-gap still instruct agents to write reports the harness blocks
**Where:** `commands/optimize.md:121-128`, `commands/feature-gap.md:123`;
agents `optimization-lens-agent.md` §Output, `feature-gap-agent.md` §Output.
**Measured:** both templates tell the agent to write `*-report.md` to the
session dir, then Read that file. The harness note measured on `analyze`
(M-BUG-18) blocks agent writes of exactly the report/findings file class;
`analyze` was converted to orchestrator-writes, these two arms were left open
(tracked in STATE as the open M-BUG-18 class — this rating is its severity call).
**What breaks:** the Read step fails or the model improvises a rescue; the
command's documented artifact (`optimization-lens-report.md`) may never exist.
User-visible flow breakage, no data corruption.
**Guard shape:** the `analyze` pattern, already proven: agent returns inline,
command persists; `analyze-report-persistence.test.mjs` is the template to copy.
### R4 — verifier-agent contradicts itself, and its "Read-Only Guarantee" is unenforced
**Where:** `agents/verifier-agent.md` — §Output Format says "Append to:
implementation-log.md"; §Read-Only Guarantee says "only uses Read, Glob, Grep /
never modifies any files"; frontmatter `tools:` lists no write tool.
**Measured:** `implement.md` Step 5 was already fixed to return-inline and
append orchestrator-side — so the agent's system prompt and the spawn prompt now
give OPPOSITE instructions to the same agent. Harness enforcement of `tools:`
is measured absent (memory: verifier/implement-log writes went through live).
**What breaks:** which instruction wins is nondeterministic; if the agent
improvises a write to satisfy its own §Output Format, a full-file Write on the
shared log clobbers parallel implementer entries — the precise defect
`implement-log-append.test.mjs` exists to prevent, entering through the file
that test does not read.
**Guard shape:** rewrite verifier-agent's Output section to return-inline (one
file), and extend `implement-log-append`/`agent-prompt-shape` to assert no agent
prompt instructs appending to the shared log. Cheap.
### R5 — session state (`state.yaml`, `scope.yaml`) is a model-written machine contract with no schema anywhere
**Where:** every phase template ("Write scope.yaml and state.yaml", "append —
never replace — completed_phases"), `.claude/rules/state-management.md`,
readers in `hooks/scripts/session-start.mjs` + `stop-session-reminder.mjs` +
every session-aware command.
**Measured:** the phase vocabulary (`discover``verify`) appears as a shared
constant in **zero** code files — it lives only in prose copies (status.md's
table, state-management.md, each template). Hooks parse with a line-grep
(`parseYamlValue`) and print whatever they find. The existing guard
(`command-shell-state-shape`: "phase commands name all four fields") checks the
template *text*, not the written *file*.
**What breaks:** resume-by-recency picks wrong sessions, status misnarrates,
session-start reminders go quiet — degradation, not corruption, but it erodes
exactly the "can resume if interrupted" promise the rule exists for.
**Guard shape:** either a state-write CLI (heavy) or a defensive reader: a lib
that validates phase tokens + required fields and *flags* malformed state, used
by hooks and dogfooded in a test. The plugin flags drift in everyone else's
config; its own session state deserves the same reader.
### R6 — both "Required Frontmatter" contracts are unguarded (currently compliant)
**Where:** `.claude/rules/agent-development.md`, `.claude/rules/command-development.md`.
**Measured:** no test outside fixtures matches `allowed-tools`;
`agent-prompt-shape` asserts only `name:` on **3 of 7** agents. Measured today:
7/7 agent frontmatters match the CLAUDE.md table; duplicate colors: **0**. So —
compliant, unwatched. Every MUST in those two rules is enforced by nothing.
**What breaks:** a new agent/command ships with missing `allowed-tools` or a
duplicate color; nothing fails; the rules files become fiction one file at a
time (the exemption-table lesson from Q1: what nothing declares, everyone
re-answers by reading).
**Guard shape:** near-free shape test walking `agents/*.md` + `commands/*.md`
asserting the two rules' required keys, name conventions, color uniqueness.
Note the irony budget: `plugin-health-scanner` already audits *other* plugins'
frontmatter — pointing it at its own repo in a test is the dogfood version.
### R7 — secret detection exists only as agent prose, in a domain STATE has parked elsewhere
**Where:** `agents/scanner-agent.md` §Secret Detection Patterns (xoxb/sk-/ghp_
regexes); `agents/verifier-agent.md` Check 7 ("Secrets Scan ✓").
**Measured:** `xoxb`/`ghp_` appear in **zero** files under `scanners/`;
`mcp-config-validator.mjs` contains the string "secret" **zero** times. The
deterministic pipeline has no secret scanning at all; the agent path claims it
in prose, and the verifier's report template renders "Secrets Scan ✓ Pass" as a
table row regardless. STATE parks secrets as the `llm-security` plugin's domain.
**What breaks:** a user reads "Secrets Scan ✓" as an executed check. The lie is
in the reporting, not in a missing feature — the feature is deliberately owned
elsewhere.
**Guard shape:** this is a *removal* candidate, not a gate (measurement can
decline the feature): strip the prose secret patterns + the verifier's Check 7,
say "secrets: out of scope, see llm-security" where the row used to be. If the
capability is ever wanted deterministically, it starts life as a scanner with a
finding code, not as agent prose.
### R8 — knowledge tables duplicated between prose and code
**Where/measured:** managed-path table — **three** copies (scanner-agent prose +
`file-discovery.mjs` + `active-config-reader.mjs`). Precedence — analyzer prose
("global beats managed (user preference)") vs `conflict-detector.mjs:138`
("local > project > user"): different vocabularies for the same claim, no link.
Optimization-lens agent's mechanism table restates register entries
(BP-MECH-001/002/004, BP-SUB-001) that live as data in
`knowledge/best-practices.json`.
**What breaks:** slow divergence — an agent narrates precedence or hierarchy the
deterministic scanners no longer implement. Confusing, not corrupting.
**Guard shape:** two-copies rule applies but *measure first* (#67): the two code
copies may legitimately differ; the prose copies should cite the code as owner
("hierarchy per `file-discovery.mjs`") rather than restate values.
### R9 — knowledge-refresh applies approved writes by model edit, validated only afterwards
**Where:** `commands/knowledge-refresh.md` Step 6 (model `Edit` of
`best-practices.json`, then run the schema test and revert on failure).
**Measured:** already tracked open in STATE ("knowledge-refresh skrive-CLI");
path anchoring is guarded (`knowledge-refresh-write-target.test.mjs`), the write
itself is not. The post-hoc validation step is real mitigation — this ranks
last *because* the failure is loud (a failing schema test in the same flow).
**Guard shape:** the campaign pattern, which this command's own sibling already
implements: every mutation a subcommand of a write-CLI. `campaign.md` is the
in-repo proof that "human-approved writes" and "CLI-executed writes" compose.
## 4. What this changes in the plan
- **Q3 (severity axis)** should carry R1/R2 as its first two rows — recovery-path
items never live in a backlog paragraph (plan §2 property 2).
- **Q4 (v6.0.0 release)** is NOT blocked by this list (the release gate is Q1 +
green suite); but R1+R2 are the strongest candidates for the first post-v6
chunk, as one chunk: a rollback/backup CLI closes both, and converts R2's
guard from "pin the prose format" to "delete the prose format".
- **R4 + R6** are lunch-sized; they can ride along with any adjacent session the
way M-BUG fixes have.
- **R7** is an operator decision (removing a claimed capability): propose, don't do.
## Appendix — measurement log
Every number above, and the command that produced it (run at `30c78ae`):
| # | Number | Command |
|---|--------|---------|
| 1 | 21 / 7 / 4 files | `ls commands/*.md \| wc -l` etc. |
| 2 | 4,950 lines | `wc -l commands/*.md agents/*.md .claude/rules/*.md` |
| 3 | 17 guard files | `find tests/commands tests/agents -name '*.test.mjs' \| wc -l` |
| 4 | 18 markers | `grep -cE '\b(MUST\|NEVER\|ALWAYS\|DO NOT\|…)\b' commands/*.md agents/*.md` |
| 5 | 16 CLIs, rollback-engine absent | `grep -ln "process.argv" scanners/*.mjs` |
| 6 | parseManifest dual-format + its cause | `sed -n '140,220p' scanners/lib/backup.mjs` (comment at :172) |
| 7 | hand-written fixture | `grep -n "sha256" tests/scanners/rollback-paths.test.mjs` (:159-186) |
| 8 | fix uses code backup | `grep -n "backup" scanners/fix-engine.mjs` (:10) |
| 9 | 0 secret patterns in code | `grep -rln "xoxb\|ghp_" scanners/` (empty); `grep -n "secret" scanners/mcp-config-validator.mjs` (empty) |
| 10 | 3 copies managed-path table | `grep -rln "Library/Application Support" scanners/` (2) + scanner-agent prose (1) |
| 11 | 0 phase-vocabulary constants | `grep -rn "'discover'" scanners/lib/*.mjs hooks/scripts/*.mjs` (empty) |
| 12 | 3/7 agents in shape guard | `AGENT_FILES` array read in `tests/agents/agent-prompt-shape.test.mjs` |
| 13 | 0 frontmatter guards | `grep -rln "allowed-tools" tests/` (fixtures + yaml-parser only) |
| 14 | 0 duplicate colors | `grep -h "^color:" agents/*.md \| sort \| uniq -d \| wc -l` |
| 15 | verifier self-contradiction | Read `agents/verifier-agent.md` (:143 append vs :242 read-only) |
Not covered by this sweep (deliberate): `knowledge/*.md` content freshness
(knowledge-refresh's domain), README/plugin.json surface (repo-standard's
domain), the deterministic scanners themselves (Q1/Q2 territory, already coded).