docs(plan): the prose-invariant sweep finds the class's third instance (Q_AUDIT)
Q1 was gate-in-prose, Q2 was argv-in-prose; the sweep's answer to 'what is the third' is data-contract-in-prose: templates hand-build files (backup manifest, session state) that engines and hooks later parse. Rated list of 9, topped by two recovery-path findings: the rollback engine has no CLI entry (restore runs as model prose with a pre-rendered 'checksum verified' line), and implement's hand-built manifest format is pinned only by a hand-written fixture — the seam that already produced one success-shaped no-op. Every number in the doc is command-produced (Fable session, no advisor). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BeVSUuNqSCDUgTvfLKv4fK
This commit is contained in:
parent
30c78aeda0
commit
dc800560b7
1 changed files with 242 additions and 0 deletions
242
docs/q-audit-prose-invariants.md
Normal file
242
docs/q-audit-prose-invariants.md
Normal file
|
|
@ -0,0 +1,242 @@
|
|||
# Q_AUDIT — the prose-invariant sweep (v6 quality plan §3)
|
||||
|
||||
**Session #68, 2026-08-12, Fable 5/xhigh (no advisor — every number below is
|
||||
command-produced; the commands are in the appendix).** Q1 closed the write gate
|
||||
in code, Q2 closed the argv contract in code. This sweep asked what ELSE is
|
||||
enforced only by prose across the three surfaces the plan names: **21** command
|
||||
templates, **7** agent prompts, **4** `.claude/rules/` files (4,950 lines), against
|
||||
the **17** guard files that already exist under `tests/commands/` + `tests/agents/`.
|
||||
|
||||
Output is a rated list, not code. Nothing here was fixed in this session.
|
||||
|
||||
## 1. The taxonomy — what counts as a prose invariant
|
||||
|
||||
The open decision this session owned. A statement counts when all three hold:
|
||||
|
||||
1. **Something else relies on it** — another component (code, hook, agent,
|
||||
downstream command) behaves correctly only while the statement is true.
|
||||
2. **Its violation is silent** — nothing fails loudly when it stops holding.
|
||||
3. **No code or test detects the violation.**
|
||||
|
||||
What deliberately does NOT count (the negative list matters — Q1's own lesson is
|
||||
that a gate firing on legitimate writes gets switched off):
|
||||
|
||||
- **Judgment rubrics** given to agents (analyzer's 100-point CLAUDE.md rubric,
|
||||
planner's risk formula) — prose is the *medium* of a judgment task, not a bug.
|
||||
- **Narration/UX rules** (ux-rules.md) — degraded output, self-correcting.
|
||||
- **Output budgets** ("MUST NOT exceed 300 lines") — worst case is a long report.
|
||||
- **Harness facts** ("the Write tool requires a prior Read") — enforced upstream.
|
||||
|
||||
Marker-grep is a non-detector here: only **18** MUST/NEVER/ALWAYS-class markers
|
||||
exist across all 28 command+agent files. The invariants are procedural steps
|
||||
whose omission is silent, not shouted rules. They are found by reading, which is
|
||||
why this was a session, not a script.
|
||||
|
||||
## 2. The class finding — the THIRD instance
|
||||
|
||||
Q1 was *gate-in-prose* (policy paraphrased in five templates). Q2 was
|
||||
*argv-in-prose* (caller contract unchecked, 54 pairs). The third instance this
|
||||
sweep asked for is:
|
||||
|
||||
**Class 3 — data-contract-in-prose: prose WRITES what code READS.**
|
||||
|
||||
Q2's mirror image. Templates instruct the model to hand-build files —
|
||||
`manifest.yaml`, `state.yaml`, `scope.yaml`, register edits — that engines,
|
||||
hooks, and later commands then parse. The writer side is a prose schema; the
|
||||
reader side either trusts it or has quietly learned one measured variant of it.
|
||||
The class has already bitten once: `parseManifest` grew its second format branch
|
||||
*after* implement-produced backups made `restoreBackup` "a success-shaped no-op"
|
||||
(comment in `scanners/lib/backup.mjs:172`).
|
||||
|
||||
Two further classes surfaced (the sweep found things it wasn't looking for,
|
||||
again): **Class 4 — contracts an agent cannot honor** (instructions colliding
|
||||
with harness behavior or the agent's own declared tools), and **Class 5 —
|
||||
knowledge tables duplicated between prose and code**.
|
||||
|
||||
## 3. The rated list
|
||||
|
||||
Rated by what breaks if the invariant silently stops holding. R1/R2 outrank
|
||||
everything because they sit on the *recovery* path — the surface that runs
|
||||
exactly when the user is already in trouble, and the least-exercised one.
|
||||
|
||||
### R1 — the restore flow is model-executed prose; the code engine has no CLI entry
|
||||
**Where:** `commands/rollback.md` §Implementation; `scanners/rollback-engine.mjs`.
|
||||
**Measured:** 16 scanner CLIs carry a `process.argv` entry — `rollback-engine.mjs`
|
||||
is not one of them. The template's "Implementation" shows an ESM `import` block a
|
||||
command template cannot execute, then offers ad-hoc `cp` as the runnable
|
||||
alternative. The engine's `restoreBackup()` verifies checksums before AND after
|
||||
each write and returns `createdNotRemoved` — none of it reachable from the
|
||||
command without `node -e`. The success template pre-renders "`(checksum
|
||||
verified)`" — a claim the runnable path never establishes.
|
||||
**What breaks:** the recovery path for every other write the plugin makes. A
|
||||
half-restore or a stale-backup restore lands on user config at the worst
|
||||
possible moment, reported as verified.
|
||||
**Also unguardable as-is:** `command-cli-contract.test.mjs` probes CLIs; with no
|
||||
CLI here, the whole Q2 guard class is structurally blind to this command.
|
||||
**Guard shape:** a thin `rollback-cli.mjs` over `listBackups`/`restoreBackup`/
|
||||
`deleteBackup`; template calls it like every other command; the contract test
|
||||
then covers it for free. The "(checksum verified)" line becomes payload-driven.
|
||||
|
||||
### R2 — implement's backup manifest is hand-built prose; the parser knows one frozen sample of it
|
||||
**Where:** `commands/implement.md` Step 3 (mkdir/cp + hand-written
|
||||
`manifest.yaml` with sha256 lines); `scanners/lib/backup.mjs` `parseManifest`;
|
||||
`tests/scanners/rollback-paths.test.mjs:157-186`.
|
||||
**Measured:** the fix pipeline's backups go through code (`fix-engine.mjs:10`
|
||||
imports `createBackup`); the implement pipeline's backup is template prose —
|
||||
two copies of the backup policy, one per pipeline. `parseManifest`'s
|
||||
implement-format branch exists because the seam already failed silently once.
|
||||
The test fixture pinning that format is **hand-written**, not derived from the
|
||||
template's own example block — the exact #63 defect shape ("a hand-typed call is
|
||||
a path no user takes").
|
||||
**What breaks:** implement.md's example drifts (a key renamed, quoting added) →
|
||||
`parseManifest` finds 0 files → every implement backup is unrestorable while
|
||||
`rollback` reports success. Second occurrence of a failure that already happened.
|
||||
**Guard shape:** minimum — derive the parser fixture from `implement.md`'s own
|
||||
fenced YAML (the Q2 extractor trick pointed at an example block). Real fix —
|
||||
implement's Step 3 calls the same `createBackup` code path fix already uses
|
||||
(via the R1 CLI), and the prose format dies entirely.
|
||||
|
||||
### R3 — optimize + feature-gap still instruct agents to write reports the harness blocks
|
||||
**Where:** `commands/optimize.md:121-128`, `commands/feature-gap.md:123`;
|
||||
agents `optimization-lens-agent.md` §Output, `feature-gap-agent.md` §Output.
|
||||
**Measured:** both templates tell the agent to write `*-report.md` to the
|
||||
session dir, then Read that file. The harness note measured on `analyze`
|
||||
(M-BUG-18) blocks agent writes of exactly the report/findings file class;
|
||||
`analyze` was converted to orchestrator-writes, these two arms were left open
|
||||
(tracked in STATE as the open M-BUG-18 class — this rating is its severity call).
|
||||
**What breaks:** the Read step fails or the model improvises a rescue; the
|
||||
command's documented artifact (`optimization-lens-report.md`) may never exist.
|
||||
User-visible flow breakage, no data corruption.
|
||||
**Guard shape:** the `analyze` pattern, already proven: agent returns inline,
|
||||
command persists; `analyze-report-persistence.test.mjs` is the template to copy.
|
||||
|
||||
### R4 — verifier-agent contradicts itself, and its "Read-Only Guarantee" is unenforced
|
||||
**Where:** `agents/verifier-agent.md` — §Output Format says "Append to:
|
||||
implementation-log.md"; §Read-Only Guarantee says "only uses Read, Glob, Grep /
|
||||
never modifies any files"; frontmatter `tools:` lists no write tool.
|
||||
**Measured:** `implement.md` Step 5 was already fixed to return-inline and
|
||||
append orchestrator-side — so the agent's system prompt and the spawn prompt now
|
||||
give OPPOSITE instructions to the same agent. Harness enforcement of `tools:`
|
||||
is measured absent (memory: verifier/implement-log writes went through live).
|
||||
**What breaks:** which instruction wins is nondeterministic; if the agent
|
||||
improvises a write to satisfy its own §Output Format, a full-file Write on the
|
||||
shared log clobbers parallel implementer entries — the precise defect
|
||||
`implement-log-append.test.mjs` exists to prevent, entering through the file
|
||||
that test does not read.
|
||||
**Guard shape:** rewrite verifier-agent's Output section to return-inline (one
|
||||
file), and extend `implement-log-append`/`agent-prompt-shape` to assert no agent
|
||||
prompt instructs appending to the shared log. Cheap.
|
||||
|
||||
### R5 — session state (`state.yaml`, `scope.yaml`) is a model-written machine contract with no schema anywhere
|
||||
**Where:** every phase template ("Write scope.yaml and state.yaml", "append —
|
||||
never replace — completed_phases"), `.claude/rules/state-management.md`,
|
||||
readers in `hooks/scripts/session-start.mjs` + `stop-session-reminder.mjs` +
|
||||
every session-aware command.
|
||||
**Measured:** the phase vocabulary (`discover`…`verify`) appears as a shared
|
||||
constant in **zero** code files — it lives only in prose copies (status.md's
|
||||
table, state-management.md, each template). Hooks parse with a line-grep
|
||||
(`parseYamlValue`) and print whatever they find. The existing guard
|
||||
(`command-shell-state-shape`: "phase commands name all four fields") checks the
|
||||
template *text*, not the written *file*.
|
||||
**What breaks:** resume-by-recency picks wrong sessions, status misnarrates,
|
||||
session-start reminders go quiet — degradation, not corruption, but it erodes
|
||||
exactly the "can resume if interrupted" promise the rule exists for.
|
||||
**Guard shape:** either a state-write CLI (heavy) or a defensive reader: a lib
|
||||
that validates phase tokens + required fields and *flags* malformed state, used
|
||||
by hooks and dogfooded in a test. The plugin flags drift in everyone else's
|
||||
config; its own session state deserves the same reader.
|
||||
|
||||
### R6 — both "Required Frontmatter" contracts are unguarded (currently compliant)
|
||||
**Where:** `.claude/rules/agent-development.md`, `.claude/rules/command-development.md`.
|
||||
**Measured:** no test outside fixtures matches `allowed-tools`;
|
||||
`agent-prompt-shape` asserts only `name:` on **3 of 7** agents. Measured today:
|
||||
7/7 agent frontmatters match the CLAUDE.md table; duplicate colors: **0**. So —
|
||||
compliant, unwatched. Every MUST in those two rules is enforced by nothing.
|
||||
**What breaks:** a new agent/command ships with missing `allowed-tools` or a
|
||||
duplicate color; nothing fails; the rules files become fiction one file at a
|
||||
time (the exemption-table lesson from Q1: what nothing declares, everyone
|
||||
re-answers by reading).
|
||||
**Guard shape:** near-free shape test walking `agents/*.md` + `commands/*.md`
|
||||
asserting the two rules' required keys, name conventions, color uniqueness.
|
||||
Note the irony budget: `plugin-health-scanner` already audits *other* plugins'
|
||||
frontmatter — pointing it at its own repo in a test is the dogfood version.
|
||||
|
||||
### R7 — secret detection exists only as agent prose, in a domain STATE has parked elsewhere
|
||||
**Where:** `agents/scanner-agent.md` §Secret Detection Patterns (xoxb/sk-/ghp_
|
||||
regexes); `agents/verifier-agent.md` Check 7 ("Secrets Scan ✓").
|
||||
**Measured:** `xoxb`/`ghp_` appear in **zero** files under `scanners/`;
|
||||
`mcp-config-validator.mjs` contains the string "secret" **zero** times. The
|
||||
deterministic pipeline has no secret scanning at all; the agent path claims it
|
||||
in prose, and the verifier's report template renders "Secrets Scan ✓ Pass" as a
|
||||
table row regardless. STATE parks secrets as the `llm-security` plugin's domain.
|
||||
**What breaks:** a user reads "Secrets Scan ✓" as an executed check. The lie is
|
||||
in the reporting, not in a missing feature — the feature is deliberately owned
|
||||
elsewhere.
|
||||
**Guard shape:** this is a *removal* candidate, not a gate (measurement can
|
||||
decline the feature): strip the prose secret patterns + the verifier's Check 7,
|
||||
say "secrets: out of scope, see llm-security" where the row used to be. If the
|
||||
capability is ever wanted deterministically, it starts life as a scanner with a
|
||||
finding code, not as agent prose.
|
||||
|
||||
### R8 — knowledge tables duplicated between prose and code
|
||||
**Where/measured:** managed-path table — **three** copies (scanner-agent prose +
|
||||
`file-discovery.mjs` + `active-config-reader.mjs`). Precedence — analyzer prose
|
||||
("global beats managed (user preference)") vs `conflict-detector.mjs:138`
|
||||
("local > project > user"): different vocabularies for the same claim, no link.
|
||||
Optimization-lens agent's mechanism table restates register entries
|
||||
(BP-MECH-001/002/004, BP-SUB-001) that live as data in
|
||||
`knowledge/best-practices.json`.
|
||||
**What breaks:** slow divergence — an agent narrates precedence or hierarchy the
|
||||
deterministic scanners no longer implement. Confusing, not corrupting.
|
||||
**Guard shape:** two-copies rule applies but *measure first* (#67): the two code
|
||||
copies may legitimately differ; the prose copies should cite the code as owner
|
||||
("hierarchy per `file-discovery.mjs`") rather than restate values.
|
||||
|
||||
### R9 — knowledge-refresh applies approved writes by model edit, validated only afterwards
|
||||
**Where:** `commands/knowledge-refresh.md` Step 6 (model `Edit` of
|
||||
`best-practices.json`, then run the schema test and revert on failure).
|
||||
**Measured:** already tracked open in STATE ("knowledge-refresh skrive-CLI");
|
||||
path anchoring is guarded (`knowledge-refresh-write-target.test.mjs`), the write
|
||||
itself is not. The post-hoc validation step is real mitigation — this ranks
|
||||
last *because* the failure is loud (a failing schema test in the same flow).
|
||||
**Guard shape:** the campaign pattern, which this command's own sibling already
|
||||
implements: every mutation a subcommand of a write-CLI. `campaign.md` is the
|
||||
in-repo proof that "human-approved writes" and "CLI-executed writes" compose.
|
||||
|
||||
## 4. What this changes in the plan
|
||||
|
||||
- **Q3 (severity axis)** should carry R1/R2 as its first two rows — recovery-path
|
||||
items never live in a backlog paragraph (plan §2 property 2).
|
||||
- **Q4 (v6.0.0 release)** is NOT blocked by this list (the release gate is Q1 +
|
||||
green suite); but R1+R2 are the strongest candidates for the first post-v6
|
||||
chunk, as one chunk: a rollback/backup CLI closes both, and converts R2's
|
||||
guard from "pin the prose format" to "delete the prose format".
|
||||
- **R4 + R6** are lunch-sized; they can ride along with any adjacent session the
|
||||
way M-BUG fixes have.
|
||||
- **R7** is an operator decision (removing a claimed capability): propose, don't do.
|
||||
|
||||
## Appendix — measurement log
|
||||
|
||||
Every number above, and the command that produced it (run at `30c78ae`):
|
||||
|
||||
| # | Number | Command |
|
||||
|---|--------|---------|
|
||||
| 1 | 21 / 7 / 4 files | `ls commands/*.md \| wc -l` etc. |
|
||||
| 2 | 4,950 lines | `wc -l commands/*.md agents/*.md .claude/rules/*.md` |
|
||||
| 3 | 17 guard files | `find tests/commands tests/agents -name '*.test.mjs' \| wc -l` |
|
||||
| 4 | 18 markers | `grep -cE '\b(MUST\|NEVER\|ALWAYS\|DO NOT\|…)\b' commands/*.md agents/*.md` |
|
||||
| 5 | 16 CLIs, rollback-engine absent | `grep -ln "process.argv" scanners/*.mjs` |
|
||||
| 6 | parseManifest dual-format + its cause | `sed -n '140,220p' scanners/lib/backup.mjs` (comment at :172) |
|
||||
| 7 | hand-written fixture | `grep -n "sha256" tests/scanners/rollback-paths.test.mjs` (:159-186) |
|
||||
| 8 | fix uses code backup | `grep -n "backup" scanners/fix-engine.mjs` (:10) |
|
||||
| 9 | 0 secret patterns in code | `grep -rln "xoxb\|ghp_" scanners/` (empty); `grep -n "secret" scanners/mcp-config-validator.mjs` (empty) |
|
||||
| 10 | 3 copies managed-path table | `grep -rln "Library/Application Support" scanners/` (2) + scanner-agent prose (1) |
|
||||
| 11 | 0 phase-vocabulary constants | `grep -rn "'discover'" scanners/lib/*.mjs hooks/scripts/*.mjs` (empty) |
|
||||
| 12 | 3/7 agents in shape guard | `AGENT_FILES` array read in `tests/agents/agent-prompt-shape.test.mjs` |
|
||||
| 13 | 0 frontmatter guards | `grep -rln "allowed-tools" tests/` (fixtures + yaml-parser only) |
|
||||
| 14 | 0 duplicate colors | `grep -h "^color:" agents/*.md \| sort \| uniq -d \| wc -l` |
|
||||
| 15 | verifier self-contradiction | Read `agents/verifier-agent.md` (:143 append vs :242 read-only) |
|
||||
|
||||
Not covered by this sweep (deliberate): `knowledge/*.md` content freshness
|
||||
(knowledge-refresh's domain), README/plugin.json surface (repo-standard's
|
||||
domain), the deterministic scanners themselves (Q1/Q2 territory, already coded).
|
||||
Loading…
Add table
Add a link
Reference in a new issue