docs: add full-depth plugin review (2026-06-20)
Grade B — clean mechanics; MEDIUM content-governance finding (Layer-4 promotion on emotional-state trigger). Part of the marketplace-wide review (config-audit v5.4.0 + llm-security + structure + version). Read-only; this file is the only artifact. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
This commit is contained in:
parent
9f5d8f9ce4
commit
18b0df9a24
1 changed files with 34 additions and 0 deletions
34
docs/review-2026-06-20.md
Normal file
34
docs/review-2026-06-20.md
Normal file
|
|
@ -0,0 +1,34 @@
|
|||
# Plugin review — ai-psychosis (2026-06-20)
|
||||
|
||||
> Full-depth review (part of the marketplace-wide sweep; pilot was okr). Tooling: config-audit
|
||||
> v5.4.0 scanners (from source) + llm-security posture assessor + structure/version checks.
|
||||
> Read-only; this file is the only artifact.
|
||||
|
||||
## Verdict
|
||||
|
||||
**Grade B — trustworthy mechanics, editorially self-interested.** No code-execution, no network
|
||||
egress, no credential access; the central privacy claim ("prompt text never written to disk") is
|
||||
real and test-enforced (canary + matched-phrase assertions across the hook lifecycle). The concerns
|
||||
are **content governance**, not technical exfiltration.
|
||||
|
||||
## Results by dimension
|
||||
|
||||
| Dimension | Result |
|
||||
|-----------|--------|
|
||||
| config-audit posture | **A** (Feature Coverage F 36 — expected) |
|
||||
| config-audit plugin-health | 2 findings: "CLAUDE.md missing commands/hooks section" — **legitimate** (ships 1 command + a hook). |
|
||||
| llm-security posture | **B** — see findings. Zero npm deps; no `child_process`/`eval`/network anywhere; prompt variable explicitly cleared (`prompt-analyzer.mjs:290`). |
|
||||
| structure / hygiene | README ✓, CHANGELOG ✓, CLAUDE.md ✓, LICENSE ✓ |
|
||||
| version consistency | **OK** (gate) |
|
||||
|
||||
## Findings
|
||||
|
||||
| ID | Severity | Location | Finding |
|
||||
|----|----------|----------|---------|
|
||||
| F-1 | **Medium** | `commands/interaction-report.md:382-391` | Layer-4 instructs Claude to append a verbatim, change-prohibited paragraph promoting an external commercial wellness program (Sadhguru "Miracle of Mind"), auto-triggered when `total flags >= 5 OR fatigue >= 2` — i.e. gated on the user's inferred emotional state, in a plugin marketed as "observation, not intervention." Opt-in (`layer4:false` default) and README-disclosed, which lowers severity. **This is the item to make an explicit accept/reject call on.** Recommend: gate/remove the promotion, or at least strip the emotional-state trigger + the "do not modify" lock. |
|
||||
| F-3 | Low (misinformation) | `README.md:544-552`, `SKILL.md:51-108` | Research citations presented as load-bearing authority that cannot be verified (future-dated arXiv IDs, an "April 2026 Anthropic guidance" quoted verbatim); the report command itself admits its "5-scale" is paraphrased, not a real Anthropic metric. **Recommend:** verify-or-remove. |
|
||||
| F-2 | Low | `skills/ai-psychosis/SKILL.md:3-13` | "MANDATORY OVERRIDE … takes precedence over being helpful" auto-loads every conversation. Content is benign/pro-safety; flagged because the *structural pattern* (a skill claiming blanket precedence) is what a malicious skill would use. Governance note. |
|
||||
| F-5 | Low (defense-in-depth) | `lib.mjs:233,59` | `session_id`/`cwd` interpolated into state-file paths without validation. Harness-supplied (not user-controlled) → not currently exploitable. Cheap fix: allowlist `^[A-Za-z0-9_-]+$` before path use. |
|
||||
|
||||
`/interaction-report` reading JSONL into context (F-4) is currently safe — records hold only a
|
||||
tool-name enum + domain labels, no free text. Noted only as a future sink.
|
||||
Loading…
Add table
Add a link
Reference in a new issue