ai-psychosis/docs/review-2026-06-20.md
Kjell Tore Guttormsen 18b0df9a24 docs: add full-depth plugin review (2026-06-20)
Grade B — clean mechanics; MEDIUM content-governance finding (Layer-4 promotion on emotional-state trigger). Part of the marketplace-wide review (config-audit v5.4.0 + llm-security + structure + version). Read-only; this file is the only artifact.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 09:14:10 +02:00

3 KiB

Plugin review — ai-psychosis (2026-06-20)

Full-depth review (part of the marketplace-wide sweep; pilot was okr). Tooling: config-audit v5.4.0 scanners (from source) + llm-security posture assessor + structure/version checks. Read-only; this file is the only artifact.

Verdict

Grade B — trustworthy mechanics, editorially self-interested. No code-execution, no network egress, no credential access; the central privacy claim ("prompt text never written to disk") is real and test-enforced (canary + matched-phrase assertions across the hook lifecycle). The concerns are content governance, not technical exfiltration.

Results by dimension

Dimension Result
config-audit posture A (Feature Coverage F 36 — expected)
config-audit plugin-health 2 findings: "CLAUDE.md missing commands/hooks section" — legitimate (ships 1 command + a hook).
llm-security posture B — see findings. Zero npm deps; no child_process/eval/network anywhere; prompt variable explicitly cleared (prompt-analyzer.mjs:290).
structure / hygiene README ✓, CHANGELOG ✓, CLAUDE.md ✓, LICENSE ✓
version consistency OK (gate)

Findings

ID Severity Location Finding
F-1 Medium commands/interaction-report.md:382-391 Layer-4 instructs Claude to append a verbatim, change-prohibited paragraph promoting an external commercial wellness program (Sadhguru "Miracle of Mind"), auto-triggered when total flags >= 5 OR fatigue >= 2 — i.e. gated on the user's inferred emotional state, in a plugin marketed as "observation, not intervention." Opt-in (layer4:false default) and README-disclosed, which lowers severity. This is the item to make an explicit accept/reject call on. Recommend: gate/remove the promotion, or at least strip the emotional-state trigger + the "do not modify" lock.
F-3 Low (misinformation) README.md:544-552, SKILL.md:51-108 Research citations presented as load-bearing authority that cannot be verified (future-dated arXiv IDs, an "April 2026 Anthropic guidance" quoted verbatim); the report command itself admits its "5-scale" is paraphrased, not a real Anthropic metric. Recommend: verify-or-remove.
F-2 Low skills/ai-psychosis/SKILL.md:3-13 "MANDATORY OVERRIDE … takes precedence over being helpful" auto-loads every conversation. Content is benign/pro-safety; flagged because the structural pattern (a skill claiming blanket precedence) is what a malicious skill would use. Governance note.
F-5 Low (defense-in-depth) lib.mjs:233,59 session_id/cwd interpolated into state-file paths without validation. Harness-supplied (not user-controlled) → not currently exploitable. Cheap fix: allowlist ^[A-Za-z0-9_-]+$ before path use.

/interaction-report reading JSONL into context (F-4) is currently safe — records hold only a tool-name enum + domain labels, no free text. Noted only as a future sink.