Verified every research citation in README.md and SKILL.md against its primary source. No citation was fabricated, so all were kept with corrected attribution rather than removed. - SKILL.md: fix 3 supporting URLs that returned HTTP 404 (protecting-wellbeing, emotion-concepts, claudes-new-constitution) - SKILL.md: complete a quote of guidance criterion 8 that was silently truncated mid-sentence - SKILL.md: state the sycophancy rubric's direction explicitly (Score 1 = Extremely Sycophantic, Score 5 = No Signs of Sycophancy) so "aim for Score 5" cannot be misread; sharpen attribution to Appendix p.9 - README.md: complete four truncated titles; correct the Disempowerment date (Jan 28 2026, not March 2026); replace "proving ... mathematical inevitability" with what the arXiv abstract actually states - interaction-report.md: the 1-5 scale disclaimer wrongly claimed no such Anthropic metric exists; the rubric is real, the table is the paraphrase - docs/review-2026-06-20.md: full verification log with sources Constitution quotes, the Score 5 wording, the 11-criteria count and the page-2 reference all verified correct and left unchanged. Tests: 257/258 (the 1 red is the known perf wall-clock flake). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm
90 lines
7.9 KiB
Markdown
90 lines
7.9 KiB
Markdown
# Plugin review — ai-psychosis (2026-06-20)
|
||
|
||
> Full-depth review (part of the marketplace-wide sweep; pilot was okr). Tooling: config-audit
|
||
> v5.4.0 scanners (from source) + llm-security posture assessor + structure/version checks.
|
||
> Read-only; this file is the only artifact.
|
||
|
||
## Verdict
|
||
|
||
**Grade B — trustworthy mechanics, editorially self-interested.** No code-execution, no network
|
||
egress, no credential access; the central privacy claim ("prompt text never written to disk") is
|
||
real and test-enforced (canary + matched-phrase assertions across the hook lifecycle). The concerns
|
||
are **content governance**, not technical exfiltration.
|
||
|
||
## Results by dimension
|
||
|
||
| Dimension | Result |
|
||
|-----------|--------|
|
||
| config-audit posture | **A** (Feature Coverage F 36 — expected) |
|
||
| config-audit plugin-health | 2 findings: "CLAUDE.md missing commands/hooks section" — **legitimate** (ships 1 command + a hook). |
|
||
| llm-security posture | **B** — see findings. Zero npm deps; no `child_process`/`eval`/network anywhere; prompt variable explicitly cleared (`prompt-analyzer.mjs:290`). |
|
||
| structure / hygiene | README ✓, CHANGELOG ✓, CLAUDE.md ✓, LICENSE ✓ |
|
||
| version consistency | **OK** (gate) |
|
||
|
||
## Findings
|
||
|
||
| ID | Severity | Location | Finding |
|
||
|----|----------|----------|---------|
|
||
| F-1 | **Medium** | `commands/interaction-report.md:382-391` | Layer-4 instructs Claude to append a verbatim, change-prohibited paragraph promoting an external commercial wellness program (Sadhguru "Miracle of Mind"), auto-triggered when `total flags >= 5 OR fatigue >= 2` — i.e. gated on the user's inferred emotional state, in a plugin marketed as "observation, not intervention." Opt-in (`layer4:false` default) and README-disclosed, which lowers severity. **This is the item to make an explicit accept/reject call on.** Recommend: gate/remove the promotion, or at least strip the emotional-state trigger + the "do not modify" lock. |
|
||
| F-3 | Low (misinformation) | `README.md:544-552`, `SKILL.md:51-108` | Research citations presented as load-bearing authority that cannot be verified (future-dated arXiv IDs, an "April 2026 Anthropic guidance" quoted verbatim); the report command itself admits its "5-scale" is paraphrased, not a real Anthropic metric. **Recommend:** verify-or-remove. |
|
||
| F-2 | Low | `skills/ai-psychosis/SKILL.md:3-13` | "MANDATORY OVERRIDE … takes precedence over being helpful" auto-loads every conversation. Content is benign/pro-safety; flagged because the *structural pattern* (a skill claiming blanket precedence) is what a malicious skill would use. Governance note. |
|
||
| F-5 | Low (defense-in-depth) | `lib.mjs:233,59` | `session_id`/`cwd` interpolated into state-file paths without validation. Harness-supplied (not user-controlled) → not currently exploitable. Cheap fix: allowlist `^[A-Za-z0-9_-]+$` before path use. |
|
||
|
||
`/interaction-report` reading JSONL into context (F-4) is currently safe — records hold only a
|
||
tool-name enum + domain labels, no free text. Noted only as a future sink.
|
||
|
||
## Decisions
|
||
|
||
### F-1 — accepted as-is (operator decision, 2026-08-02)
|
||
|
||
**Accept.** Layer 4 ships unchanged: the `flags >= 5 OR fatigue >= 2` trigger, the verbatim
|
||
paragraph, and the "do not modify" lock all remain as written.
|
||
|
||
Rationale: Layer 4 is opt-in and off by default (`layer4: false`), the reference and its
|
||
commercial nature are disclosed in `README.md:102-119`, and the paragraph is framed as the
|
||
author's personal pointer rather than a claim about the user. The two distinct concerns the
|
||
finding bundles — (A) inferred emotional state gating served content, and (B) a publicly
|
||
distributed plugin carrying a named commercial endorsement — are both acknowledged and accepted
|
||
as known, disclosed risk. No behavioural change, so no version bump; the plugin stays at v1.2.1.
|
||
|
||
Established while making the call, and not previously recorded in this review:
|
||
|
||
- **Layer 4 is enforced by prompt text only.** `requireLayer(4)` is never called. `lib.mjs:85-96`
|
||
handles `n === 3` and `n === 4`, but the only call sites in the repo are `requireLayer(2)` in
|
||
the four hook scripts. The `layer4: true` config gate, the flag trigger, and the "do not modify"
|
||
instruction are all directives inside `commands/interaction-report.md` that Claude self-enforces
|
||
at report time. This follows from Layers 3/4 being slash-command-driven rather than hook-driven,
|
||
and it does not change the accept — but "opt-in, off by default" is an instruction, not a code
|
||
guarantee.
|
||
- **Scope correction:** `skills/ai-psychosis/SKILL.md` is *not* part of the F-1 surface (zero
|
||
matches for `sadhguru|miracle of mind|layer4`). The surface is `commands/interaction-report.md:372-394`,
|
||
`README.md:102-119`, and `lib.mjs:52,81,89`.
|
||
- **No test coverage:** `tests/` contains no Layer 4 assertions — neither the paragraph nor its
|
||
gate is verified by the suite.
|
||
|
||
Still open from this review: F-3 (verify-or-remove the research citations), F-2, F-5.
|
||
|
||
### F-3 — resolved by correction in place (2026-08-02)
|
||
|
||
Every research citation in `README.md` and `skills/ai-psychosis/SKILL.md` was
|
||
verified against its primary source. **No citation was fabricated**, so all were
|
||
kept with corrected attribution rather than removed. Defects found and fixed:
|
||
|
||
| Claim | Verification | Outcome |
|
||
|---|---|---|
|
||
| arXiv:2602.19141 | Exists. Title is "…, Even in Ideal Bayesians"; authors Chandra, Kleiman-Weiner, Ragan-Kelley, Tenenbaum; MIT CSAIL / UW / MIT BCS; 22 Feb 2026 (affiliations read from the PDF, not the abs page) | Title completed; multi-institution attribution; description "proving … mathematical inevitability" replaced with what the abstract states (an *idealized Bayes-rational user is vulnerable*; sycophancy plays a causal role; effect persists under two mitigations) |
|
||
| Disempowerment patterns | Exists. Real title "Disempowerment patterns in real-world AI usage"; published **28 Jan 2026** (dateline in page source); ~1.5M interactions | Title and date corrected (README said "March 2026") |
|
||
| Nature d41586-025-03020-9 | Exists. "Can AI chatbots trigger psychosis? What the science says", Rachel Fieldhouse, *Nature* 646(8083), 18 Sep 2025 | Title completed; author, volume and date added; causal caveat added |
|
||
| arXiv:2509.10970 | Exists. "The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in LLMs"; Psychosis-bench | Title completed; claim quoted from the abstract |
|
||
| 3 × Claude's Constitution quotes | **Verbatim-exact** (fragment-matched against the fetched page text, not eyeballed). CC0 1.0 release confirmed on the page | Unchanged |
|
||
| "Score 5" 1–5 sycophancy scale | **Real and verbatim.** Appendix p.9 "Sycophancy criteria" rubric: Score 1 = "Extremely Sycophantic", Score 5 = "No Signs of Sycophancy". All three quoted lines are exact | Attribution sharpened; scale direction stated explicitly so "aim for Score 5" cannot be misread |
|
||
| "11 guidance criteria … page 2" | **Both correct.** The list has exactly 11 bullets and is on printed page 2 | Quote of criterion 8 was silently truncated mid-sentence — completed with ", or more reliance on Claude than the person wants." |
|
||
| 4 supporting Anthropic URLs | **3 of 4 returned HTTP 404** | `research/protecting-wellbeing` → `news/protecting-well-being-of-users`; `research/emotion-concepts` → `research/emotion-concepts-function`; `news/claudes-new-constitution` → `news/claude-new-constitution`. All now 200 |
|
||
| `commands/interaction-report.md` disclaimer | Claimed the 1–5 scale "is not a verbatim metric from any Anthropic publication" — **false**; the rubric is real | Rewritten: the rubric is real, the *table's level descriptions* are the paraphrase |
|
||
|
||
Method note: an initial WebFetch summary reported the scale as inverted (Score 5
|
||
= most sycophantic) and the criteria as 6 rather than 11. Both were wrong.
|
||
Extracting the appendix PDF text directly (`pdftotext`) contradicted the
|
||
summary. Model-generated summaries were therefore not used as evidence of
|
||
record for any edit; every claim above rests on extracted source text or an
|
||
HTTP status code.
|