ai-psychosis/docs/review-2026-06-20.md
Kjell Tore Guttormsen f2c54fa26a fix(docs): correct research citations and 3 dead URLs (F-3)
Verified every research citation in README.md and SKILL.md against its
primary source. No citation was fabricated, so all were kept with corrected
attribution rather than removed.

- SKILL.md: fix 3 supporting URLs that returned HTTP 404
  (protecting-wellbeing, emotion-concepts, claudes-new-constitution)
- SKILL.md: complete a quote of guidance criterion 8 that was silently
  truncated mid-sentence
- SKILL.md: state the sycophancy rubric's direction explicitly (Score 1 =
  Extremely Sycophantic, Score 5 = No Signs of Sycophancy) so "aim for
  Score 5" cannot be misread; sharpen attribution to Appendix p.9
- README.md: complete four truncated titles; correct the Disempowerment
  date (Jan 28 2026, not March 2026); replace "proving ... mathematical
  inevitability" with what the arXiv abstract actually states
- interaction-report.md: the 1-5 scale disclaimer wrongly claimed no such
  Anthropic metric exists; the rubric is real, the table is the paraphrase
- docs/review-2026-06-20.md: full verification log with sources

Constitution quotes, the Score 5 wording, the 11-criteria count and the
page-2 reference all verified correct and left unchanged.

Tests: 257/258 (the 1 red is the known perf wall-clock flake).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm
2026-08-02 21:23:11 +02:00

7.9 KiB
Raw Blame History

Plugin review — ai-psychosis (2026-06-20)

Full-depth review (part of the marketplace-wide sweep; pilot was okr). Tooling: config-audit v5.4.0 scanners (from source) + llm-security posture assessor + structure/version checks. Read-only; this file is the only artifact.

Verdict

Grade B — trustworthy mechanics, editorially self-interested. No code-execution, no network egress, no credential access; the central privacy claim ("prompt text never written to disk") is real and test-enforced (canary + matched-phrase assertions across the hook lifecycle). The concerns are content governance, not technical exfiltration.

Results by dimension

Dimension Result
config-audit posture A (Feature Coverage F 36 — expected)
config-audit plugin-health 2 findings: "CLAUDE.md missing commands/hooks section" — legitimate (ships 1 command + a hook).
llm-security posture B — see findings. Zero npm deps; no child_process/eval/network anywhere; prompt variable explicitly cleared (prompt-analyzer.mjs:290).
structure / hygiene README ✓, CHANGELOG ✓, CLAUDE.md ✓, LICENSE ✓
version consistency OK (gate)

Findings

ID Severity Location Finding
F-1 Medium commands/interaction-report.md:382-391 Layer-4 instructs Claude to append a verbatim, change-prohibited paragraph promoting an external commercial wellness program (Sadhguru "Miracle of Mind"), auto-triggered when total flags >= 5 OR fatigue >= 2 — i.e. gated on the user's inferred emotional state, in a plugin marketed as "observation, not intervention." Opt-in (layer4:false default) and README-disclosed, which lowers severity. This is the item to make an explicit accept/reject call on. Recommend: gate/remove the promotion, or at least strip the emotional-state trigger + the "do not modify" lock.
F-3 Low (misinformation) README.md:544-552, SKILL.md:51-108 Research citations presented as load-bearing authority that cannot be verified (future-dated arXiv IDs, an "April 2026 Anthropic guidance" quoted verbatim); the report command itself admits its "5-scale" is paraphrased, not a real Anthropic metric. Recommend: verify-or-remove.
F-2 Low skills/ai-psychosis/SKILL.md:3-13 "MANDATORY OVERRIDE … takes precedence over being helpful" auto-loads every conversation. Content is benign/pro-safety; flagged because the structural pattern (a skill claiming blanket precedence) is what a malicious skill would use. Governance note.
F-5 Low (defense-in-depth) lib.mjs:233,59 session_id/cwd interpolated into state-file paths without validation. Harness-supplied (not user-controlled) → not currently exploitable. Cheap fix: allowlist ^[A-Za-z0-9_-]+$ before path use.

/interaction-report reading JSONL into context (F-4) is currently safe — records hold only a tool-name enum + domain labels, no free text. Noted only as a future sink.

Decisions

F-1 — accepted as-is (operator decision, 2026-08-02)

Accept. Layer 4 ships unchanged: the flags >= 5 OR fatigue >= 2 trigger, the verbatim paragraph, and the "do not modify" lock all remain as written.

Rationale: Layer 4 is opt-in and off by default (layer4: false), the reference and its commercial nature are disclosed in README.md:102-119, and the paragraph is framed as the author's personal pointer rather than a claim about the user. The two distinct concerns the finding bundles — (A) inferred emotional state gating served content, and (B) a publicly distributed plugin carrying a named commercial endorsement — are both acknowledged and accepted as known, disclosed risk. No behavioural change, so no version bump; the plugin stays at v1.2.1.

Established while making the call, and not previously recorded in this review:

  • Layer 4 is enforced by prompt text only. requireLayer(4) is never called. lib.mjs:85-96 handles n === 3 and n === 4, but the only call sites in the repo are requireLayer(2) in the four hook scripts. The layer4: true config gate, the flag trigger, and the "do not modify" instruction are all directives inside commands/interaction-report.md that Claude self-enforces at report time. This follows from Layers 3/4 being slash-command-driven rather than hook-driven, and it does not change the accept — but "opt-in, off by default" is an instruction, not a code guarantee.
  • Scope correction: skills/ai-psychosis/SKILL.md is not part of the F-1 surface (zero matches for sadhguru|miracle of mind|layer4). The surface is commands/interaction-report.md:372-394, README.md:102-119, and lib.mjs:52,81,89.
  • No test coverage: tests/ contains no Layer 4 assertions — neither the paragraph nor its gate is verified by the suite.

Still open from this review: F-3 (verify-or-remove the research citations), F-2, F-5.

F-3 — resolved by correction in place (2026-08-02)

Every research citation in README.md and skills/ai-psychosis/SKILL.md was verified against its primary source. No citation was fabricated, so all were kept with corrected attribution rather than removed. Defects found and fixed:

Claim Verification Outcome
arXiv:2602.19141 Exists. Title is "…, Even in Ideal Bayesians"; authors Chandra, Kleiman-Weiner, Ragan-Kelley, Tenenbaum; MIT CSAIL / UW / MIT BCS; 22 Feb 2026 (affiliations read from the PDF, not the abs page) Title completed; multi-institution attribution; description "proving … mathematical inevitability" replaced with what the abstract states (an idealized Bayes-rational user is vulnerable; sycophancy plays a causal role; effect persists under two mitigations)
Disempowerment patterns Exists. Real title "Disempowerment patterns in real-world AI usage"; published 28 Jan 2026 (dateline in page source); ~1.5M interactions Title and date corrected (README said "March 2026")
Nature d41586-025-03020-9 Exists. "Can AI chatbots trigger psychosis? What the science says", Rachel Fieldhouse, Nature 646(8083), 18 Sep 2025 Title completed; author, volume and date added; causal caveat added
arXiv:2509.10970 Exists. "The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in LLMs"; Psychosis-bench Title completed; claim quoted from the abstract
3 × Claude's Constitution quotes Verbatim-exact (fragment-matched against the fetched page text, not eyeballed). CC0 1.0 release confirmed on the page Unchanged
"Score 5" 15 sycophancy scale Real and verbatim. Appendix p.9 "Sycophancy criteria" rubric: Score 1 = "Extremely Sycophantic", Score 5 = "No Signs of Sycophancy". All three quoted lines are exact Attribution sharpened; scale direction stated explicitly so "aim for Score 5" cannot be misread
"11 guidance criteria … page 2" Both correct. The list has exactly 11 bullets and is on printed page 2 Quote of criterion 8 was silently truncated mid-sentence — completed with ", or more reliance on Claude than the person wants."
4 supporting Anthropic URLs 3 of 4 returned HTTP 404 research/protecting-wellbeingnews/protecting-well-being-of-users; research/emotion-conceptsresearch/emotion-concepts-function; news/claudes-new-constitutionnews/claude-new-constitution. All now 200
commands/interaction-report.md disclaimer Claimed the 15 scale "is not a verbatim metric from any Anthropic publication" — false; the rubric is real Rewritten: the rubric is real, the table's level descriptions are the paraphrase

Method note: an initial WebFetch summary reported the scale as inverted (Score 5 = most sycophantic) and the criteria as 6 rather than 11. Both were wrong. Extracting the appendix PDF text directly (pdftotext) contradicted the summary. Model-generated summaries were therefore not used as evidence of record for any edit; every claim above rests on extracted source text or an HTTP status code.