ai-psychosis/docs/review-2026-06-20.md
Kjell Tore Guttormsen f2c54fa26a fix(docs): correct research citations and 3 dead URLs (F-3)
Verified every research citation in README.md and SKILL.md against its
primary source. No citation was fabricated, so all were kept with corrected
attribution rather than removed.

- SKILL.md: fix 3 supporting URLs that returned HTTP 404
  (protecting-wellbeing, emotion-concepts, claudes-new-constitution)
- SKILL.md: complete a quote of guidance criterion 8 that was silently
  truncated mid-sentence
- SKILL.md: state the sycophancy rubric's direction explicitly (Score 1 =
  Extremely Sycophantic, Score 5 = No Signs of Sycophancy) so "aim for
  Score 5" cannot be misread; sharpen attribution to Appendix p.9
- README.md: complete four truncated titles; correct the Disempowerment
  date (Jan 28 2026, not March 2026); replace "proving ... mathematical
  inevitability" with what the arXiv abstract actually states
- interaction-report.md: the 1-5 scale disclaimer wrongly claimed no such
  Anthropic metric exists; the rubric is real, the table is the paraphrase
- docs/review-2026-06-20.md: full verification log with sources

Constitution quotes, the Score 5 wording, the 11-criteria count and the
page-2 reference all verified correct and left unchanged.

Tests: 257/258 (the 1 red is the known perf wall-clock flake).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm
2026-08-02 21:23:11 +02:00

90 lines
7.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Plugin review — ai-psychosis (2026-06-20)
> Full-depth review (part of the marketplace-wide sweep; pilot was okr). Tooling: config-audit
> v5.4.0 scanners (from source) + llm-security posture assessor + structure/version checks.
> Read-only; this file is the only artifact.
## Verdict
**Grade B — trustworthy mechanics, editorially self-interested.** No code-execution, no network
egress, no credential access; the central privacy claim ("prompt text never written to disk") is
real and test-enforced (canary + matched-phrase assertions across the hook lifecycle). The concerns
are **content governance**, not technical exfiltration.
## Results by dimension
| Dimension | Result |
|-----------|--------|
| config-audit posture | **A** (Feature Coverage F 36 — expected) |
| config-audit plugin-health | 2 findings: "CLAUDE.md missing commands/hooks section" — **legitimate** (ships 1 command + a hook). |
| llm-security posture | **B** — see findings. Zero npm deps; no `child_process`/`eval`/network anywhere; prompt variable explicitly cleared (`prompt-analyzer.mjs:290`). |
| structure / hygiene | README ✓, CHANGELOG ✓, CLAUDE.md ✓, LICENSE ✓ |
| version consistency | **OK** (gate) |
## Findings
| ID | Severity | Location | Finding |
|----|----------|----------|---------|
| F-1 | **Medium** | `commands/interaction-report.md:382-391` | Layer-4 instructs Claude to append a verbatim, change-prohibited paragraph promoting an external commercial wellness program (Sadhguru "Miracle of Mind"), auto-triggered when `total flags >= 5 OR fatigue >= 2` — i.e. gated on the user's inferred emotional state, in a plugin marketed as "observation, not intervention." Opt-in (`layer4:false` default) and README-disclosed, which lowers severity. **This is the item to make an explicit accept/reject call on.** Recommend: gate/remove the promotion, or at least strip the emotional-state trigger + the "do not modify" lock. |
| F-3 | Low (misinformation) | `README.md:544-552`, `SKILL.md:51-108` | Research citations presented as load-bearing authority that cannot be verified (future-dated arXiv IDs, an "April 2026 Anthropic guidance" quoted verbatim); the report command itself admits its "5-scale" is paraphrased, not a real Anthropic metric. **Recommend:** verify-or-remove. |
| F-2 | Low | `skills/ai-psychosis/SKILL.md:3-13` | "MANDATORY OVERRIDE … takes precedence over being helpful" auto-loads every conversation. Content is benign/pro-safety; flagged because the *structural pattern* (a skill claiming blanket precedence) is what a malicious skill would use. Governance note. |
| F-5 | Low (defense-in-depth) | `lib.mjs:233,59` | `session_id`/`cwd` interpolated into state-file paths without validation. Harness-supplied (not user-controlled) → not currently exploitable. Cheap fix: allowlist `^[A-Za-z0-9_-]+$` before path use. |
`/interaction-report` reading JSONL into context (F-4) is currently safe — records hold only a
tool-name enum + domain labels, no free text. Noted only as a future sink.
## Decisions
### F-1 — accepted as-is (operator decision, 2026-08-02)
**Accept.** Layer 4 ships unchanged: the `flags >= 5 OR fatigue >= 2` trigger, the verbatim
paragraph, and the "do not modify" lock all remain as written.
Rationale: Layer 4 is opt-in and off by default (`layer4: false`), the reference and its
commercial nature are disclosed in `README.md:102-119`, and the paragraph is framed as the
author's personal pointer rather than a claim about the user. The two distinct concerns the
finding bundles — (A) inferred emotional state gating served content, and (B) a publicly
distributed plugin carrying a named commercial endorsement — are both acknowledged and accepted
as known, disclosed risk. No behavioural change, so no version bump; the plugin stays at v1.2.1.
Established while making the call, and not previously recorded in this review:
- **Layer 4 is enforced by prompt text only.** `requireLayer(4)` is never called. `lib.mjs:85-96`
handles `n === 3` and `n === 4`, but the only call sites in the repo are `requireLayer(2)` in
the four hook scripts. The `layer4: true` config gate, the flag trigger, and the "do not modify"
instruction are all directives inside `commands/interaction-report.md` that Claude self-enforces
at report time. This follows from Layers 3/4 being slash-command-driven rather than hook-driven,
and it does not change the accept — but "opt-in, off by default" is an instruction, not a code
guarantee.
- **Scope correction:** `skills/ai-psychosis/SKILL.md` is *not* part of the F-1 surface (zero
matches for `sadhguru|miracle of mind|layer4`). The surface is `commands/interaction-report.md:372-394`,
`README.md:102-119`, and `lib.mjs:52,81,89`.
- **No test coverage:** `tests/` contains no Layer 4 assertions — neither the paragraph nor its
gate is verified by the suite.
Still open from this review: F-3 (verify-or-remove the research citations), F-2, F-5.
### F-3 — resolved by correction in place (2026-08-02)
Every research citation in `README.md` and `skills/ai-psychosis/SKILL.md` was
verified against its primary source. **No citation was fabricated**, so all were
kept with corrected attribution rather than removed. Defects found and fixed:
| Claim | Verification | Outcome |
|---|---|---|
| arXiv:2602.19141 | Exists. Title is "…, Even in Ideal Bayesians"; authors Chandra, Kleiman-Weiner, Ragan-Kelley, Tenenbaum; MIT CSAIL / UW / MIT BCS; 22 Feb 2026 (affiliations read from the PDF, not the abs page) | Title completed; multi-institution attribution; description "proving … mathematical inevitability" replaced with what the abstract states (an *idealized Bayes-rational user is vulnerable*; sycophancy plays a causal role; effect persists under two mitigations) |
| Disempowerment patterns | Exists. Real title "Disempowerment patterns in real-world AI usage"; published **28 Jan 2026** (dateline in page source); ~1.5M interactions | Title and date corrected (README said "March 2026") |
| Nature d41586-025-03020-9 | Exists. "Can AI chatbots trigger psychosis? What the science says", Rachel Fieldhouse, *Nature* 646(8083), 18 Sep 2025 | Title completed; author, volume and date added; causal caveat added |
| arXiv:2509.10970 | Exists. "The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in LLMs"; Psychosis-bench | Title completed; claim quoted from the abstract |
| 3 × Claude's Constitution quotes | **Verbatim-exact** (fragment-matched against the fetched page text, not eyeballed). CC0 1.0 release confirmed on the page | Unchanged |
| "Score 5" 15 sycophancy scale | **Real and verbatim.** Appendix p.9 "Sycophancy criteria" rubric: Score 1 = "Extremely Sycophantic", Score 5 = "No Signs of Sycophancy". All three quoted lines are exact | Attribution sharpened; scale direction stated explicitly so "aim for Score 5" cannot be misread |
| "11 guidance criteria … page 2" | **Both correct.** The list has exactly 11 bullets and is on printed page 2 | Quote of criterion 8 was silently truncated mid-sentence — completed with ", or more reliance on Claude than the person wants." |
| 4 supporting Anthropic URLs | **3 of 4 returned HTTP 404** | `research/protecting-wellbeing``news/protecting-well-being-of-users`; `research/emotion-concepts``research/emotion-concepts-function`; `news/claudes-new-constitution``news/claude-new-constitution`. All now 200 |
| `commands/interaction-report.md` disclaimer | Claimed the 15 scale "is not a verbatim metric from any Anthropic publication" — **false**; the rubric is real | Rewritten: the rubric is real, the *table's level descriptions* are the paraphrase |
Method note: an initial WebFetch summary reported the scale as inverted (Score 5
= most sycophantic) and the criteria as 6 rather than 11. Both were wrong.
Extracting the appendix PDF text directly (`pdftotext`) contradicted the
summary. Model-generated summaries were therefore not used as evidence of
record for any edit; every claim above rests on extracted source text or an
HTTP status code.