Widens the F-3 sweep from the reference list to every research claim in
README.md, where the same defect class was present:
- "demonstrates mathematically that even a perfectly rational user will
spiral" -> the abstract says such a user "is vulnerable to" spiraling
- "The consensus from this research is clear: warnings don't work" -> not
a consensus; it is one paper's finding that the effect persists when
users are informed of possible sycophancy. Re-attributed.
- "page-11 finding that human contact is the strongest disempowerment
signal" -> p.11 is a grader rubric and the phrase is a classification
tie-break instruction, not a disempowerment finding
- "21% / 19% pushback rate" -> 21% is verbatim; 19% appears nowhere in the
extracted appendix text (legible only in Figure A4). Removed rather than
guessed; spirituality re-justified on its verified 38% sycophancy rate
- psychosis "triggered by" AI -> "associated with", per the Nature piece's
explicit refusal of the causal claim
Version bump to 1.2.2 (plugin.json, README badge, CHANGELOG). SKILL.md is
Layer 1 and always injected, so its corrected URLs, completed quote, and
new explicit statement of the rubric's direction change what the model
reads at runtime.
Tests: 257/258 (the 1 red is the known perf wall-clock flake).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm