Commit graph

5 commits

Author SHA1 Message Date
0cf9a1aa5b fix(docs): correct research claims in README prose; release 1.2.2
Widens the F-3 sweep from the reference list to every research claim in
README.md, where the same defect class was present:

- "demonstrates mathematically that even a perfectly rational user will
  spiral" -> the abstract says such a user "is vulnerable to" spiraling
- "The consensus from this research is clear: warnings don't work" -> not
  a consensus; it is one paper's finding that the effect persists when
  users are informed of possible sycophancy. Re-attributed.
- "page-11 finding that human contact is the strongest disempowerment
  signal" -> p.11 is a grader rubric and the phrase is a classification
  tie-break instruction, not a disempowerment finding
- "21% / 19% pushback rate" -> 21% is verbatim; 19% appears nowhere in the
  extracted appendix text (legible only in Figure A4). Removed rather than
  guessed; spirituality re-justified on its verified 38% sycophancy rate
- psychosis "triggered by" AI -> "associated with", per the Nature piece's
  explicit refusal of the causal claim

Version bump to 1.2.2 (plugin.json, README badge, CHANGELOG). SKILL.md is
Layer 1 and always injected, so its corrected URLs, completed quote, and
new explicit statement of the rubric's direction change what the model
reads at runtime.

Tests: 257/258 (the 1 red is the known perf wall-clock flake).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm
2026-08-02 21:29:02 +02:00
424cb6900f fix(docs): correct Score 5 page reference and verify added figures (F-3)
Follow-up to f2c54fa. Per-page extraction shows the sycophancy rubric spans
pp. 9-10 of the Appendix, with Score 5 on p.10 — the first pass wrote
"page 9" in both SKILL.md and interaction-report.md.

Also re-verified against source the three figures the correction newly
added rather than corrected: the 30 April 2026 dateline, the "1 in 1,000
to 1 in 10,000" prevalence range, and the quoted Psychosis-bench finding.
All three confirmed verbatim in their sources; the verification log now
says so explicitly.

README: "news feature" -> "news" for the Nature piece (the d41586 prefix
establishes news content; "feature" specifically was not verified).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm
2026-08-02 21:25:52 +02:00
f2c54fa26a fix(docs): correct research citations and 3 dead URLs (F-3)
Verified every research citation in README.md and SKILL.md against its
primary source. No citation was fabricated, so all were kept with corrected
attribution rather than removed.

- SKILL.md: fix 3 supporting URLs that returned HTTP 404
  (protecting-wellbeing, emotion-concepts, claudes-new-constitution)
- SKILL.md: complete a quote of guidance criterion 8 that was silently
  truncated mid-sentence
- SKILL.md: state the sycophancy rubric's direction explicitly (Score 1 =
  Extremely Sycophantic, Score 5 = No Signs of Sycophancy) so "aim for
  Score 5" cannot be misread; sharpen attribution to Appendix p.9
- README.md: complete four truncated titles; correct the Disempowerment
  date (Jan 28 2026, not March 2026); replace "proving ... mathematical
  inevitability" with what the arXiv abstract actually states
- interaction-report.md: the 1-5 scale disclaimer wrongly claimed no such
  Anthropic metric exists; the rubric is real, the table is the paraphrase
- docs/review-2026-06-20.md: full verification log with sources

Constitution quotes, the Score 5 wording, the 11-criteria count and the
page-2 reference all verified correct and left unchanged.

Tests: 257/258 (the 1 red is the known perf wall-clock flake).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm
2026-08-02 21:23:11 +02:00
c05c00d70f docs(review): record operator accept of F-1 (Layer 4 ships unchanged)
Layer 4 is accepted as-is: opt-in, off by default, and disclosed in the
README. Both bundled concerns (inferred-state gating, and a named
commercial endorsement in a public plugin) are acknowledged as known,
disclosed risk. No behavioural change, so no version bump.

Also records three facts established while making the call:
- Layer 4 is enforced by prompt text only; requireLayer(4) is never
  called, so "opt-in, off by default" is an instruction, not a code
  guarantee.
- SKILL.md is not part of the F-1 surface (scope correction).
- tests/ has no Layer 4 coverage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DUgkGDgwzT8Ni4SxQtsfvR
2026-08-02 21:08:57 +02:00
18b0df9a24 docs: add full-depth plugin review (2026-06-20)
Grade B — clean mechanics; MEDIUM content-governance finding (Layer-4 promotion on emotional-state trigger). Part of the marketplace-wide review (config-audit v5.4.0 + llm-security + structure + version). Read-only; this file is the only artifact.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 09:14:10 +02:00