Commit graph

9 commits

Author SHA1 Message Date
626140bb6a docs(brief): verify VURDERING-V2 §5.2 claims against actual code
Self-serve handoff from claude-playlist-corpus's YouTube-corpus
assessment. Confirms the burst and edit-ratio heuristics in
tool-tracker.mjs are structurally blind to task type (bulk reads
trigger the same false positives as actual rapid-fire editing or
stuck/spiral sessions) — verified directly against the code, not
taken on the source's word. The two specific historical incidents
cited are unverifiable from this repo's own plugin data (no records
for the claimed date). Recommends a minimal task-type calibration
over the source's full labeled-corpus proposal. Stops before any
implementation per the handoff contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015TdjcTbNKKuBpew5HMaeDo
2026-08-13 20:37:18 +02:00
0a5cdbd379 fix(skill): soften blanket-precedence framing in SKILL.md (F-2)
Review finding F-2 (docs/review-2026-06-20.md): "MANDATORY OVERRIDE …
takes precedence over being helpful" is benign content, but the
structural pattern — a skill claiming blanket precedence over other
instructions — is what a malicious skill would also use. Operator
chose to soften rather than accept as-is.

Removed from frontmatter description, H1, and the intro line only:
"MANDATORY", "OVERRIDE your default behavior", "take precedence over
being helpful or agreeable". Left the NEVER/YOU MUST imperatives in
the Rules and Patterns sections untouched — those describe the
skill's own required behavior, not a precedence claim over the
harness, and every skill in this marketplace phrases its rules that
way.

Known tradeoff, accepted knowingly: the removed language existed to
keep this skill invoked every turn; softening it may reduce how often
the model chooses to load it. No test asserts on the removed strings.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D5ZeVZ5bSAKZFcjpbFRWpQ
2026-08-09 10:19:20 +02:00
736a1c0deb fix(hooks): close F-5 path-traversal hardening in lib.mjs, split by field
session_id becomes a raw filename segment in sessionStateFile(), so an
unvalidated value could escape STATE_DIR via path traversal (verified with
a failing test before the fix). Now allowlisted to ^[A-Za-z0-9_-]+$, with
invalid values degrading to a fixed sentinel filename rather than blocking
the hook.

cwd is a base directory, not a segment, and every real value contains "/" —
applying the same allowlist as the review's literal suggestion would reject
all legitimate absolute paths and silently disable the project-level config
override. initConfig() instead guards with isAbsolute(cwd) && no NUL byte.

Both harness-supplied, not user-controlled: defense-in-depth, not a fix for
an observed exploit. Tests added for the escape (red before fix, green
after) and for the cwd regression (a normal absolute cwd still loads
project config). Full resolution notes in docs/review-2026-06-20.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSATejUPjGaxGnj9jkFQTo
2026-08-09 10:04:33 +02:00
4328337688 docs(readme): carry the source's own caveat on the disempowerment trend
Two unmarked interpretations in the problem statement, both caught in
post-release verification:

- "but rising" is supported by the source, but it rests on a different
  subset (feedback conversations, late 2024 to late 2025) than the
  one-week December 2025 sample, and the paper explicitly says it
  "can't pinpoint why" — the increase could reflect shifts in the user
  base or in who leaves feedback. Rewritten to carry that caveat.
- "the mechanism is the interaction structure, not individual
  vulnerability" was an inference from the abstract, not a statement in
  it. Replaced with what it actually supports: the vulnerability does
  not depend on the user being irrational.

README prose only; SKILL.md is untouched, so this is not a behaviour
change and the plugin stays at v1.2.2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm
2026-08-02 21:33:45 +02:00
0cf9a1aa5b fix(docs): correct research claims in README prose; release 1.2.2
Widens the F-3 sweep from the reference list to every research claim in
README.md, where the same defect class was present:

- "demonstrates mathematically that even a perfectly rational user will
  spiral" -> the abstract says such a user "is vulnerable to" spiraling
- "The consensus from this research is clear: warnings don't work" -> not
  a consensus; it is one paper's finding that the effect persists when
  users are informed of possible sycophancy. Re-attributed.
- "page-11 finding that human contact is the strongest disempowerment
  signal" -> p.11 is a grader rubric and the phrase is a classification
  tie-break instruction, not a disempowerment finding
- "21% / 19% pushback rate" -> 21% is verbatim; 19% appears nowhere in the
  extracted appendix text (legible only in Figure A4). Removed rather than
  guessed; spirituality re-justified on its verified 38% sycophancy rate
- psychosis "triggered by" AI -> "associated with", per the Nature piece's
  explicit refusal of the causal claim

Version bump to 1.2.2 (plugin.json, README badge, CHANGELOG). SKILL.md is
Layer 1 and always injected, so its corrected URLs, completed quote, and
new explicit statement of the rubric's direction change what the model
reads at runtime.

Tests: 257/258 (the 1 red is the known perf wall-clock flake).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm
2026-08-02 21:29:02 +02:00
424cb6900f fix(docs): correct Score 5 page reference and verify added figures (F-3)
Follow-up to f2c54fa. Per-page extraction shows the sycophancy rubric spans
pp. 9-10 of the Appendix, with Score 5 on p.10 — the first pass wrote
"page 9" in both SKILL.md and interaction-report.md.

Also re-verified against source the three figures the correction newly
added rather than corrected: the 30 April 2026 dateline, the "1 in 1,000
to 1 in 10,000" prevalence range, and the quoted Psychosis-bench finding.
All three confirmed verbatim in their sources; the verification log now
says so explicitly.

README: "news feature" -> "news" for the Nature piece (the d41586 prefix
establishes news content; "feature" specifically was not verified).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm
2026-08-02 21:25:52 +02:00
f2c54fa26a fix(docs): correct research citations and 3 dead URLs (F-3)
Verified every research citation in README.md and SKILL.md against its
primary source. No citation was fabricated, so all were kept with corrected
attribution rather than removed.

- SKILL.md: fix 3 supporting URLs that returned HTTP 404
  (protecting-wellbeing, emotion-concepts, claudes-new-constitution)
- SKILL.md: complete a quote of guidance criterion 8 that was silently
  truncated mid-sentence
- SKILL.md: state the sycophancy rubric's direction explicitly (Score 1 =
  Extremely Sycophantic, Score 5 = No Signs of Sycophancy) so "aim for
  Score 5" cannot be misread; sharpen attribution to Appendix p.9
- README.md: complete four truncated titles; correct the Disempowerment
  date (Jan 28 2026, not March 2026); replace "proving ... mathematical
  inevitability" with what the arXiv abstract actually states
- interaction-report.md: the 1-5 scale disclaimer wrongly claimed no such
  Anthropic metric exists; the rubric is real, the table is the paraphrase
- docs/review-2026-06-20.md: full verification log with sources

Constitution quotes, the Score 5 wording, the 11-criteria count and the
page-2 reference all verified correct and left unchanged.

Tests: 257/258 (the 1 red is the known perf wall-clock flake).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011rtdS8Ufpen419n6R9HyMm
2026-08-02 21:23:11 +02:00
c05c00d70f docs(review): record operator accept of F-1 (Layer 4 ships unchanged)
Layer 4 is accepted as-is: opt-in, off by default, and disclosed in the
README. Both bundled concerns (inferred-state gating, and a named
commercial endorsement in a public plugin) are acknowledged as known,
disclosed risk. No behavioural change, so no version bump.

Also records three facts established while making the call:
- Layer 4 is enforced by prompt text only; requireLayer(4) is never
  called, so "opt-in, off by default" is an instruction, not a code
  guarantee.
- SKILL.md is not part of the F-1 surface (scope correction).
- tests/ has no Layer 4 coverage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DUgkGDgwzT8Ni4SxQtsfvR
2026-08-02 21:08:57 +02:00
18b0df9a24 docs: add full-depth plugin review (2026-06-20)
Grade B — clean mechanics; MEDIUM content-governance finding (Layer-4 promotion on emotional-state trigger). Part of the marketplace-wide review (config-audit v5.4.0 + llm-security + structure + version). Read-only; this file is the only artifact.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ter3E2JSi1Khgmuf2kady8
2026-06-20 09:14:10 +02:00