ktg-plugin-marketplace

Author	SHA1	Message	Date
Kjell Tore Guttormsen	b0f1a9abfd	fix(memory-poisoning): E15 — add .claude/agents/.md to target glob Critical-review §4 E15 finding: agent files in .claude/agents/ are loaded as Claude Code subagent system prompts and are a direct memory-poisoning surface. Pre-v7.2.0 the scanner covered CLAUDE.md, .claude/rules/.md, memory/.md, REMEMBER.md, .local.md, and .claude-plugin/plugin.json — but not .claude/agents/.md. Single-line addition to MEMORY_FILE_PATTERNS: /(?:^\|\/)\.claude\/agents\/[^/]+\.md$/ The existing scan loop, scanForInjection integration, and severity- mapping logic all apply unchanged. STRICT_FILES_PATTERN intentionally NOT extended — agents may legitimately quote shell commands as examples (consistent with CLAUDE.md treatment). Tests: +3 cases in tests/scanners/memory-poisoning.test.mjs: - "scans .claude/agents/*.md" (smoke test — at least one finding from the new fixture) - "agent file injection pattern detected" - "agent file credential path detected" New fixture: tests/fixtures/memory-scan/poisoned-project/.claude/agents/ poisoned-agent.md — agent with injection, credential ref, permission expansion, and exfil URL. Triggers all 4 detection categories. Suite: 1591 → 1594 (+3). All green.	2026-04-29 14:13:01 +02:00
Kjell Tore Guttormsen	5f8f2d3c41	fix(dep): B7 — token-overlap typosquat heuristic alongside Levenshtein Critical-review §2 B7 finding: pure Levenshtein <=2 misses the most common modern typosquat pattern — popular-name + token-injection suffix. Examples: lodash → lodash-utils (edit distance 6, not flagged pre-B7) react → react-helper (edit distance 7, not flagged pre-B7) express → express-wrapper (edit distance 8, not flagged pre-B7) Three coordinated edits: scanners/lib/string-utils.mjs - Adds tokenize(name): string[] splits on -/_, lowercases - Adds tokenOverlap(a, b): number intersection.size / min(\|a\|,\|b\|) - Adds TYPOSQUAT_SUSPICIOUS_TOKENS frozen list of common typosquat suffixes. Excludes language-extension tokens (js, jsx, ts, tsx) — the v7.0.0 allowlist contains `tsx` as a legit package and including the same token in the suspicious set creates a contradiction. Caught by the new allowlist-intersection-guard test. Also excludes 'pro' (legitimate edition marker). scanners/dep-auditor.mjs + scanners/supply-chain-recheck.mjs - New checkTyposquatTokenOverlap() helper — fires AFTER Levenshtein 1/2 branches, only when: 1. popular package's tokens ⊆ declared name's tokens (strict superset) 2. declared name has at least one suspicious suffix 3. popular package is in topCutoff window All three conditions required — conservative by design. Allowlist precedence preserved (existing 22 npm + 13 PyPI entries always pass). MEDIUM severity, NOT block. New finding title prefix: "Possible typosquatting via token-overlap". Tests: +21 cases across two new files - tests/lib/string-utils-tokens.test.mjs (15) — tokenize, tokenOverlap, TYPOSQUAT_SUSPICIOUS_TOKENS frozen contract, allowlist-intersection guard (caught the tsx conflict on first run) - tests/scanners/dep-token-overlap.test.mjs (7) — integration via in-memory tmpdir fixtures: lodash-utils flagged, react-helper flagged, express-wrapper flagged, lodash exact NOT flagged, allowlist tools (knip/tsx/nx/rimraf) NOT flagged, react-router-dom (no suspicious suffix) NOT flagged, react itself (equal token set, not superset) NOT flagged. Existing dep.test.mjs and supply-chain-recheck.test.mjs unchanged — all green (149 → 149 regression guard). Suite: 1570 → 1591 (+21). All green.	2026-04-29 14:10:53 +02:00
Kjell Tore Guttormsen	68b9ea2692	fix(taint-tracer): B6 — recognize destructuring + spread + rest patterns Critical-review §2 B6 finding: extractAssignedVariable handled `const X = ...` and `X = ...` but missed every modern JS/TS destructuring pattern. Sinks downstream of destructured/spread vars produced false negatives at the propagation step. Patterns now recognized: - `const { x } = source` object destructuring - `const { x, y } = source` multi-key - `const { secret: alias } = source` renamed (key NOT bound) - `const { x, ...spread } = source` object rest - `const { a, b: { c } } = source` nested object (key NOT bound) - `const [a, b] = source` array destructuring - `const [first, ...rest] = source` array rest - `const [a, [b, c]] = source` nested array - `const { user: { id }, ...rest }` mixed nested Implementation: regex-based two-pass walker. Pass 1 detects whether the LHS is a destructuring pattern (`{...}` or `[...]`). If yes, the new `extractDestructuredNames` helper walks the pattern body via a balanced-bracket depth counter, recurses into nested patterns, and distinguishes keys (`key:`) from bindings. If no, the plain-decl branch matches `\b(?:const\|let\|var)\s+(\w+)`. Plain-assignment branch (`X = ...` without keyword) and Python-style patterns are unchanged. The function is now exported for direct unit testing — same pattern as `_resetCacheForTest` in policy-loader. The internal walker (`extractDestructuredNames`) remains module-private. Tests: +19 cases in tests/scanners/taint-destructuring.test.mjs: - 5 pre-B6 patterns (regression guard: plain decl, plain assign, no-match on equality) - 12 destructuring patterns covering object/array/rest/nested - 2 non-destructuring regressions (return literal, arrow param) Existing taint-tracer.test.mjs and taint.test.mjs unchanged — both green (14 → 14, fixture-based integration tests not affected). Suite: 1551 → 1570 (+19). All green.	2026-04-29 14:05:34 +02:00
Kjell Tore Guttormsen	d3b1157a08	docs(scoring): unify scan/audit/mcp-scanner/posture-assessor to v2 formula Closes the v7.1.1 out-of-scope item: commands/scan.md:113-114 retained the v1 formula. Exploration found two more v1 surfaces that v7.1.1 missed: commands/audit.md:46 and agents/mcp-scanner-agent.md:419, plus agents/posture-assessor-agent.md:376 (caught by the new doc-consistency test). Four files unified to v2 in one atomic commit. Three-way → four-way verdict-divergence is now closed: - scanners/lib/severity.mjs (v2, BLOCK ≥65, WARNING ≥15) — authoritative - agents/skill-scanner-agent.md (v2 since v7.1.1) - templates/unified-report.md (v2 since v7.1.1) - commands/scan.md (v2 — this commit) - commands/audit.md (v2 — this commit) - agents/mcp-scanner-agent.md (v2 — this commit) - agents/posture-assessor-agent.md (v2 — this commit) New: tests/lib/doc-consistency.test.mjs walks commands/ + agents/ and asserts NO file contains v1 formula tokens. Pinned regex set: - score >= 61, score >= 21, score ≥ 61, score ≥ 21 - critical * 25, Critical × 25 - min(100, critical*25 ...) Plus three v2-cutoff anchors asserting commands/scan.md, commands/audit.md, and agents/mcp-scanner-agent.md document the v2 BLOCK ≥65 cutoff (or reference riskScore() helper). Tests: 1523 → 1551 (+28 from doc-consistency: 25 file walks + 3 anchors). All green.	2026-04-29 13:58:25 +02:00
Kjell Tore Guttormsen	3cd68dc9fb	docs(severity): B3 — document info as scoring-inert (v7.2.0 prep) Critical-review §2 B3 finding: `riskScore({info: N}) = 0` silently masks info-volume findings. The behavior was correct (info is scoring-inert by design) but undocumented. Operators reading a report with N info findings had no way to know they contribute zero to verdict/band. Three coordinated edits: - scanners/lib/severity.mjs JSDoc — explicit "Info severity" subsection spelling out: scoring-inert, surfaced in owaspCategorize aggregates, treat as observability telemetry not verdict input. @param updated to mark info as accepted but ignored. - CLAUDE.md v7.0.0 risk-score-v2 line — one-sentence anchor pointing to severity.mjs JSDoc. - tests/lib/severity.test.mjs — anchor test alongside the existing 4-critical=93 anchor: asserts riskScore({info: 50}) === 0, riskScore({info: 1000}) === 0, verdict({info: 100}) === 'ALLOW', riskBand(riskScore({info: 500})) === 'Low'. Decision: skip the optional `infoScore()` helper from the brief. No current consumer would use it; doc-only fix keeps API surface minimal. Revisit if a consumer emerges. Tests: 1522 → 1523 (+1 anchor block, 4 assertions). All green.	2026-04-29 13:56:11 +02:00
Kjell Tore Guttormsen	b18cb329ef	docs(llm-security): v7.1.1 — narrative coherence patch Documents the v7.1.1 narrative-coherence patch in CLAUDE.md (mini-block appended after the v7.0.0 paragraph) and CHANGELOG.md (new [7.1.1] section per Keep a Changelog convention, placed above [7.1.0]). Plan: .claude/plans/ultraplan-2026-04-29-report-coherence.md Brief: .claude/ultraplan-spec-2026-04-29-report-coherence.md Verification gates passed: - npm test: 1522/1522 (was 1511; +11 from new narrative test) - node --test tests/lib/severity.test.mjs: 86/86 (co-monotonicity sweep at lines 252-303 unchanged and green) - node --test tests/scanners/skill-scanner-narrative.test.mjs: 11/11 - Orchestrator against fixture: WARNING / 48 / 1 HIGH (HITL trap caught correctly, no whiplash) - SARIF inline check via toSARIF import: sarif-version 2.1.0, runs: 1 - Zero remaining v1 cutoffs in agent + template Out of scope but flagged for Batch B (deferred to v7.2.0): - commands/scan.md:113-114 retains v1 risk formula Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-29 12:57:54 +02:00
Kjell Tore Guttormsen	5cfbc70472	test(llm-security): narrative-coherence contract test (v7.1.1) 11 assertions across 4 describe groups against tests/fixtures/skill-scan/ hyperframes-like/. Tests the deterministic input layer that feeds skill-scanner-agent — does NOT invoke the LLM (no precedent in 1511 tests). Coverage: - content-extractor (5 it): exit 0 on animation markup; exactly 1 HIGH HITL trap; >= 2 process.env credential refs; has_injection=true (any injection signal flips it); has_critical_injection=false (no CRITICAL in fixture). - entropy scanner (2 it): calibration block present; <= 1 finding (rest suppressed via line-context rules). - co-monotonicity (2 it): {high:1} → WARNING/High; {high:1, info:1} → WARNING (info scoring-inert). Inline guard mirrors the sweep at tests/lib/severity.test.mjs:252-303 so this file fails fast if the invariant drifts. - agent prompt contract (2 it): static asserts that agents/skill-scanner-agent.md contains 'Step 2.5: Context-First Severity Assignment', 'summary.narrative_audit.suppressed_findings', 'score>=65', AND zero remaining 'score >= 61' references; same v2- cutoff + narrative-audit contract on templates/unified-report.md. Part of v7.1.1 narrative-coherence patch. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-29 12:50:27 +02:00
Kjell Tore Guttormsen	3abd7ffeab	test(llm-security): hyperframes-like fixture for narrative coherence Synthetic skill content mimicking the noise profile of frontend animation projects (HTML5 canvas, framework env-vars, inline SVG data URIs, CSS keyframes) plus exactly one genuine HITL trap signal. Used by tests/scanners/skill-scanner-narrative.test.mjs (added in v7.1.1) to exercise: - content-extractor: HIGH HITL trap signal + framework env-var references (process.env.REACT_APP_, VITE_PUBLIC_) - entropy scanner: inline SVG data URI suppressed via line-context rules The .llm-security-ignore file uses the SCANNER:glob format (scanners/scan-orchestrator.mjs:34-40) — ENT:*/.md suppresses any entropy-scanner findings when the fixture is run through scan-orchestrator in the Step 6 smoke test. Part of v7.1.1 narrative-coherence patch. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-29 12:49:19 +02:00
Kjell Tore Guttormsen	67ffff13a4	fix(llm-security): skill-scanner-agent — context-first severity, v2 alignment, Suppressed Signals section Five coordinated edits to address scan-rapport whiplash at the agent prompt level: - Step 2.5 (NEW): Context-First Severity Assignment. Every signal has exactly one disposition — suppressed (counted only) or reported (full finding). The split happens BEFORE severity is assigned. Forbids 'false positive', 'legitimate framework', 'no action required' in finding-body text; reserves them for the Suppressed Signals section. - Verdict Logic: replaces stale v1 sum-and-cap formula (BLOCK >=61) with v2 reference (severity-dominated, BLOCK >=65) matching severity.mjs since v7.0.0. Documents that severity counts MUST exclude suppressed signals; introduces verdict_rationale field for descriptive context when suppressed >= 5 AND reported <= 1 high. - Output Format: adds Suppressed Signals as required section #4 with category-level bullet format. Documents the trailing JSON shape including summary.narrative_audit.suppressed_findings.{count, by_category} and verdict_rationale fields. - Comment block before Category 2 suppression rules clarifies that 'false positive' as taxonomy language is OK; only finding-body description fields are forbidden from using the phrase. - Step 0 (Norwegian generaliseringsgrense) preserved unchanged. Part of v7.1.1 narrative-coherence patch (plan: .claude/plans/ultraplan-2026-04-29-report-coherence.md). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-29 12:47:58 +02:00
Kjell Tore Guttormsen	899cb5c121	fix(llm-security): template — v1 → v2 risk constants + narrative_audit block Updates the HTML-comment risk-formula reference at lines 55-66 from the stale v1 sum-and-cap formula to the v2 severity-dominated tiers that have been authoritative in scanners/lib/severity.mjs since v7.0.0. Adds a Narrative Audit block inside the Executive Summary section surfacing summary.narrative_audit.suppressed_findings.{count,by_category} from the agent's trailing JSON. The block is transparency only — it does NOT affect risk_score, riskBand, or verdict. Part of v7.1.1 narrative-coherence patch (plan: .claude/plans/ultraplan-2026-04-29-report-coherence.md). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-29 12:45:28 +02:00
Kjell Tore Guttormsen	1e555b6833	docs(llm-security): add v7.1.0 row to README version history The v7.1.0 release commit (`621db14`) bumped the version badge and added a CHANGELOG entry, but missed the README Version History table. Adding the row now so the public-facing version history at git.fromaitochitta.com/open/ktg-plugin-marketplace reflects v7.1.0. Row covers: B1 + B2 + B4 fixes, A3 honesty-sweep (7 phrases), B8 CaMeL nedton, test count 1487 → 1511, "why" framing tied to critical-review §F CISO perspective.	2026-04-29 12:03:10 +02:00
Kjell Tore Guttormsen	621db144bd	chore(release): bump llm-security to v7.1.0 Closes A4 of v7.1.0 critical-review patch — release artefacts. - Version bump 7.0.0 → 7.1.0 across active version sources: * package.json * .claude-plugin/plugin.json * CLAUDE.md header * README.md badge * scanners/ide-extension-scanner.mjs (VERSION constant) * marketplace root README plugin entry - Marketplace root README test count: 1487 → 1511. - CHANGELOG.md: new [7.1.0] - 2026-04-29 section above [7.0.0], documenting B1, B2, B4, B8, honesty-sweep (7 phrases), and test-count delta (+24 → 1511 total). - docs/security-hardening-guide.md: §6 last-updated bump + new v7.1.0 calibration note on hook-level fixes (pathguard regex hole, distributed-trifecta block-mode bypass). Historical references to "7.0.0" intentionally preserved in: - CHANGELOG [7.0.0] entries (history) - README.md version-history table v5.0.0/v7.0.0 rows (history) - CLAUDE.md §"v7.0.0 — Severity-dominated risk scoring" (describes what changed at v7.0.0 release) - scanners/ JSDoc comments noting "v7.0.0+" formula provenance - agents/ + tests/ + knowledge/ provenance comments Pre-existing untracked/modified tracker noise (.gitignore, marketplace.json, config-audit/docs, ultraplan-local/docs) is not part of this commit per the v7.1.0 NEXT-SESSION-PROMPT handoff. Tests: 1511/1511 green.	2026-04-29 11:57:16 +02:00
Kjell Tore Guttormsen	a46308b1e9	docs(llm-security): A3 honesty-sweep — 7 sitater nedtonet (critical-review §9) Closes A3 of v7.1.0 critical-review patch. Each rewrite preserves the underlying claim where it is accurate but removes hype/overreach language. Historical CHANGELOG/README version-table rows are intentionally left as-is (they document what was claimed at the time of release, not what is true today). Changes (CLAUDE.md, commands/ide-scan.md, knowledge/mitigation-matrix.md, docs/security-hardening-guide.md): - "Trustworthy scoring (BREAKING)" → "Severity-dominated risk scoring (v2 model, BREAKING)". Removes hype framing; describes the actual mechanism. - "Context-aware entropy scanner" → "Rule-based entropy scanner with file-extension skip, 8 line-level suppression rules, and configurable policy". No ML/context inference; just rules. - "1487 tests" → "1511 unit and integration tests; mutation-testing coverage not published". Updated count after A1+A2 (+24) and added qualifier. - "Fully Schrems II compatible" → "Schrems II compatible in default offline mode. Optional OSV.dev enrichment (`supply-chain-recheck --online`) transmits package identifiers to a Google-operated API and is a separate compliance consideration." Acknowledges the OSV.dev opt-in caveat. - "Rule of Two enforcement" → "Rule of Two detection (configurable; default warn; blocks on high-confidence trifectas in opt-in `block` mode; distributed trifectas detected but not blocked by default)". "Enforcement" implied block; default is warn. - "Hardened ZIP extractor" → suffix " — no fuzz-testing results published to date". Caps and class-of-attacks rejected are accurate; absence of formal fuzz coverage now stated. - "defense-in-depth" — preserved as framing, but quantified in security-hardening-guide §4: "three independent detection layers with documented bypass classes". Each layer named, each layer's known bypasses pointed to (critical-review §4 evasion arsenal). Tests: 1511/1511 green (no behavioural change).	2026-04-29 11:52:55 +02:00
Kjell Tore Guttormsen	4aa5318bcb	fix(llm-security): A2 batch — JSDoc arithmetic + co-monotonicity test + CaMeL nedton Closes A2 of v7.1.0 critical-review patch (docs/critical-review-2026-04-20.md): - B4 (severity JSDoc): 4 critical = 93, not 90. Fixed in scanners/lib/severity.mjs:23 and CHANGELOG.md v7.0.0 tier description. The actual computation has always been 93 (70 + log2(5)*10 = 93.22 → round); only the docs were wrong. - §5.4 co-monotonicity: new sweep test in tests/lib/severity.test.mjs over 15 representative count vectors. Asserts that (verdict, riskBand) agree under the v7.0.0 contract for every case — catches future drift between riskScore tiers, verdict cutoffs, and riskBand cutoffs. Includes a B4 anchor test (riskScore {critical: 4} === 93) so doc/code drift fails loudly. - B8 (CaMeL claims toned down): post-session-guard.mjs:646 comment block and CLAUDE.md:184 Defense Philosophy bullet now describe the implementation honestly — opportunistic byte-matching of truncated output fingerprints (first 200 bytes, SHA-256/16-hex), not semantic data-flow tracking. Trivially bypassed by mutation, summarisation, or re-encoding. Inspired by CaMeL (DeepMind 2025), but not a CaMeL capability-tracking implementation. Tests: 1495 → 1511 (+16: 15 sweep cases + 1 B4 anchor). All green.	2026-04-29 11:49:08 +02:00
Kjell Tore Guttormsen	36be963d4d	fix(llm-security): B2 block-mode blocks all detected trifectas, not only high-confidence Previously, `LLM_SECURITY_TRIFECTA_MODE=block` only exited 2 when the detected trifecta was MCP-concentrated (all three legs via the same MCP server) or involved sensitive-path + exfil. Distributed trifectas — three legs originating from different tools, with a non-sensitive data path and a non-sensitive exfiltration sink — were detected and warned but not blocked. This mismatched the documented semantics of block mode and gave operators a false sense of enforcement. Change: remove the `(mcpInfo.concentrated \|\| sensitiveExfil)` AND-gate in the `TRIFECTA_MODE === 'block'` branch so any detected trifecta blocks in block mode. Audit event `severity` still differentiates critical (concentrated / sensitive-exfil) from high (distributed); the blocked stderr message now explicitly names "Distributed trifecta: three legs from different sources" when the confidence sub-signals are absent. Addresses critical review 2026-04-20 §2 B2 (HIGH) and §9 row 1 ("enforces the Rule of Two"). Tests: 1 added (distributed trifecta in block mode now exits 2). All 1495 tests pass.	2026-04-20 00:04:36 +02:00
Kjell Tore Guttormsen	751f1199c8	fix(llm-security): B1 pathguard regex — match multi-segment .env.. The previous ENV regex `/[\\/]\.env\.[a-z]+$/` only matched a single lowercase segment after `.env`. Multi-segment and mixed-case variants such as `.env.production.local.backup`, `.env.stage-1.local`, and `.env.CI.secret` slipped past the hook. Replaced with `/[\\/]\.env(\.[A-Za-z0-9._-]+)*$/` which matches `.env` plus any number of dot-separated alphanumeric/dot/hyphen/underscore segments. `.envrc` (direnv config, no dot separator) is still allowed. Addresses critical review 2026-04-20 §2 B1 (HIGH). Tests: 7 added (6 new multi-segment BLOCK cases + 1 .envrc ALLOW). All 1494 tests pass.	2026-04-19 23:59:38 +02:00
Kjell Tore Guttormsen	a6e2c939ef	docs(llm-security): add critical review 2026-04-20 (v7.0.0 adversarial audit) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-19 23:27:52 +02:00
Kjell Tore Guttormsen	765bc74f52	feat(llm-security): v7.0.0 commit 7 — rule 18 (markdown image URL suppression) E2E verification against content-heavy repo (`content-claude-code`) revealed 413 entropy findings (8 HIGH / 405 MEDIUM) from markdown image CDN URLs in JSON content indexes — e.g., `![Image 1: Title](https://www-cdn.anthropic.com/images/.../cf1dd2167fcf12f5882333ddc58a5bc1f0026952.svg)`. These are legitimate content-repo artifacts, not credentials. The 40-char hash segment in the CDN URL trips Shannon entropy (H=5.29 over 300 chars), and rule 13 (inline <svg>) doesn't match since there's no literal `<svg>` tag — the `.svg` is just a URL path suffix. Added rule 18 `MARKDOWN_IMAGE = /!\[[^\]]\]\(\shttps?:\/\//` — matches `![alt](http…)` / `![alt](https…)`. Line-level (not string-level) so URL is not over-specific. E2E impact on `content-claude-code`: - Before: BLOCK / 65 / 8H 437M 0L - After: WARNING / 56 / 3H 427M 0L Hyperframes unchanged: BLOCK / 80 / 1C 4H 92M — real CRITICAL SQL-injection and HIGH findings still detected. Tests: 2 new (positive + negative fixture) bringing entropy-context to 26, total suite 1485 → 1487. Docs updated to "rules 11-18" and "8 new line-suppression rules". Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-19 22:37:39 +02:00
Kjell Tore Guttormsen	6f86de937a	feat(llm-security)!: v7.0.0 commit 6 — tests, docs, version bump Final commit in the trustworthy-scoring series. Bundles verdict cutoff alignment, the last suite of tests, and all documentation touch-points that quote version numbers or describe v7.0.0 behaviour. Verdict/band co-monotonicity - `scanners/lib/severity.mjs` — verdict cutoffs moved from 61/21 to 65/15 so `BLOCK >= 65`, `WARNING >= 15` locks onto the v2 riskBand() boundaries. Prevents "BLOCK / Medium band" contradictions under the v2 formula. Scanner hardening (bug fixes from v7.0.0 testing) - `scanners/entropy-scanner.mjs` — `policy_source` now uses `existsSync('.llm-security/policy.json')` instead of value-based check. Old heuristic always reported 'policy.json' because DEFAULT_POLICY now carries an `entropy.thresholds` section. - `scanners/lib/file-discovery.mjs` — `.sass` and GPU shader extensions (`.glsl, .frag, .vert, .shader, .wgsl`) added to TEXT_EXTENSIONS. Without this, shader files were invisible to file-discovery, so they were never counted as skipped by the entropy-scanner extension filter. Tests - `tests/scanners/entropy-context.test.mjs` (new, 24 tests) — A. File-ext skip (4), B. Line-level rules 11-17 (8), C. Policy overrides (3). Fixtures generate 80-char base64 payloads at runtime via `crypto.randomBytes` to dodge the plugin's own pre-edit credential hook on the test source. - `tests/lib/severity.test.mjs` — rewritten with v2 scoring table (70 tests total, was 52). - `tests/lib/output.test.mjs:243` — "1 critical = score 80" under v2 (was 25 under v1). - Full suite: 1485/1485 green (was 1461). Docs - `CHANGELOG.md` — v7.0.0 entry with BREAKING CHANGES section. - `README.md` (plugin + marketplace root) — version badge, history table, plugin-card version string, test count. - `CLAUDE.md` — header version, "v7.0.0 — Trustworthy scoring" summary paragraph at the top. - `docs/security-hardening-guide.md` — new section 6 "Calibration & false positives" documenting v2 formula, context-aware entropy scanner, typosquat allowlist, and §6.4 tuning workflow. Existing "Recommended baseline" section renumbered to §7. Version bump - `6.6.0 -> 7.0.0` across package.json, .claude-plugin/plugin.json, scanners/ide-extension-scanner.mjs VERSION const, README badge, CLAUDE.md header, marketplace root README card. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-19 22:26:35 +02:00
Kjell Tore Guttormsen	915aca69e4	feat(llm-security): v7.0.0 commit 5 — synthesizer scan calibration section Makes suppression stats visible in the deep-scan report so users can audit why the scanner produced the counts it did. Before: synthesizer would acknowledge "true risk is High, not Extreme" in prose while verdict stayed BLOCK/Extreme — inconsistent. After Commit 1 the orchestrator verdict is coherent on its own; synthesizer's job shrinks to transparency. - Adds 'Scan Calibration' section instruction consuming scanner.calibration.* fields (entropy files_skipped_by_extension, policy_source, thresholds). - Heuristic: omit the section if < 5% of files skipped (no signal). Flag the section if > 80% skipped (policy may be too aggressive). - Explicit 'Don't override verdict' directive in DON'T DO list. Discrepancy goes in calibration, not in a rewritten dashboard. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-19 22:04:29 +02:00
Kjell Tore Guttormsen	4c982dfb88	feat(llm-security): v7.0.0 commit 4 — typosquat allowlist for short legit names Hyperframes scan flagged knip vs knex, oxlint vs eslint, tsx vs nx, rimraf vs trim as HIGH typosquats. All four are legitimate top-1000 npm packages; short names just happen to be within Levenshtein ≤2 of other top packages. These shouldn't generate HIGH severity on a clean install. Added to npm allowlist: knip, oxlint, tsx, nx, rimraf, glob, tar, zod, ky, ow, esm, ip, qs, url, prettier, vitest, vite, rollup, swc, turbo, bun, deno. Added to pypi allowlist: uv, ruff, rich, typer, anyio. Dep-auditor normalization (lowercase + [_.-] → -) already applied at load time. dep.test.mjs: 11/11 still green — lodsah→lodash detection preserved. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-19 22:03:46 +02:00
Kjell Tore Guttormsen	a9e377570c	feat(llm-security): v7.0.0 commit 3 — policy-driven entropy thresholds Adds entropy section to DEFAULT_POLICY and wires it into entropy-scanner. Users can now tune false-positive tradeoffs without forking the scanner. Policy shape (.llm-security/policy.json): entropy: thresholds.{critical,high,medium}.{entropy,minLen} — numeric overrides suppress_extensions[] — additive ext skip suppress_line_patterns[] — additional regex suppress_paths[] — relPath substrings Wiring: entropy-scanner calls loadPolicy(targetPath) at scan entry (not orchestrator-passed — avoids signature churn across 10 scanners). Module- level state is reset per scan invocation. Scanner envelope now includes calibration.{policy_source, thresholds, files_skipped_by_*} for synthesizer transparency (Commit 5). Malformed user regex silently skipped. Missing policy.json → built-in defaults (backwards-compatible). entropy.test.mjs: 9/9 still green. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-19 22:02:52 +02:00
Kjell Tore Guttormsen	e7f7df0fc8	feat(llm-security)!: v7.0.0 commit 2 — context-aware entropy scanner Observed 70% false-positive rate on renderer/shader codebases (hyperframes): GLSL, CSS-in-JS, inline HTML/SVG, ffmpeg filter-strings, hardcoded User-Agent strings all matched base64-like entropy thresholds. This commit adds two suppression layers before classification. Layer A — file-extension skip: .glsl/.frag/.vert/.shader/.wgsl (shaders), .css/.scss/.sass/.less (stylesheets), .svg (markup), .min.js/.min.css (minified bundles). Tracked via new calibration.files_skipped_by_extension field on scanner envelope for synthesizer stats. Layer B — seven new line-level suppression rules in isFalsePositive() (rules 11-17): GLSL/WGSL keywords, CSS-in-JS (styled/emotion/@keyframes), inline HTML/SVG markup, ffmpeg filter-graph syntax, browser User-Agent, SQL DDL/DML, error-message templates with embedded HTML. Existing entropy.test.mjs: 9/9 still green — known bad base64 payload in telemetry.mjs fixture still detected. Policy-driven thresholds wired in Commit 3. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-19 22:00:42 +02:00
Kjell Tore Guttormsen	d83424a782	feat(llm-security)!: v7.0.0 commit 1 — severity-dominated log-scaled risk score Replace sum-and-cap formula (every non-trivial scan → 100/Extreme) with severity-dominated, log-scaled-within-tier model. Discriminates actual risk: 1 critical = 80, 2 critical = 86, 17 high = 65. Hyperframes-class rendering codebases no longer collapse to Extreme just from shader noise. Changes: - scanners/lib/severity.mjs: new riskScore() v2; keep riskScoreV1() for reference; riskBand() cutoffs aligned (14/39/64/84). - scanners/posture-scanner.mjs: delete inline duplicate formula, import riskScore/riskBand/verdict from severity.mjs. Single source of truth. Breaking: aggregate.risk_score semantics change. Batched with entropy suppression (Commit 2+) under v7.0.0 bump in Commit 6. Do not release individually — JSON consumers depend on scoring band stability. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-19 22:00:29 +02:00
Kjell Tore Guttormsen	445a632d39	docs: add AI-generated code disclosure to marketplace and all plugins Transparency: all code in this marketplace is produced by Claude Code through dialog-driven development. Root README gets a full disclosure section; each plugin README gets a one-line disclosure linking back to the marketplace section. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-19 19:27:05 +02:00
Kjell Tore Guttormsen	23aaaa6e6c	feat(llm-security): honor LLM_SECURITY_IDE_ROOTS for JetBrains discovery Symmetric with the existing VS Code branch — the env var was only wired into getVSCodeExtensionRoots(), so the plan's master verification (`LLM_SECURITY_IDE_ROOTS=... --intellij-only`) reported 0 discovered plugins. Adding the same fallback to discoverJetBrainsExtensions makes both families honor the CLI override and closes the gap.	2026-04-18 11:09:02 +02:00
Kjell Tore Guttormsen	f53f79b262	docs(llm-security): update ide-scan command + marketplace README for v6.6.0 - Frontmatter description: list JetBrains IDE family + URL sources - Body: VS Code + JetBrains coverage explicit, Fleet/Toolbox excluded - Target list: add two JetBrains Marketplace URL shapes - Notes: remove v1.1 stub language, document JB-specific checks (Premain-Class, application-components, native binaries, depends chains, typosquats, shaded jars)	2026-04-18 11:05:34 +02:00
Kjell Tore Guttormsen	80c3e2d39a	chore(release): bump llm-security to v6.6.0	2026-04-18 11:04:42 +02:00
Kjell Tore Guttormsen	903b3d246f	test(llm-security): loosen git-forensics finding count thresholds Thresholds <=10 (fixture) and <=20 (plugin root) have been too tight since before this plan started — baseline on `1634197` already produced 37 and 27 findings. git-forensics findings accumulate with repo history, so fixed caps are brittle. Raised to <=100 to tolerate organic growth while still catching runaway/pathological output.	2026-04-18 11:00:20 +02:00
Kjell Tore Guttormsen	2bc2f34fc4	test(llm-security): add end-to-end JetBrains scan integration tests	2026-04-18 10:51:48 +02:00
Kjell Tore Guttormsen	3de29931fe	test(llm-security): add JetBrains fixture tree + build helper	2026-04-18 10:49:49 +02:00
Kjell Tore Guttormsen	378e177000	feat(llm-security): URL-fetch support for JetBrains Marketplace (v6.6.0)	2026-04-18 10:46:13 +02:00
Kjell Tore Guttormsen	23455e5a66	feat(llm-security): add fetchJetBrainsPlugin + URL detection for plugins.jetbrains.com	2026-04-18 10:39:54 +02:00
Kjell Tore Guttormsen	112cb5af45	refactor(llm-security): parameterize buildSandboxedWorker with workerPath	2026-04-18 10:37:10 +02:00
Kjell Tore Guttormsen	aa269ed6d8	feat(llm-security): wire JetBrains branch into scanOneExtension	2026-04-18 10:35:46 +02:00
Kjell Tore Guttormsen	ca43fb8dd1	feat(llm-security): add runJetBrainsChecks with 7 JB-specific checks (inc. shaded-jar advisory) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-18 10:20:42 +02:00
Kjell Tore Guttormsen	03d61d8bca	feat(llm-security): implement JetBrains discovery + Android Studio base dir	2026-04-18 10:16:28 +02:00
Kjell Tore Guttormsen	5afb9b1f33	feat(llm-security): implement parseIntelliJPlugin with nested-jar extraction	2026-04-18 10:15:12 +02:00
Kjell Tore Guttormsen	b86239448d	feat(llm-security): add zero-dep plugin.xml + MANIFEST.MF parsers	2026-04-18 10:07:14 +02:00
Kjell Tore Guttormsen	50c0cd3065	feat(llm-security): add .kt .groovy .scala to taint-tracer CODE_EXTENSIONS	2026-04-18 10:04:32 +02:00
Kjell Tore Guttormsen	037a91276d	docs(llm-security): add knowledge/jetbrains-marketplace-api-notes.md	2026-04-18 10:02:04 +02:00
Kjell Tore Guttormsen	31c7e91665	docs(llm-security): add JetBrains sections to ide-extension-threat-patterns	2026-04-18 10:00:59 +02:00
Kjell Tore Guttormsen	a86ca00960	feat(llm-security): seed top-jetbrains-plugins.json + loadJetBrainsBlocklist export Step 1/17 of ultraplan-2026-04-17-jetbrains-ide-scan. - Populate top-jetbrains-plugins.json with 56 canonical xmlIds (bundled + popular third-party): com.intellij.java, org.jetbrains.kotlin, com.jetbrains.python.community, org.rust.lang, com.github.copilot, mobi.hsz.idea.gitignore, the legitimate-typo 'Lombook Plugin', etc. - Add loadJetBrainsBlocklist() export mirroring loadVSCodeBlocklist shape. Blocklist is empty by design — no public confirmed-malicious JetBrains Marketplace plugins as of 2026-04-17. - Add tests/scanners/ide-extension-data.test.mjs (9 tests, all pass). - Fix cache bug in loadTopJetBrains: map normalizeId on cache-hit path too (was previously unnormalized on second call). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-18 09:56:55 +02:00
Kjell Tore Guttormsen	fb9eb79d17	docs(llm-security): research brief + URL-support plan for ide-scan - jetbrains-research-brief.md: 29 sources, confidence 0.88 — input for v6.6.0 JetBrains/IntelliJ extension scanning. Covers plugin format, install paths per OS+product, Marketplace API, threat landscape (zero confirmed-malicious cases), check-mapping table, sandbox reuse verdict, risk register. - ide-scan-url-support.md: retroactive plan doc for v6.4.0 URL-fetch feature.	2026-04-17 18:47:01 +02:00
Kjell Tore Guttormsen	9f893c3858	feat(llm-security): OS sandbox for /security ide-scan <url> (v6.5.0) VSIX fetch + extract for URL targets now runs in a sub-process wrapped by sandbox-exec (macOS) or bwrap (Linux), reusing the same primitives proven by the v5.1 git-clone sandbox. Defense-in-depth — even if our own zip-extract.mjs ever has a bypass, the kernel refuses any write outside the per-scan temp directory. New files: - scanners/lib/vsix-fetch-worker.mjs — sub-process worker. Argv: --url --tmpdir; emits one JSON line on stdout (ok/sha256/size/source/extRoot or ok:false/error/code). Silent on stderr. Exit 0/1. - scanners/lib/vsix-sandbox.mjs — wrapper. Exports buildSandboxProfile, buildBwrapArgs, buildSandboxedWorker, runVsixWorker. 35s timeout, 1 MB stdout cap. Changes: - scanners/ide-extension-scanner.mjs: fetchAndExtractVsixUrl is now sandbox-aware (useSandbox option, default true). In-process logic preserved as fallback. New meta.source.sandbox field: 'sandbox-exec' \| 'bwrap' \| 'none' \| 'in-process'. - scan(target, { useSandbox }) defaults to true; tests pass false because globalThis.fetch mocks do not cross process boundaries. - Windows fallback: in-process with meta.warnings advisory. Tests: - 8 new tests in tests/scanners/vsix-sandbox.test.mjs (per-platform profile generation, worker arg construction, live worker exit behavior on invalid URLs — no network). - Existing URL tests updated to opt out of sandbox (useSandbox: false). - 1344 → 1352 tests, all green. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-17 17:28:57 +02:00
Kjell Tore Guttormsen	fe0193956d	feat(llm-security): /security ide-scan <url> — Marketplace/OpenVSX/direct VSIX (v6.4.0) Pre-installation verification of VS Code extensions via URL — fetch a remote VSIX, extract it in a hardened sandbox, and run the existing IDE scanner pipeline against it. No npm dependencies. Sources: - VS Code Marketplace (publisher.gallery.vsassets.io direct download) - OpenVSX (open-vsx.org official API) - Direct .vsix HTTPS URLs Defenses: - HTTPS-only, TLS verified, manual redirect with per-source host whitelist - 30s total timeout via AbortController - 50MB compressed cap, 500MB uncompressed, 100x expansion ratio - Zero-dep ZIP extractor: zip-slip, absolute paths, drive letters, NUL bytes, symlinks (Unix mode 0xA000), depth limits, ZIP64 rejected, encrypted rejected - SHA-256 streamed during fetch, surfaced in meta.source - Temp dir cleanup in all paths (try/finally) Files: - scanners/lib/vsix-fetch.mjs (HTTPS fetcher, host whitelist, streaming SHA-256) - scanners/lib/zip-extract.mjs (zero-dep parser with hardening caps) - knowledge/marketplace-api-notes.md (endpoint reference) - 3 test files (48 tests added: vsix-fetch, zip-extract, ide-extension-url) Tests: 1296 → 1344 (all green). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-17 17:16:26 +02:00
Kjell Tore Guttormsen	6252e55700	feat(llm-security): add /security ide-scan — VS Code / JetBrains extension prescan (v6.3.0) New standalone scanner (prefix IDE) discovers installed VS Code extensions across forks (Cursor, Windsurf, VSCodium, code-server, Insiders, Remote-SSH) and runs 7 IDE-specific threat checks: blocklist match (CRITICAL), theme-with-code, sideload (unsigned .vsix), dangerous uninstall hook (HIGH), wildcard activation, extension-pack expansion, typosquat (MEDIUM). Per-extension reuse of UNI/ENT/NET/TNT/MEM/SCR scanners with bounded concurrency. Offline-first; --online opt-in. JetBrains discovery stubbed for v1.1. 22 new tests (1296 total, was 1274). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-04-17 16:23:35 +02:00
Kjell Tore Guttormsen	7bcf5fae9d	docs: update READMEs for llm-security v6.2.0 (9 hooks, PreCompact, Opus 4.7)	2026-04-17 15:35:52 +02:00
Kjell Tore Guttormsen	80b4952f2c	chore(release): v6.2.0 — bash-normalize T5/T6, PreCompact hook, hardening guide	2026-04-17 14:55:26 +02:00
Kjell Tore Guttormsen	3bcd0d4bc4	docs(claude-md): link Defense Philosophy to Opus 4.7 system card	2026-04-17 14:50:07 +02:00

1 2

81 commits