llm-security/docs/ci-cd-guide.md
Kjell Tore Guttormsen b4d9f83521
docs: knowledge-file count from source, no offline claim for the CLI
PLAN § v8.1.3 tillegg b and c.
- b: scanner-reference.md said Knowledge Files (20); knowledge/ holds
  22. Added typosquat-allowlist.json and workflow-injection-patterns.md;
  the test pins header and table to the directory.
- c: ci-cd-guide.md claimed zero network calls, OSV opt-in and "no
  cross-border data transfer". Measured in the code: dep runs npm audit
  (package.json) and pip-audit (requirements.txt, if installed), network
  resolves found domains over DNS, supply-chain queries OSV.dev, none
  with a switch. The guide now says so. scanner-reference.md carried the
  same claim plus a `--online` flag that does not exist (0 hits in
  scanners/); fixed in the same commit because the same gate covers it
  (chosen over leaving a known-false line in a file already edited here).
  That extra check was red on fd7de23's text, verified.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 11:37:58 +02:00

7.6 KiB

CI/CD Integration Guide

Integrate llm-security into your CI/CD pipeline for automated security scanning of AI/LLM projects. The standalone CLI runs 14 deterministic Node.js scanners — no AI models and no Anthropic API. It is not air-gapped: three scanners reach the network (see Data Sovereignty).

Data Sovereignty

The standalone CLI is not offline. The scanners analyze your source code locally (Shannon entropy, regex patterns, AST traversal, git log parsing) and do not upload it. Three of them reach the network whenever their input is present, and there is no switch to turn any of them off:

  • dep runs npm audit when the target has a package.json and pip-audit (if installed) when it has a requirements.txt. Both send package names and versions to their registry's advisory service (npm registry, PyPI).
  • network resolves the domains it finds in the code over DNS, so those names reach your configured resolver.
  • supply-chain queries the OSV.dev batch API with package names and versions from the lockfiles, over HTTPS.

Block egress at the network layer if the run must be air-gapped.

What about Claude Code integration? The Claude Code plugin (hooks, agents, commands) uses AI models and sends data to Anthropic. These components are not included in the standalone CLI. When you run npx llm-security scan, only deterministic scanners execute.

Schrems II / NSM Compliance

  • Standalone CLI: source code stays in the run, but package metadata and domain names leave it (above) — evaluate per your organization's data classification
  • OSV.dev queries (always on when lockfiles are present): package names and versions to a Google-operated API
  • npm audit / pip-audit (always on when package.json / requirements.txt is present): package names and versions to the npm registry and PyPI
  • DNS lookups (always on when the code names a domain): domain names found in the code, to your resolver
  • Claude Code plugin: sends code context to Anthropic (US) — requires data processing agreement for regulated environments

Norwegian Regulatory Context

  • NSM Grunnprinsipper: Automated security scanning fulfills GP 3.1 (vulnerability management) and GP 2.4 (secure development)
  • Digitaliseringsdirektoratet: Aligns with recommended practices for AI system development lifecycle security
  • EU AI Act: Directly supports Art. 9 (risk management) and Art. 15 (cybersecurity) requirements. Note the Digital Omnibus (European Parliament 16 June 2026, Council 29 June 2026) deferred the high-risk obligations these articles sit under — to 2 December 2027 for stand-alone Annex III systems and 2 August 2028 for AI embedded in Annex I regulated products. Transparency obligations still apply from 2 August 2026, so this remains preparatory rather than deadline-driven work

5-Minute Setup

GitHub Actions

Copy ci/github-action.yml to .github/workflows/llm-security.yml:

name: LLM Security Scan
on: [push, pull_request]
jobs:
  security-scan:
    runs-on: ubuntu-latest
    permissions:
      security-events: write
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: '18'
      - run: npx llm-security scan . --fail-on high --format sarif --output-file results.sarif
      - uses: github/codeql-action/upload-sarif@v3
        if: always()
        with:
          sarif_file: results.sarif

SARIF results appear in the repository's Security tab under Code scanning alerts.

Azure DevOps

Copy ci/azure-pipelines.yml to your pipeline, or include the scan step in an existing pipeline:

steps:
  - task: NodeTool@0
    inputs:
      versionSpec: '18.x'
  - script: npx llm-security scan . --fail-on high --format sarif --output-file $(Build.ArtifactStagingDirectory)/results.sarif
    displayName: Run llm-security scan
  - task: PublishBuildArtifacts@1
    condition: always()
    inputs:
      pathToPublish: $(Build.ArtifactStagingDirectory)/results.sarif
      artifactName: llm-security-scan

For Azure DevOps Advanced Security, replace PublishBuildArtifacts@1 with AdvancedSecurity-Publish@1.

GitLab CI

Add to .gitlab-ci.yml:

llm-security-scan:
  image: node:18-alpine
  stage: test
  script:
    - npx llm-security scan . --fail-on high --format sarif --output-file results.sarif
  artifacts:
    paths:
      - results.sarif
    reports:
      sast: results.sarif
    when: always

SAST report parsing requires GitLab Ultimate. On Free/Premium tiers, download the SARIF artifact manually.

Configuration

CLI Flags

Flag Description
--fail-on <severity> Exit 1 if any finding at or above severity exists. Values: critical, high, medium, low
--compact One-liner per finding format. Reduces CI log noise
--format sarif Output OASIS SARIF 2.1.0 (default: JSON)
--output-file <path> Write full results to file. Stdout gets compact aggregate
--baseline Diff against stored baseline (show new/resolved findings)
--save-baseline Save current results as baseline for future diffs

Policy File

Configure defaults in .llm-security/policy.json:

{
  "ci": {
    "failOn": "high",
    "compact": true
  }
}

CLI flags always take precedence over policy file values.

Environment Variables

Variable Description
LLM_SECURITY_PRECOMPACT_MODE block|warn|off for the PreCompact transcript scan (default warn)
LLM_SECURITY_UPDATE_CHECK=off Disable the daily update-check HTTP call

The audit trail moved to the policy file in v8.0.0: set audit.log_path in .llm-security/policy.json instead of LLM_SECURITY_AUDIT_LOG.

Exit Codes

Code Meaning When
0 Clean / below threshold No findings at or above --fail-on level, or ALLOW verdict
1 Threshold exceeded Findings at or above --fail-on level, or WARNING verdict (without --fail-on)
2 Block BLOCK verdict (only without --fail-on)

With --fail-on, exit codes are binary: 0 (clean) or 1 (threshold exceeded). Without --fail-on, the legacy tri-state (0/1/2) is preserved.

What Gets Scanned

The 14 deterministic scanners cover:

Scanner Detects
Unicode Zero-width characters, homoglyphs, Unicode Tag steganography
Entropy High-entropy strings (potential secrets/tokens)
Permission Overly broad permissions, missing tool justification
Dependency Known vulnerable packages, typosquats
Taint Untrusted input flows to sensitive operations
Git forensics Force pushes, sensitive file history, author anomalies
Network Suspicious URLs, exfiltration endpoints, C2 patterns
Memory poisoning Injection patterns in CLAUDE.md, memory files, rules
Supply chain Lockfile audit, blocklists, OSV.dev (always on)
Workflow CI/CD workflow injection — untrusted triggers, unpinned actions, spoofed bots
Trigger abuse Activation-surface abuse in command/agent/skill frontmatter — shadowing, baiting, overly broad triggers
Signature Known-malware identity match (webshells, reverse shells, cryptominers, hacktools)
AST taint Scope-aware Python taint analysis (parse-only), falls back to regex taint tracing
Toxic flow Lethal trifecta correlation (input + access + exfil)

Local Testing

Test the exact same command locally before adding to CI:

# With npx (requires npm publish)
npx llm-security scan . --fail-on high --compact

# With local clone
node bin/llm-security.mjs scan . --fail-on high --compact