PLAN § v8.1.3 tillegg b and c. - b: scanner-reference.md said Knowledge Files (20); knowledge/ holds 22. Added typosquat-allowlist.json and workflow-injection-patterns.md; the test pins header and table to the directory. - c: ci-cd-guide.md claimed zero network calls, OSV opt-in and "no cross-border data transfer". Measured in the code: dep runs npm audit (package.json) and pip-audit (requirements.txt, if installed), network resolves found domains over DNS, supply-chain queries OSV.dev, none with a switch. The guide now says so. scanner-reference.md carried the same claim plus a `--online` flag that does not exist (0 hits in scanners/); fixed in the same commit because the same gate covers it (chosen over leaving a known-false line in a file already edited here). That extra check was red on fd7de23's text, verified. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
178 lines
7.6 KiB
Markdown
178 lines
7.6 KiB
Markdown
# CI/CD Integration Guide
|
|
|
|
Integrate llm-security into your CI/CD pipeline for automated security scanning of AI/LLM projects. The standalone CLI runs 14 deterministic Node.js scanners — no AI models and no Anthropic API. It is not air-gapped: three scanners reach the network (see Data Sovereignty).
|
|
|
|
## Data Sovereignty
|
|
|
|
**The standalone CLI is not offline.** The scanners analyze your source code locally (Shannon entropy, regex patterns, AST traversal, git log parsing) and do not upload it. Three of them reach the network whenever their input is present, and there is no switch to turn any of them off:
|
|
|
|
- **dep** runs `npm audit` when the target has a `package.json` and `pip-audit` (if installed) when it has a `requirements.txt`. Both send package names and versions to their registry's advisory service (npm registry, PyPI).
|
|
- **network** resolves the domains it finds in the code over DNS, so those names reach your configured resolver.
|
|
- **supply-chain** queries the [OSV.dev](https://osv.dev/) batch API with package names and versions from the lockfiles, over HTTPS.
|
|
|
|
Block egress at the network layer if the run must be air-gapped.
|
|
|
|
**What about Claude Code integration?** The Claude Code plugin (hooks, agents, commands) uses AI models and sends data to Anthropic. These components are **not included** in the standalone CLI. When you run `npx llm-security scan`, only deterministic scanners execute.
|
|
|
|
### Schrems II / NSM Compliance
|
|
|
|
- Standalone CLI: source code stays in the run, but package metadata and domain names leave it (above) — evaluate per your organization's data classification
|
|
- OSV.dev queries (always on when lockfiles are present): package names and versions to a Google-operated API
|
|
- `npm audit` / `pip-audit` (always on when `package.json` / `requirements.txt` is present): package names and versions to the npm registry and PyPI
|
|
- DNS lookups (always on when the code names a domain): domain names found in the code, to your resolver
|
|
- Claude Code plugin: sends code context to Anthropic (US) — requires data processing agreement for regulated environments
|
|
|
|
### Norwegian Regulatory Context
|
|
|
|
- **NSM Grunnprinsipper:** Automated security scanning fulfills GP 3.1 (vulnerability management) and GP 2.4 (secure development)
|
|
- **Digitaliseringsdirektoratet:** Aligns with recommended practices for AI system development lifecycle security
|
|
- **EU AI Act:** Directly supports Art. 9 (risk management) and Art. 15 (cybersecurity) requirements. Note the Digital Omnibus (European Parliament 16 June 2026, Council 29 June 2026) deferred the high-risk obligations these articles sit under — to 2 December 2027 for stand-alone Annex III systems and 2 August 2028 for AI embedded in Annex I regulated products. Transparency obligations still apply from 2 August 2026, so this remains preparatory rather than deadline-driven work
|
|
|
|
## 5-Minute Setup
|
|
|
|
### GitHub Actions
|
|
|
|
Copy `ci/github-action.yml` to `.github/workflows/llm-security.yml`:
|
|
|
|
```yaml
|
|
name: LLM Security Scan
|
|
on: [push, pull_request]
|
|
jobs:
|
|
security-scan:
|
|
runs-on: ubuntu-latest
|
|
permissions:
|
|
security-events: write
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
- uses: actions/setup-node@v4
|
|
with:
|
|
node-version: '18'
|
|
- run: npx llm-security scan . --fail-on high --format sarif --output-file results.sarif
|
|
- uses: github/codeql-action/upload-sarif@v3
|
|
if: always()
|
|
with:
|
|
sarif_file: results.sarif
|
|
```
|
|
|
|
SARIF results appear in the repository's **Security** tab under **Code scanning alerts**.
|
|
|
|
### Azure DevOps
|
|
|
|
Copy `ci/azure-pipelines.yml` to your pipeline, or include the scan step in an existing pipeline:
|
|
|
|
```yaml
|
|
steps:
|
|
- task: NodeTool@0
|
|
inputs:
|
|
versionSpec: '18.x'
|
|
- script: npx llm-security scan . --fail-on high --format sarif --output-file $(Build.ArtifactStagingDirectory)/results.sarif
|
|
displayName: Run llm-security scan
|
|
- task: PublishBuildArtifacts@1
|
|
condition: always()
|
|
inputs:
|
|
pathToPublish: $(Build.ArtifactStagingDirectory)/results.sarif
|
|
artifactName: llm-security-scan
|
|
```
|
|
|
|
For Azure DevOps Advanced Security, replace `PublishBuildArtifacts@1` with `AdvancedSecurity-Publish@1`.
|
|
|
|
### GitLab CI
|
|
|
|
Add to `.gitlab-ci.yml`:
|
|
|
|
```yaml
|
|
llm-security-scan:
|
|
image: node:18-alpine
|
|
stage: test
|
|
script:
|
|
- npx llm-security scan . --fail-on high --format sarif --output-file results.sarif
|
|
artifacts:
|
|
paths:
|
|
- results.sarif
|
|
reports:
|
|
sast: results.sarif
|
|
when: always
|
|
```
|
|
|
|
SAST report parsing requires GitLab Ultimate. On Free/Premium tiers, download the SARIF artifact manually.
|
|
|
|
## Configuration
|
|
|
|
### CLI Flags
|
|
|
|
| Flag | Description |
|
|
|------|-------------|
|
|
| `--fail-on <severity>` | Exit 1 if any finding at or above severity exists. Values: `critical`, `high`, `medium`, `low` |
|
|
| `--compact` | One-liner per finding format. Reduces CI log noise |
|
|
| `--format sarif` | Output OASIS SARIF 2.1.0 (default: JSON) |
|
|
| `--output-file <path>` | Write full results to file. Stdout gets compact aggregate |
|
|
| `--baseline` | Diff against stored baseline (show new/resolved findings) |
|
|
| `--save-baseline` | Save current results as baseline for future diffs |
|
|
|
|
### Policy File
|
|
|
|
Configure defaults in `.llm-security/policy.json`:
|
|
|
|
```json
|
|
{
|
|
"ci": {
|
|
"failOn": "high",
|
|
"compact": true
|
|
}
|
|
}
|
|
```
|
|
|
|
CLI flags always take precedence over policy file values.
|
|
|
|
### Environment Variables
|
|
|
|
| Variable | Description |
|
|
|----------|-------------|
|
|
| `LLM_SECURITY_PRECOMPACT_MODE` | `block\|warn\|off` for the PreCompact transcript scan (default `warn`) |
|
|
| `LLM_SECURITY_UPDATE_CHECK=off` | Disable the daily update-check HTTP call |
|
|
|
|
The audit trail moved to the policy file in v8.0.0: set `audit.log_path` in
|
|
`.llm-security/policy.json` instead of `LLM_SECURITY_AUDIT_LOG`.
|
|
|
|
## Exit Codes
|
|
|
|
| Code | Meaning | When |
|
|
|------|---------|------|
|
|
| `0` | Clean / below threshold | No findings at or above `--fail-on` level, or ALLOW verdict |
|
|
| `1` | Threshold exceeded | Findings at or above `--fail-on` level, or WARNING verdict (without `--fail-on`) |
|
|
| `2` | Block | BLOCK verdict (only without `--fail-on`) |
|
|
|
|
With `--fail-on`, exit codes are binary: 0 (clean) or 1 (threshold exceeded). Without `--fail-on`, the legacy tri-state (0/1/2) is preserved.
|
|
|
|
## What Gets Scanned
|
|
|
|
The 14 deterministic scanners cover:
|
|
|
|
| Scanner | Detects |
|
|
|---------|---------|
|
|
| Unicode | Zero-width characters, homoglyphs, Unicode Tag steganography |
|
|
| Entropy | High-entropy strings (potential secrets/tokens) |
|
|
| Permission | Overly broad permissions, missing tool justification |
|
|
| Dependency | Known vulnerable packages, typosquats |
|
|
| Taint | Untrusted input flows to sensitive operations |
|
|
| Git forensics | Force pushes, sensitive file history, author anomalies |
|
|
| Network | Suspicious URLs, exfiltration endpoints, C2 patterns |
|
|
| Memory poisoning | Injection patterns in CLAUDE.md, memory files, rules |
|
|
| Supply chain | Lockfile audit, blocklists, OSV.dev (always on) |
|
|
| Workflow | CI/CD workflow injection — untrusted triggers, unpinned actions, spoofed bots |
|
|
| Trigger abuse | Activation-surface abuse in command/agent/skill frontmatter — shadowing, baiting, overly broad triggers |
|
|
| Signature | Known-malware identity match (webshells, reverse shells, cryptominers, hacktools) |
|
|
| AST taint | Scope-aware Python taint analysis (parse-only), falls back to regex taint tracing |
|
|
| Toxic flow | Lethal trifecta correlation (input + access + exfil) |
|
|
|
|
## Local Testing
|
|
|
|
Test the exact same command locally before adding to CI:
|
|
|
|
```bash
|
|
# With npx (requires npm publish)
|
|
npx llm-security scan . --fail-on high --compact
|
|
|
|
# With local clone
|
|
node bin/llm-security.mjs scan . --fail-on high --compact
|
|
```
|