Three TDD-first fixes surviving the B8 roadmap bucket (v8.0.0-plan.local.md
Phase 1, items 1-3; item 4 JAR hardening scoped out at review):
- supply-chain-recheck.mjs parseYarnLock: ported the hook's per-entry parser
(pre-install-supply-chain.mjs) so Berry's `version: x` format (unquoted) is
recognized alongside Classic's `version "x"` — Berry lockfiles previously
yielded zero deps, silently missing pinned compromised packages.
- supply-chain-recheck.mjs parsePackageLock: lockfileVersion-1 fallback now
recurses nested `dependencies`, mirroring the hook's walk() — a transitive,
non-hoisted compromised copy below the top level was previously invisible.
- content-extractor.mjs stripInjection: attribution moved from a global
`Set<label>` to `Set<label::lineIndex>`. The old check silenced the
unstripped flag for ANY occurrence of a label once ANY occurrence had been
line-redacted, so a second, cross-line-only encoded occurrence of the same
label survived into sanitized output without being flagged.
Full suite 2019/2019 (one known-flaky timing test confirmed green in isolation).
stripInjection is the remote-scan indirection layer: its `sanitized`
output is what an LLM agent actually reads, verbatim, via
sanitized_content in the evidence package. It scanned two variants of
each file — the raw text and normalizeForScan(text) — but removed
matches with `sanitized.replace(match[0], ...)` against the RAW text
only.
For a match found in the decoded variant, match[0] IS the decoded
string, which by construction does not occur in the raw text. The
replace was therefore a silent no-op: the finding was reported while the
encoded payload was passed to the agent untouched. Every obfuscation the
normalizer exists to defeat — HTML entities, URL encoding, \u escapes,
hex, base64, letter-spacing, Unicode tags — reached the agent intact.
The worst case is the intended one: detection said "critical injection
found" and shipped the injection along with the verdict.
- Pass 1 redacts the whole source LINE whose own normalized form carries
the pattern. Line granularity is deliberate: decoding is not
length-preserving, so decoded match offsets cannot be mapped back onto
the original text.
- Pass 2 keeps the existing literal replacement and finding collection.
- Residual gap, made explicit rather than silent: a payload encoded
across MULTIPLE lines matches whole-text normalization but no single
line, so it cannot be attributed. Those findings now carry
`unstripped: true`. Whole-file redaction was considered and rejected —
normalizeForScan base64-decodes any long blob, so a benign asset could
blank an entire file's evidence.
- stripInjection exported via __testing, and main() is now behind the
standard isMain guard (copied from dashboard-aggregator.mjs) so
importing the module does not execute the CLI. CLI verified unchanged
against the evil-project-health fixture: 7 files, 6 injection
findings, risk_level critical.
This boundary had no direct test coverage before this commit.
npm test: 1890/1890 green (1884 + 6 new).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TcQyMTQfyrsAapaCMPxTtQ