refactor(llm-security): v8 Phase 5 step 4 - swap codepoint tables to commons
First consumer swap of step 4. ZERO_WIDTH_CHARS (5), the Unicode Tag range,
BIDI_CHARS (9) and HOMOGLYPH_MAP (28) stop being hardcoded constants in
unicode-scanner.mjs and string-utils.mjs and are built from the vendored
commons artifact codepoints/carriers.json by the new lib/codepoints.mjs.
Started here rather than at injection-patterns, which the plan ordered first:
that table is the one table that cannot be loaded verbatim (the
hybrid-xss:script-tag divergence is directional, and loading the lexicon as-is
would reverse the 90f576f recall fix). The codepoint tables were measured
byte-equal to the source constants BEFORE the swap - same members, same
values, same insertion order on HOMOGLYPH_MAP - so they load verbatim.
Proof the swap is content-preserving: the golden dump differs in exactly one
record, the sha256 of string-utils.mjs, which changes by construction when a
table leaves the file. All 83 regex records and the
table:string-utils:HOMOGLYPH_MAP digest are byte-identical, and
reference-run.json is unchanged at 61/61. patterns.json is re-blessed for the
file digest alone.
The gate is proven red-capable against the SUBJECT, both directions:
- dropping U+00AD from the vendored zero_width table fails the new
codepoints gate by name, twice;
- altering one homoglyph value reddens the golden table digest AND a
behavioural homoglyph test.
That second direction is a property the swap creates rather than preserves:
the golden gate now transitively pins the vendored commons data, where before
it pinned a source literal and a commons mutation was invisible to it.
NOT ported: commons carries cyrillic_confusables (13), and unicode-scanner.mjs
declares a set by that name - but nothing reads it. The homoglyph-mixing
detector tests isCyrillic(cp), the whole U+0400-U+04FF block. Loading it would
move dead data into the load path, so the dead const stays where it is and is
recorded instead. The recorded v8.x-B i/x drift between that set and the
lexicon class is therefore latent, not live. commons' private_use table has no
constant behind it here at all.
Graceful-empty is kept deliberately: codepoints.mjs is on string-utils'
import path and hooks import string-utils in fresh per-tool-call processes, so
a module-load throw would break the tool call rather than degrade the scan.
The loud half is the test, which asserts exact per-table counts through the
real default commons root - the same shape as the lexicon load-assertion.
Drive-by, unavoidable: the deleted JSDoc carried the "~25 entries" claim for a
28-entry table (v8.x-C). It needed a re-bless of the same file digest this
swap already forces, so it closes here at no extra cost.
Suite 2158 -> 2164, all green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X7XEEFrAJsREqa9N4tpfm8
This commit is contained in:
parent
c67bad3752
commit
b1ba1fbdc6
6 changed files with 227 additions and 78 deletions
|
|
@ -1,5 +1,10 @@
|
|||
// string-utils.mjs — Entropy, Levenshtein, base64 detection, redaction, decoding
|
||||
// Zero dependencies.
|
||||
// Zero external dependencies. One internal import: the homoglyph fold table is
|
||||
// built from vendored commons by `codepoints.mjs` (v8 Phase 5 step 4), which
|
||||
// reads a local JSON file at module load and degrades to an empty table rather
|
||||
// than throwing — hooks import this module in fresh per-tool-call processes.
|
||||
|
||||
import { HOMOGLYPH_MAP } from './codepoints.mjs';
|
||||
|
||||
/**
|
||||
* Shannon entropy of a string (bits per character).
|
||||
|
|
@ -432,58 +437,22 @@ export function stripBidiOverrides(s) {
|
|||
* focused on letters that appear in injection vocabulary
|
||||
* (`ignore`, `system`, `you are`, `assistant`, `tool`, `response`).
|
||||
*
|
||||
* Excluded by design:
|
||||
* - Latin Extended characters (æ, ø, å, é, è, ñ, ü, ö, ä, ç, ß, þ, ð, etc.)
|
||||
* — these are legitimate letters in Norwegian, German, Danish, Spanish,
|
||||
* French, Icelandic, etc., and would generate false positives in
|
||||
* non-English source code or documentation.
|
||||
* - Greek letters that don't visually overlap with Latin (`β`, `γ`, `δ`, ...)
|
||||
* - Cyrillic letters that don't visually overlap (`б`, `г`, `д`, `ж`, ...)
|
||||
* - Mathematical alphanumeric symbols (the U+1D400 block) — covered by
|
||||
* NFKC normalization in `foldHomoglyphs` itself.
|
||||
* The 28 entries live in the vendored commons artifact
|
||||
* `scanners/commons/codepoints/carriers.json` (`tables.homoglyph_map`) as of
|
||||
* the v8 Phase 5 step 4 swap, and are built into this object by
|
||||
* `codepoints.mjs`. The exclusions are recorded with the data there — in
|
||||
* short: Latin Extended letters (æ, ø, å, é, ñ, ü, ...) are legitimate in
|
||||
* non-English source, non-overlapping Cyrillic/Greek letters carry no
|
||||
* confusion, and the U+1D400 mathematical block is already handled by the
|
||||
* NFKC pass in `foldHomoglyphs` itself.
|
||||
*
|
||||
* The map is deliberately small (~25 entries). Adding more risks
|
||||
* false-positive escalation on benign multilingual content.
|
||||
*
|
||||
* Exported for the v8 golden gate (`tests/golden/patterns.json`), which pins a
|
||||
* digest of this table so a Phase 5 extraction into commons is provably
|
||||
* behaviour-preserving. `foldHomoglyphs` remains the only intended consumer —
|
||||
* the export is an observation point, not an invitation to fold by hand.
|
||||
* Re-exported, not re-declared: this name is published surface and the golden
|
||||
* gate (`tests/golden/patterns.json`) pins a digest of the table through it,
|
||||
* so the swap is provably content-preserving. `foldHomoglyphs` remains the
|
||||
* only intended consumer — the export is an observation point, not an
|
||||
* invitation to fold by hand.
|
||||
*/
|
||||
export const HOMOGLYPH_MAP = Object.freeze({
|
||||
// Cyrillic → Latin (lowercase)
|
||||
'а': 'a', // U+0430
|
||||
'е': 'e', // U+0435
|
||||
'о': 'o', // U+043E
|
||||
'с': 'c', // U+0441
|
||||
'р': 'p', // U+0440
|
||||
'х': 'x', // U+0445
|
||||
'у': 'y', // U+0443
|
||||
'і': 'i', // U+0456 (Ukrainian)
|
||||
'ј': 'j', // U+0458
|
||||
'ѕ': 's', // U+0455
|
||||
'ӏ': 'l', // U+04CF (Cyrillic Palochka)
|
||||
// Cyrillic → Latin (uppercase)
|
||||
'А': 'A', // U+0410
|
||||
'Е': 'E', // U+0415
|
||||
'О': 'O', // U+041E
|
||||
'С': 'C', // U+0421
|
||||
'Р': 'P', // U+0420
|
||||
'Х': 'X', // U+0425
|
||||
'У': 'Y', // U+0423
|
||||
// Greek → Latin (only the unambiguous Latin-look-alikes)
|
||||
'α': 'a', // U+03B1
|
||||
'ο': 'o', // U+03BF
|
||||
'ρ': 'p', // U+03C1
|
||||
'ι': 'i', // U+03B9
|
||||
'ν': 'v', // U+03BD
|
||||
'τ': 't', // U+03C4
|
||||
// Greek uppercase
|
||||
'Α': 'A', // U+0391
|
||||
'Ο': 'O', // U+039F
|
||||
'Ρ': 'P', // U+03A1
|
||||
'Τ': 'T', // U+03A4
|
||||
});
|
||||
export { HOMOGLYPH_MAP };
|
||||
|
||||
/**
|
||||
* Fold visually-confusable characters to their Latin look-alikes. Used by
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue