llm-ingestion-pipeline-secu.../CONTRIBUTING.md
Kjell Tore Guttormsen 1625f3893b docs: version-sync + SECURITY/CONTRIBUTING + honest-limits (Session E)
- README: tests badge 275->357; status v0.1->v0.2 (repo is 0.2.0; the v1.0
  bump belongs to the Session G freeze, not this docs pass); add three
  honest-limits — lone-HIGH-in-trusted-prose->WARN, vacuous quarantine-floor,
  Cyrillic/Latin homoglyph-mix false positive.
- docs/BRIEF.md: drop "No code yet" pre-implementation framing -> implemented v0.2.
- docs/OKF-INGESTION-BRIEF.md 4: correct cross-link control language —
  absolute https / references/ targets are spec-permitted, not rejected.
- Add SECURITY.md (private Forgejo disclosure) + CONTRIBUTING.md (stdlib-only
  core, Iron-Law TDD, no trailers, Forgejo-only invariants).
2026-07-15 10:08:24 +02:00

3.2 KiB

Contributing

Thanks for looking. This is a small, security-focused library; contributions are welcome as long as they hold the invariants below. When in doubt, open a discussion before a large change.

Where the project lives

  • Canonical repository: Forgejogit.fromaitochitta.com/open/llm-ingestion-pipeline-security.
  • There is no GitHub repository; do not open pull requests there.

Non-negotiable invariants

A change that breaks one of these will not be merged, however useful it is otherwise:

  1. Test-first (the Iron Law). No production code without a failing test first. Add the red test, then the minimal implementation that makes it green. A pure refactor keeps the existing suite green unchanged as its proof of no behaviour change.
  2. Stdlib-only core. pyproject.toml [project].dependencies stays []. The deterministic core imports only the standard library. Extraction / ML / judge libraries (pypdf, python-docx, python-pptx, openpyxl, ML detectors) live only in an extra ([dev], [ml], [judge]), never in core dependencies. The upload front-end and its parsers are dev-scoped showcase code under tests/.
  3. Report / mutation separation. Detectors report findings; they never mutate the scanned text. neutralize is the only defanging surface and it is opt-in and byte-identical on clean input.
  4. Fail-secure / fail-closed stays total. guard() and screen_output must never turn an un-scannable or malformed input into a silent persist.
  5. Calibration lives in one place. Tunable thresholds (entropy floors, MAX_SCAN_CHARS, disposition ranks, cognitive-load lengths, severities) belong in src/llm_ingestion_guard/calibration.py, so the planned Node/TS port can mirror the exact numbers. Do not scatter magic numbers back into detectors.
  6. The shared lexicon is never split. injection_lexicon.json is the single pattern source of truth for both the Python core and the future TS port.
  7. Never ship a live payload. Build attack strings in tests from chr(0x…) fragments so the gitleaks pre-commit hook and the repo itself stay clean. Never bypass the hook with --no-verify.

Running the tests

python -m venv .venv
.venv/bin/pip install -e '.[dev]'        # dev extra: extraction libs for the front-end showcase
PYTHONPATH=src .venv/bin/pytest          # the full suite must be green

The suite is the release gate: a change is not done until the whole suite is green.

Commit style

  • Conventional Commits: type(scope): description (e.g. feat(egress): decode-rescan feeds base64 plaintext to secret-egress).
  • No trailers. Do not add Co-Authored-By / Signed-off-by / tool-attribution trailers.
  • Keep commits scoped: one logical change per commit; version bumps sync every file that names the version in the same commit.

Scope and honest limits

New detection is welcome, but the project ships its limitations as a control (see README.mdHonest limitations). If a change narrows a stated gap, update that section. If it introduces a new deliberate boundary, document it there rather than leaving a silent miss. Absolute claims ("catches all …", "cannot be bypassed") do not belong in this codebase.