AppSecNews
AI Security Open source Growing

LLM Guard

by Protect AI

Python library of composable input and output scanners that sanitize prompts and validate model responses entirely within your own environment.

Visit github.com (leaves AppSecNews, opens in a new tab) Leaves AppSecNews for the vendor's own site.

No endorsements yet

Run LLM Guard in production? A named recommendation helps the next team shortlisting it.

Recommend this tool

Endorsers verify their identity through LinkedIn. Titles and companies are self declared, shown as they were when each person signed, and reviewed by an editor before anything is published. Endorsements are never paid for.

What we still need to verify : 2 points in this profile are not yet confirmed against vendor documentation.
  • Current maintainership and release cadence following vendor acquisition, confirm
  • Complete scanner inventory, confirm against project docs

Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.

What it does

LLM Guard is a library of scanners you compose into two pipelines, one for prompts and one for responses. The input scanners cover the obvious hazards and several less obvious ones: prompt injection detection using a fine-tuned transformer classifier, secret detection, personally identifiable information anonymization, banned topics and substrings, code detection, language identification, token limit enforcement, and invisible or control character stripping that catches instructions hidden in text a human reviewer cannot see.

Output scanners run the other direction. They check for sensitive data in a response, deanonymize placeholders the input stage substituted, validate JSON structure, detect refusals, flag malicious or unreachable URLs, and score relevance and factual consistency against the supplied context. The anonymize and deanonymize pair is the neat trick: sensitive values are replaced with tokens held in a local vault before the prompt is sent and restored afterward, so the model never sees the real values and the user never sees the placeholders.

Where it fits

In the application process, in production, wherever your LLM calls originate. Application engineers wire the scanner chain in; security decides which scanners are blocking and which are advisory. There is also a container that exposes the pipeline over an API for non-Python services. Everything runs locally, so the prerequisite is capacity to host and warm transformer models, plus the patience to tune thresholds against your own traffic before turning anything to blocking mode.

Strengths

  • Fully self-hosted, so prompt content never leaves your environment, which is frequently the deciding factor over hosted detection APIs.
  • The anonymize and deanonymize pattern solves a real problem that plain redaction does not, preserving answer quality while protecting values.
  • Scanner set is wide and pragmatic, including invisible text and token limits that many guardrail projects skip.
  • Each scanner is independent, so you can run a handful rather than paying for the whole chain.

Limitations

  • Several scanners load transformer models. Memory and CPU cost is real, and cold start latency will surprise you in a serverless deployment.
  • Scanner accuracy varies widely by category. The injection classifier in particular needs evaluation against your own traffic, since false positives on legitimate technical prompts are common.
  • Open-source guardrail projects live or die on maintenance, and the project's cadence after its vendor changed hands is worth checking before you depend on it.

Who it suits

A strong fit for Python teams that need guardrails inside their own boundary and are willing to tune them, particularly in regulated environments where sending prompts to a detection vendor is not an option. Less appropriate for teams with no capacity to host models or tune thresholds, who will get more reliable results from a managed detection service.

Used LLM Guard? Recommend it under your own name and title.

Recommend this tool