AppSecNews
AI Security Commercial, free tier Growing

Lakera Guard

by Lakera

Detection API that classifies prompts and model outputs for injection, jailbreak attempts, sensitive data and policy violations before they land.

Visit lakera.ai (leaves AppSecNews, opens in a new tab) Leaves AppSecNews for the vendor's own site.

No endorsements yet

Run Lakera Guard in production? A named recommendation helps the next team shortlisting it.

Recommend this tool

Endorsers verify their identity through LinkedIn. Titles and companies are self declared, shown as they were when each person signed, and reviewed by an editor before anything is published. Endorsements are never paid for.

What we still need to verify : 3 points in this profile are not yet confirmed against vendor documentation.
  • Corporate ownership and product naming, confirm with vendor
  • Self-hosted deployment availability and terms, confirm
  • Current detector categories and policy controls, confirm

Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.

What it does

Lakera Guard is a classification service you call on the way into and out of an LLM. You send it the prompt, the retrieved context, or the model's response, and it returns a verdict across categories: prompt injection and jailbreak attempts, personally identifiable information, moderated content, and unknown or suspicious links embedded in text. Your application decides what to do with the verdict, so the enforcement logic stays yours while the detection is the vendor's problem.

What distinguishes it is how the detection models are trained. Lakera built a public prompt injection game that attracted an enormous volume of genuine human attack attempts, and that adversarial corpus feeds the classifiers. The models therefore see novel phrasings from real attackers rather than synthetic variations on a published jailbreak list, which is the usual weakness of detection built on static rules. The company also runs an offensive side, automated red teaming against a customer's own application.

Where it fits

This sits in the request path in production, added by the team that owns the GenAI application, usually with security setting the policy thresholds. It runs as a network call, so you trade a few tens of milliseconds and a data flow for detection you do not maintain. Settle two things first: what happens on a positive verdict, and whether prompt content may leave your environment.

Strengths

  • Detection models trained on a large corpus of real human attack attempts rather than a curated list of published jailbreaks.
  • Simple API surface, so adding it to an existing application is a wrapper rather than an architectural change.
  • Covers both directions, catching injected instructions on input and leaked sensitive content on output.
  • Latency is low enough to sit in a synchronous user-facing path, which is the practical bar for a guardrail.

Limitations

  • It is a classifier, and classifiers can be probed. An attacker with repeated access learns what gets through, and detection quality on genuinely novel techniques is unknowable in advance.
  • The hosted model routes prompt content to a third party, which is a data governance conversation in regulated environments.
  • Detection is not enforcement. If your application ignores the verdict or handles it badly, the guard has bought you nothing.

Who it suits

Good for product teams shipping customer-facing GenAI who want credible injection detection without staffing a research function for it. Less appropriate where every inference must stay inside your boundary and self-hosting is not available to you, or where the risk you actually have is model supply chain rather than prompt abuse.

Used Lakera Guard? Recommend it under your own name and title.

Recommend this tool