What we still need to verify : 3 points in this profile are not yet confirmed against vendor documentation.
- Corporate ownership and product naming, confirm with vendor
- Self-hosted deployment availability and terms, confirm
- Current detector categories and policy controls, confirm
Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.
What it does
Lakera Guard is a classification service you call on the way into and out of an LLM. You send it the prompt, the retrieved context, or the model's response, and it returns a verdict across categories: prompt injection and jailbreak attempts, personally identifiable information, moderated content, and unknown or suspicious links embedded in text. Your application decides what to do with the verdict, so the enforcement logic stays yours while the detection is the vendor's problem.
What distinguishes it is how the detection models are trained. Lakera built a public prompt injection game that attracted an enormous volume of genuine human attack attempts, and that adversarial corpus feeds the classifiers. The models therefore see novel phrasings from real attackers rather than synthetic variations on a published jailbreak list, which is the usual weakness of detection built on static rules. The company also runs an offensive side, automated red teaming against a customer's own application.
Where it fits
This sits in the request path in production, added by the team that owns the GenAI application, usually with security setting the policy thresholds. It runs as a network call, so you trade a few tens of milliseconds and a data flow for detection you do not maintain. Settle two things first: what happens on a positive verdict, and whether prompt content may leave your environment.
Strengths
- Detection models trained on a large corpus of real human attack attempts rather than a curated list of published jailbreaks.
- Simple API surface, so adding it to an existing application is a wrapper rather than an architectural change.
- Covers both directions, catching injected instructions on input and leaked sensitive content on output.
- Latency is low enough to sit in a synchronous user-facing path, which is the practical bar for a guardrail.
Limitations
- It is a classifier, and classifiers can be probed. An attacker with repeated access learns what gets through, and detection quality on genuinely novel techniques is unknowable in advance.
- The hosted model routes prompt content to a third party, which is a data governance conversation in regulated environments.
- Detection is not enforcement. If your application ignores the verdict or handles it badly, the guard has bought you nothing.
Who it suits
Good for product teams shipping customer-facing GenAI who want credible injection detection without staffing a research function for it. Less appropriate where every inference must stay inside your boundary and self-hosting is not available to you, or where the risk you actually have is model supply chain rather than prompt abuse.
Used Lakera Guard? Recommend it under your own name and title.
Recommend this tool