AppSecNews
AI Security Open source Emerging

Rebuff

by Protect AI

Open source prompt injection detector that layers heuristics, a classifier prompt, a vector store of known attacks and canary tokens.

Visit github.com (leaves AppSecNews, opens in a new tab) Leaves AppSecNews for the vendor's own site.

No endorsements yet

Run Rebuff in production? A named recommendation helps the next team shortlisting it.

Recommend this tool

Endorsers verify their identity through LinkedIn. Titles and companies are self declared, shown as they were when each person signed, and reviewed by an editor before anything is published. Endorsements are never paid for.

What we still need to verify : 2 points in this profile are not yet confirmed against vendor documentation.
  • Current maintenance status and activity of the project: confirm
  • Required external dependencies for each detection layer: verify against project docs

Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.

What it does

Rebuff inspects user input before it reaches your prompt and estimates whether it is an injection attempt. It is built as four independent layers because no single technique holds up alone. The first is heuristic: pattern matching for the recognizable shapes of injection, instructions to disregard prior context, attempts to elicit the system prompt, delimiter breakouts. The second sends the input to a language model with a classifier prompt asking whether the text is trying to subvert instructions. The third embeds the input and compares it against a vector store of past attacks, so one that worked before is recognized in variant form. Each layer returns a score and you set the rejection threshold.

The fourth layer works on the output side. Rebuff inserts a canary token into the prompt, a unique marker the model is told to keep private, then checks responses for it. If the token appears in output, the system prompt leaked and you have direct evidence of extraction rather than a guess. Detected attacks can be written back to the vector store, so the deployment learns from what it sees.

Where it fits

In the request path of an LLM application, called before the prompt is assembled and again on the response before it is returned. It is a library, so the application team owns it. Two of the four layers need external services, a model endpoint for the classifier and a vector store for similarity, which adds a dependency to every request you screen.

Strengths

  • Layering independent methods is the right architecture for a problem where any single filter is bypassable.
  • Canary tokens give a high confidence signal of system prompt leakage, which most detectors cannot produce at all.
  • Confirmed attacks feed the similarity store, improving detection against variants of what you have already seen.
  • Small, readable codebase, useful as a reference implementation even if you do not deploy it.

Limitations

  • Heuristic and similarity layers detect known shapes. Novel phrasing, obfuscation and non English payloads get through, and the classifier is itself a model that can be talked out of its judgment.
  • The classifier and embedding calls add latency and inference overhead to every screened request, and the project reads as a reference implementation more than a hardened control.
  • Input screening is one control among several. It does nothing about what an agent may do once an injection succeeds.

Who it suits

Teams that want to understand how injection detection works, or that need a defensible first filter in front of a low risk LLM feature. Not the right foundation for a high risk agentic system, where constraining tool permissions matters more than filtering the input.

Used Rebuff? Recommend it under your own name and title.

Recommend this tool