AppSecNews
AI Security Open source Growing

FuzzyAI

by CyberArk

Open source fuzzer that applies a catalog of published jailbreak and prompt injection techniques against local or hosted language model endpoints.

Visit github.com (leaves AppSecNews, opens in a new tab) Leaves AppSecNews for the vendor's own site.

No endorsements yet

Run FuzzyAI in production? A named recommendation helps the next team shortlisting it.

Recommend this tool

Endorsers verify their identity through LinkedIn. Titles and companies are self declared, shown as they were when each person signed, and reviewed by an editor before anything is published. Endorsements are never paid for.

What we still need to verify : 2 points in this profile are not yet confirmed against vendor documentation.
  • Complete attack method list and provider coverage: verify against project docs
  • Classifier options used to judge attack success: confirm

Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.

What it does

FuzzyAI is a command line fuzzer for language models. You give it a target endpoint, prompts you expect the model to refuse, and a list of attack methods. It then transforms your base prompt according to each method, sends the result, and classifies the response to decide whether the model complied. The value is in the catalog: the project collects jailbreak techniques from published research and implements them behind a common interface, so running twenty documented attacks is a single command rather than twenty reimplementations.

The techniques span several families. Persona and instruction override attacks try to displace the system prompt. Encoding and obfuscation attacks hide intent behind character substitution, alternative alphabets, or rendered text a vision model reads but a text filter does not. Conversational attacks escalate across turns, building context until a request that would be refused cold is accepted. Search driven attacks treat jailbreaking as optimization, mutating prompts across generations and keeping whatever scores better. Success judgment is delegated to configurable classifiers: a keyword heuristic, or a separate model acting as judge. Providers are pluggable and local serving is supported, which matters when test payloads should not leave your network.

Where it fits

This is a model and guardrail evaluation tool, operated by a security engineer or a red teamer rather than by the application team. Natural uses are comparing candidate models before selection, measuring whether a new input filter changes outcomes, and producing evidence for a model risk review. Results need interpretation, so it is not a pass or fail CI gate the way a dependency scanner is. You need endpoint access, credentials, and tolerance for the generation cost of a broad sweep.

Strengths

  • Consolidates published attack techniques into one runnable catalog with a consistent interface.
  • Provider abstraction plus local model support lets you test sensitive prompts without sending them to a third party.
  • Genetic and multi turn strategies find failures that single shot prompt lists cannot reach.
  • Attacks and classifiers are separable, so a stricter judge swaps in without touching attack code.

Limitations

  • It tests a model endpoint. It does not by itself exercise your retrieval layer, tool wiring, or agent loop, which is where most real world impact lives.
  • Classifier accuracy caps result quality. Keyword judging over reports refusals as successes, and a model judge adds cost and its own variance.
  • Published jailbreaks age quickly as providers patch them, so a clean run means less than it appears.

Who it suits

Security teams doing model selection or guardrail validation who want a broad technique library they can run themselves. Not the right tool for an application team wanting end to end assurance of a deployed assistant, which needs a harness that drives the full application.

Used FuzzyAI? Recommend it under your own name and title.

Recommend this tool