AppSecNews
AI Security Open source Growing

Garak

by NVIDIA

Command line LLM vulnerability scanner that fires a library of attack probes at a model endpoint and scores the responses with matched detectors.

Visit github.com (leaves AppSecNews, opens in a new tab) Leaves AppSecNews for the vendor's own site.

No endorsements yet

Run Garak in production? A named recommendation helps the next team shortlisting it.

Recommend this tool

Endorsers verify their identity through LinkedIn. Titles and companies are self declared, shown as they were when each person signed, and reviewed by an editor before anything is published. Endorsements are never paid for.

What we still need to verify : 1 point in this profile is not yet confirmed against vendor documentation.
  • Current generator and probe inventory changes often, confirm against the repository

Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.

What it does

Garak works like a fuzzer for language models. You point it at a generator, an OpenAI-compatible endpoint, a locally loaded Hugging Face model, an Ollama instance or a plain REST target, and it runs a catalogue of probes against it. Each probe is a family of adversarial prompts with a purpose: coax out malware, replay memorized training data, follow an instruction hidden in retrieved text, emit toxic content, or hallucinate a package name that an attacker could then register.

Every probe pairs with one or more detectors that judge the response. Detectors range from cheap string and regex matching against known compliance markers to classifier models that score toxicity or decide whether a refusal was bypassed. A run produces probe-by-probe hit rates plus a log holding the exact prompts and completions that succeeded, which is what you actually hand to the model owner.

Where it fits

This is pre-deployment testing and scheduled regression work, not inline protection. Security engineers and ML teams run it from a laptop or a CI job against a staging endpoint whenever the model, the system prompt or the guardrail configuration changes. You need a reachable endpoint and a real token budget, because a full suite issues a large number of generations.

Strengths

  • The probe and detector plugin structure means adding a house-specific attack is writing one class, not forking the tool.
  • Broad generator support lets the same suite run against a hosted API, a local weights file and an internal wrapper service, so results are comparable.
  • The hit log gives reproducible evidence: a finding arrives as a prompt someone can paste, not just a score.
  • Coverage reaches past jailbreaks into encoding tricks, data leakage, package hallucination and glitch token behavior.

Limitations

  • Results are nondeterministic. The same suite against the same model can move several points between runs, so small deltas carry no signal.
  • Detectors are heuristics. String matching over-reports and classifiers carry their own bias, so triaging the hit log stays manual.
  • A published probe corpus is a known corpus. A model or filter can be tuned to pass it, which makes a clean report weaker evidence than it appears.

Who it suits

A good fit for teams that own or integrate LLMs and want a repeatable, fully self-hosted baseline without routing prompts to a vendor, and for red teams that want a harness to hang custom probes on. It suits you less well if you want inline blocking, a managed service or a report you can put in front of an auditor, since Garak produces raw findings and leaves interpretation and remediation entirely with you.

Used Garak? Recommend it under your own name and title.

Recommend this tool