AppSecNews
AI Security Open source Growing

DeepTeam

by Confident AI

Open source Python framework that generates adversarial prompts against an LLM application and scores the responses for vulnerabilities such as bias, PII leakage and excessive agency.

Visit github.com (leaves AppSecNews, opens in a new tab) Leaves AppSecNews for the vendor's own site.

No endorsements yet

Run DeepTeam in production? A named recommendation helps the next team shortlisting it.

Recommend this tool

Endorsers verify their identity through LinkedIn. Titles and companies are self declared, shown as they were when each person signed, and reviewed by an editor before anything is published. Endorsements are never paid for.

What we still need to verify : 2 points in this profile are not yet confirmed against vendor documentation.
  • Full list of attack and vulnerability modules: verify against project docs
  • Model provider coverage for attack generation and judging: confirm

Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.

What it does

DeepTeam is a red teaming framework for LLM applications, built by the same project as DeepEval and sharing its metric machinery. You give it a callback function that takes a string and returns your application's response. DeepTeam then synthesizes adversarial inputs, sends them through that callback, and grades what comes back. Because the integration point is a function rather than a network endpoint, it tests the whole application including your system prompt, retrieval layer, and tool wiring, not just the raw model.

The framework separates two concerns that are often conflated. Vulnerabilities describe what you are testing for: toxicity, bias, PII leakage, prompt leakage, misinformation, excessive agency, and similar failure classes. Attacks describe how the payload is delivered: direct prompt injection, multi turn jailbreak strategies that escalate across a conversation, and encoding based evasions such as base64, leetspeak, or multilingual rewrites that try to slip past input filters. You compose the two, so the same PII leakage probe can be delivered plainly and then again wrapped in an encoding, which tells you whether a guardrail is matching intent or matching strings. Scoring is done by an LLM judge with a configurable model.

Where it fits

This runs pre-release and on a schedule. The natural home is a nightly or per-release job in CI that exercises a staging deployment of the assistant, with results compared against the previous run. It can also be run ad hoc by an engineer who just changed a system prompt. The prerequisite is a stable, callable interface to the application and a budget for model calls, since both the attacker and the judge are LLMs.

Strengths

  • Tests the deployed application rather than the bare model, so system prompt and guardrail changes show up in results.
  • Clean separation of vulnerability from attack technique, which makes coverage gaps visible instead of implied.
  • Multi turn attack strategies catch escalation failures that single prompt test suites miss entirely.
  • Python native with a small surface, so it drops into an existing test suite without new infrastructure.

Limitations

  • LLM as judge scoring is noisy. Run to run variance is real, and borderline findings need human review before anyone treats them as a defect.
  • Every run costs model calls for generation and grading, so thorough sweeps are slow and expensive to run on every commit.
  • Findings tell you a jailbreak worked, not which control to change. Remediation guidance is thin.

Who it suits

Engineering teams shipping an LLM feature who want repeatable safety and security regression testing they control. Less suitable for a security team that wants a managed service with reporting and an owner to call, or for organizations that cannot send test prompts to a third party model provider.

Used DeepTeam? Recommend it under your own name and title.

Recommend this tool