The leading open-source LLM red teaming tools in 2026 are Promptfoo for testing whole applications with OWASP-mapped attack plugins, NVIDIA's garak for scanning models with a large library of probes, Microsoft's PyRIT for scripted, multi-turn attack campaigns, and DeepTeam and Giskard for Python teams that want red teaming next to their evaluations. All are free; the right one depends on whether you are testing a raw model, an application with retrieval and tools, or running a bespoke security engagement.
This guide covers application security testing for LLM products: prompt injection, data leakage, jailbreaks and unsafe tool use. It explains what red teaming is, compares the tools, and gives a first command for each.
What is red teaming in AI?
Red teaming in AI means deliberately attacking your own model or application to find failures before real users or attackers do. For an LLM product that typically includes trying to override the system prompt, extract private data, make the model produce harmful content, trick an agent into calling tools it should not, and exhaust budgets with runaway requests.
Automated tools generate thousands of adversarial inputs, send them to your target, and use detectors or LLM judges to decide whether each attack succeeded. Manual red teaming by people is still valuable for creative, context-specific attacks, but automated tools give you coverage and repeatability, and they can run in CI on every prompt or model change.
The most common checklist is the OWASP Top 10 for LLM Applications 2025: prompt injection; sensitive information disclosure; supply chain; data and model poisoning; improper output handling; excessive agency; system prompt leakage; vector and embedding weaknesses; misinformation; and unbounded consumption. The first of these gets its own deep dive in our guide to prompt injection attacks.
LLM red teaming tools compared
| Tool | Maintainer | License | Language | Best at |
|---|---|---|---|---|
| Promptfoo | Promptfoo, now part of OpenAI | MIT | Node.js CLI, YAML config | Application-level testing with OWASP, NIST and EU AI Act presets |
| garak | NVIDIA | Apache 2.0 | Python CLI | Scanning models with many built-in probes |
| PyRIT | Microsoft | MIT | Python library | Custom, multi-turn attack orchestration |
| DeepTeam | Confident AI | Apache 2.0 | Python library | 50+ vulnerability types, built on DeepEval |
| Giskard | Giskard | Apache 2.0 | Python library | Scans plus RAG and agent tests in one package |
To summarize the table: Promptfoo and garak are command-line tools you can run in minutes, Promptfoo aimed at applications and garak at models; PyRIT is a toolkit for security specialists who want to script their own campaigns; and DeepTeam and Giskard fit teams that already write Python evaluation tests.
The tools in detail
Promptfoo
Promptfoo started as a prompt testing CLI and added a full red teaming module. You describe your target, which can be an HTTP endpoint, a model or a custom script, plus what the application is for, and Promptfoo generates attacks tailored to that purpose. Plugins cover harm categories such as personal data leakage, prompt injection and excessive agency, and strategies such as jailbreaks and multi-turn attacks change how those inputs are delivered. Presets map directly to frameworks, so adding the owasp:llm plugin to your config tests against the OWASP list (OWASP guide).
First run, from the quickstart:
npx promptfoo@latest redteam setup
npx promptfoo@latest redteam run
npx promptfoo@latest redteam report
The setup command opens a browser-based configurator, run executes the attacks, and report opens the results. In March 2026 Promptfoo agreed to join OpenAI and said the project remains open source under the MIT license (announcement).
garak
garak, from NVIDIA, describes itself as an LLM vulnerability scanner, and its creators compare it to network tools like nmap and Metasploit. It ships a large library of probes for jailbreaks, prompt injection, encoding tricks, data leakage, toxicity and hallucination, paired with detectors that judge the responses. It talks to Hugging Face models, OpenAI-compatible APIs, AWS Bedrock, LiteLLM, local llama.cpp models and generic REST endpoints (GitHub).
First run:
python -m pip install -U garak
python3 -m garak --list_probes
python3 -m garak --target_type openai --target_name YOUR_MODEL --spec probes.encoding
garak is best for testing a model or endpoint directly. It is less suited to testing application logic, such as whether your retrieval layer leaks another customer's documents, unless you wrap that application as a REST target.
PyRIT
PyRIT, the Python Risk Identification Tool from Microsoft's AI red team, is a framework rather than a push-button scanner (GitHub). You compose targets, converters that transform prompts, scorers that judge responses, and orchestrated attacks, including multi-turn strategies such as Crescendo, which escalates gradually over a conversation, and Tree of Attacks with Pruning. It is the most flexible option and the one with the steepest learning curve. Install it with pip install pyrit and start from the documentation notebooks.
DeepTeam
DeepTeam, from Confident AI, is an open-source Python framework built on the DeepEval evaluation library. It offers more than 50 vulnerability types, from bias and personal data leakage to SQL injection through an agent, and attack methods including jailbreaking, prompt injection and multi-turn exploitation. It runs locally and can apply framework presets for the OWASP Top 10 for LLMs, NIST AI RMF and MITRE ATLAS (GitHub). Install with pip install -U deepteam.
Giskard
Giskard version 3 is a rewrite focused on testing agents and multi-turn systems. Its scan package probes for security and quality issues, and the same library handles RAG evaluation and custom checks. It requires Python 3.12 or newer; install the scanner with pip install "giskard[scan]" (GitHub). Version 2 is still available but no longer actively maintained, so older tutorials may not match.
Promptfoo vs garak
The Promptfoo vs garak question comes down to what you are testing. Choose garak to probe a model or a model endpoint for known weaknesses with almost no setup; it is like running a vulnerability scanner against a server. Choose Promptfoo to test your application as users see it, including system prompts, retrieval, tools and business rules, because it generates attacks specific to your use case and grades them against your policies. Many teams use both: garak when evaluating a new model, Promptfoo in CI for the product.
How to red team an LLM application
- Define scope and threats. List what the system can access, what tools it can call, and what would be harmful in your context. Map these to the OWASP categories.
- Get permission and isolate. Only test systems you own or are authorized to test, and use staging environments with fake data where possible.
- Start with automated scans. Run Promptfoo against the application and garak against the underlying model to get broad coverage.
- Add targeted and multi-turn attacks. Use PyRIT or Promptfoo's multi-turn strategies for the scenarios that matter most, such as data exfiltration through an agent's tools.
- Triage results. Automated graders produce false positives; have a person confirm the important findings.
- Fix and retest. Tighten prompts, permissions and output handling, add guardrails, then rerun the same tests to confirm.
- Automate in CI. Turn confirmed attacks into regression tests so fixes stay fixed, and log production traffic with a tracing tool such as Langfuse to spot new attack patterns.
Hosted evaluation platforms such as Patronus AI add judge models for catching unsafe or hallucinated outputs, which can complement open-source attack tools when you need managed monitoring.
Pros and cons of automated red teaming
Pros: broad, repeatable coverage; cheap to rerun after every change; framework presets make compliance reporting easier; all the tools above are free and open source.
Cons: LLM-judged results need human confirmation; attack libraries lag behind new techniques; generating and grading thousands of attacks costs model API spend; no tool understands your business risks as well as your own team.
Who each tool is for
Product and platform teams shipping LLM features should start with Promptfoo. Security teams assessing models and endpoints should add garak. Dedicated AI red teams running bespoke engagements will get the most from PyRIT. Python teams already using DeepEval should try DeepTeam, and teams testing agents end to end can look at Giskard. For the quality side of testing, see our guides on LLM evaluation frameworks and how to evaluate LLM apps; for understanding why safety benchmark numbers can mislead, see our LLM benchmarks explainer.
FAQ
What is red teaming in AI?
Red teaming in AI is the practice of deliberately attacking a model or AI application to uncover failures such as prompt injection, data leakage, jailbreaks and unsafe actions, so they can be fixed before real users or attackers find them.
How do you red team an LLM?
Define what the system can access and what harms matter, then run automated tools such as Promptfoo or garak to generate attacks, add targeted multi-turn tests for high-risk scenarios, have people confirm findings, fix them, and keep the attacks as regression tests in CI.
What is the best open-source LLM red teaming tool?
Promptfoo is the most complete choice for testing applications, garak is the quickest way to scan a model, and PyRIT is the most flexible for custom campaigns. All three are free and open source.
What is the OWASP Top 10 for LLM applications?
It is OWASP's list of the most critical security risks for LLM applications. The 2025 edition covers prompt injection, sensitive information disclosure, supply chain, data and model poisoning, improper output handling, excessive agency, system prompt leakage, vector and embedding weaknesses, misinformation and unbounded consumption.
Is Promptfoo still open source after joining OpenAI?
Yes. Promptfoo announced in March 2026 that it was joining OpenAI and said the project remains open source under the MIT license.
What is PyRIT?
PyRIT is Microsoft's open-source Python Risk Identification Tool for generative AI. It is a framework for security professionals to build automated attack campaigns, including multi-turn strategies, against AI systems.