Menu
magesh.ai agent v1.0 (views are my own)
kill-chain resources about · viewing: red_team · 00:00:00
← agent.navigate: resources / assessment & red teaming
25 min read · 6 tools · 3 methodologies · 14 references

Agent Red Team Framework

Traditional pen testing doesn't cover agent-specific attack vectors. You can't nmap an agent's reasoning chain. You need tools built for this — and a methodology that tests each kill chain stage systematically.

category:
Assessment & Red Teaming · security-teams
CONTEXT Red teaming tests each Kill Chain stage — read the threat model first →

The Numbers

Historical results from AgentDojo v1 (June 2024). The cards measure different scopes and must not be read as a before/after comparison. They are published results, not experiments reproduced on this site. Original paper, sections 4.1 and 4.3.

92%
attack success rate on Slack agent suite
GPT-4o · Slack suite · v1 Figure 7
7.5%
attack success rate WITH tool-filtering defense
GPT-4o · aggregate suite · v1 section 4.3
<66%
benign task utility for the best model evaluated in v1
Historical model set; not a current capability ceiling
Scope of the finding: In this experiment, more capable models sometimes completed attacker goals more successfully. That observation depends on models, tasks and attack definitions; it is not a universal inverse-scaling law. Evaluate benign utility and attack success together.

Six Tools

Each serves a different purpose. Use them together for coverage.

⬡ AgentDojo — Agent Security Benchmark

The original AgentDojo release includes 97 user tasks and 629 security cases across workspace, Slack, banking and travel suites. It evaluates indirect injection in tool outputs and measures both user-task utility and attacker-goal success. Record the benchmark commit and task subset for any new evaluation.

Use AgentDojo to evaluate indirect-injection resilience. The paper’s 92% Slack result and 7.5% aggregate filtered result have different denominators. Compare matched model, task-suite, attack and defense conditions; report repeated trials and confidence intervals for your own runs.
Source: Debenedetti et al., ETH Zurich. github.com/ethz-spylab/agentdojo
⬡ PyRIT — Automated Red Teaming

Microsoft’s open-source framework for automated red teaming. Reusable attack strategies including single-turn, multi-turn, Crescendo (gradual escalation), and Tree of Attacks with Pruning (TAP). Now integrated into Azure AI Foundry as the "AI Red Teaming Agent."

Use for: Automated, multi-turn attack campaigns against your agent. Best for Stage 3 HIJACK testing — can the agent's behavior be redirected through conversation?
Source: Microsoft. github.com/Azure/PyRIT
⬡ Garak — LLM Vulnerability Scanner

Garak connects model targets to probes and response detectors, then reports results. It can test prompt injection, harmful outputs and leakage patterns. Its model-level results do not by themselves test the authorization behavior of a complete tool-using deployment.

Use for: CI/CD integration. Run Garak on every model update to catch regressions. Best for broad vulnerability scanning across Stages 1-2.
Source: NVIDIA. github.com/NVIDIA/garak
⬡ CyberSecEval — Progressive Security Benchmarks

Meta’s PurpleLlama includes CyberSecEval suites assessing different security capabilities and risks. Offensive-operation, SOC and patching evaluations ask different questions from prompt-injection resistance. Pin the suite and task version, and describe exactly what its success criteria measure.

Use capability benchmarks to characterize what a model or agent can do. To establish hijacking or authorization failure, separately test attacker-controlled inputs against the complete tool and permission setup.
Source: Meta AI. github.com/meta-llama/PurpleLlama
⬡ Prompt Guard 2 — Injection Detection

Meta's lightweight classifier for real-time prompt injection detection. Prompt Guard 2 (86M mDeBERTa-base / 22M DeBERTa-xsmall) uses binary classification (benign vs. malicious). The original Prompt Guard v1 used three classes (benign, injection, jailbreak). Deployable on CPU for real-time filtering. Fine-tunable to your data.

Important caveat

Meta documents that adaptive attacks may bypass Prompt Guard and that application-specific inputs affect detection. Treat the classifier as one defense layer, and evaluate it against your own threat model.

⬡ HarmBench — Attack vs Defense Comparison

Center for AI Safety benchmark. The original 2024 study compares 18 red teaming methods tested against 33 target LLMs and defenses. Performance varies by target, attack and defense; model size alone does not establish robustness. These are historical results for the evaluated systems.

Use for: Selecting which attack methods to use against your specific model/defense combination. HarmBench data tells you which attacks are most effective against which defenses.
Source: Center for AI Safety. harmbench.org

Three Methodologies

01
CSA Agentic AI Red Teaming Guide

The CSA Agentic AI Red Teaming Guide (May 2025) covers agent-focused threats including authorization, goals, memory and orchestration. Use its guidance to scope tests, execute controlled attacks, analyze evidence and report remediation; this summary does not assert an exact numbered lifecycle from the guide.

Use the guide alongside OWASP and a deployment-specific threat model. A published methodology still needs adaptation to your data, permissions and operational consequences.

02
OWASP Top 10 for Agentic Applications

Released December 2025, 100+ contributors. Ten agent-specific risks: ASI01 (Agent Goal Hijack) through ASI10 (Rogue Agents). Covers identity, tools, delegated trust boundaries, and autonomous operation risks. Use this as your risk checklist — each ASI maps to specific test cases.

03
Evidence-First Auditing (from my practice)

My assessment practice is to attach reproducible evidence to each finding: affected component, test conditions, observable outcome, impact, proposed fix and retest. There is no minimum number of findings. Record successful controls and negative results too. Prioritize by demonstrated impact, exploitability and remaining uncertainty; an architecture review may identify issues that individual file checks miss.

This isn't a published standard — it's how I run security assessments. The point is that red teaming without evidence is just a conversation.

Testing Each Stage

Map each kill chain stage to the right tool and test.

Kill Chain StageWhat to TestTool
01 RECONCan the agent's tools, permissions, and system prompt be extracted?Manual probing + Garak
02 INJECTCan indirect injection via tool responses change agent behavior?AgentDojo + PyRIT
03 HIJACKCan the agent's goal be substituted through multi-turn conversation?PyRIT + task-specific end-to-end injection tests
04 ESCALATECan the agent access tools beyond its intended scope?Manual + hook bypass testing
05 EXFILCan the agent leak data through legitimate channels?AgentDojo + manual output review
06 PERSISTCan memory or config files be poisoned for future sessions?Manual + MCP security checks

Red teaming is how you validate the kill chain's defensive controls. Combine it with hook-based guardrails (prevention) and MCP security (tool-layer defense) for defense in depth.

This work represents the author's independent research and personal views. It is not related to or endorsed by the author's employer.