Agent Red Team Framework
Traditional pen testing doesn't cover agent-specific attack vectors. You can't nmap an agent's reasoning chain. You need tools built for this — and a methodology that tests each kill chain stage systematically.
Assessment & Red Teaming · security-teams The Numbers
Historical results from AgentDojo v1 (June 2024). The cards measure different scopes and must not be read as a before/after comparison. They are published results, not experiments reproduced on this site. Original paper, sections 4.1 and 4.3.
Six Tools
Each serves a different purpose. Use them together for coverage.
The original AgentDojo release includes 97 user tasks and 629 security cases across workspace, Slack, banking and travel suites. It evaluates indirect injection in tool outputs and measures both user-task utility and attacker-goal success. Record the benchmark commit and task subset for any new evaluation.
Microsoft’s open-source framework for automated red teaming. Reusable attack strategies including single-turn, multi-turn, Crescendo (gradual escalation), and Tree of Attacks with Pruning (TAP). Now integrated into Azure AI Foundry as the "AI Red Teaming Agent."
Garak connects model targets to probes and response detectors, then reports results. It can test prompt injection, harmful outputs and leakage patterns. Its model-level results do not by themselves test the authorization behavior of a complete tool-using deployment.
Meta’s PurpleLlama includes CyberSecEval suites assessing different security capabilities and risks. Offensive-operation, SOC and patching evaluations ask different questions from prompt-injection resistance. Pin the suite and task version, and describe exactly what its success criteria measure.
Meta's lightweight classifier for real-time prompt injection detection. Prompt Guard 2 (86M mDeBERTa-base / 22M DeBERTa-xsmall) uses binary classification (benign vs. malicious). The original Prompt Guard v1 used three classes (benign, injection, jailbreak). Deployable on CPU for real-time filtering. Fine-tunable to your data.
Meta documents that adaptive attacks may bypass Prompt Guard and that application-specific inputs affect detection. Treat the classifier as one defense layer, and evaluate it against your own threat model.
Center for AI Safety benchmark. The original 2024 study compares 18 red teaming methods tested against 33 target LLMs and defenses. Performance varies by target, attack and defense; model size alone does not establish robustness. These are historical results for the evaluated systems.
Three Methodologies
The CSA Agentic AI Red Teaming Guide (May 2025) covers agent-focused threats including authorization, goals, memory and orchestration. Use its guidance to scope tests, execute controlled attacks, analyze evidence and report remediation; this summary does not assert an exact numbered lifecycle from the guide.
Use the guide alongside OWASP and a deployment-specific threat model. A published methodology still needs adaptation to your data, permissions and operational consequences.
Released December 2025, 100+ contributors. Ten agent-specific risks: ASI01 (Agent Goal Hijack) through ASI10 (Rogue Agents). Covers identity, tools, delegated trust boundaries, and autonomous operation risks. Use this as your risk checklist — each ASI maps to specific test cases.
My assessment practice is to attach reproducible evidence to each finding: affected component, test conditions, observable outcome, impact, proposed fix and retest. There is no minimum number of findings. Record successful controls and negative results too. Prioritize by demonstrated impact, exploitability and remaining uncertainty; an architecture review may identify issues that individual file checks miss.
This isn't a published standard — it's how I run security assessments. The point is that red teaming without evidence is just a conversation.
Testing Each Stage
Map each kill chain stage to the right tool and test.
| Kill Chain Stage | What to Test | Tool |
|---|---|---|
| 01 RECON | Can the agent's tools, permissions, and system prompt be extracted? | Manual probing + Garak |
| 02 INJECT | Can indirect injection via tool responses change agent behavior? | AgentDojo + PyRIT |
| 03 HIJACK | Can the agent's goal be substituted through multi-turn conversation? | PyRIT + task-specific end-to-end injection tests |
| 04 ESCALATE | Can the agent access tools beyond its intended scope? | Manual + hook bypass testing |
| 05 EXFIL | Can the agent leak data through legitimate channels? | AgentDojo + manual output review |
| 06 PERSIST | Can memory or config files be poisoned for future sessions? | Manual + MCP security checks |
Red teaming is how you validate the kill chain's defensive controls. Combine it with hook-based guardrails (prevention) and MCP security (tool-layer defense) for defense in depth.
Detection patterns, governance guides, and more practitioner content coming.
This work represents the author's independent research and personal views. It is not related to or endorsed by the author's employer.