Classic red teaming simulates attackers against networks, systems and identities – the findings are usually clearly reproducible technical vulnerabilities. AI red teaming, by contrast, targets the behaviour of the AI system itself: language models respond probabilistically, the same attack may only succeed on the tenth attempt, and a finding is often not a programming error but undesirable model behaviour. Beyond classic security objectives, testing also covers content-related harms, such as the generation of harmful or false content. Equally important is the distinction from benchmarks: static test datasets measure known capabilities, while red teaming deliberately hunts for novel, context-specific failure modes – one of the central lessons of the Microsoft AI Red Team from more than 100 GenAI products tested.
AI Red Teaming for generative AI systems
How to systematically test language models and AI applications for jailbreaks, prompt injection, data leakage and unsafe tool use – from scenario definition to a repeatable testing process.
As generative AI enters business processes, it creates an attack surface that conventional security testing does not cover: the behaviour of the model itself and its integration into applications. AI red teaming transfers the idea of the simulated attack to this layer – from jailbreaks and direct and indirect prompt injection to the abuse of connected tools. Since 2025, robust methodological anchors have been available, most notably the OWASP GenAI Red Teaming Guide and the NIST taxonomy for adversarial machine learning. At the same time, the AI Act makes adversarial testing an explicit legal obligation for certain models and requires high-risk AI systems to be resilient against AI-specific attacks.
2025: from methodological anchor to legal obligation
Three anchors from 2025 — tap a milestone for details.
OWASP GenAI Red Teaming Guide 1.0
The OWASP GenAI Security Project publishes a structured, vendor-neutral methodology with four testing layers: model, implementation, infrastructure and runtime behaviour.
NIST AI 100-2 E2025
The NIST taxonomy for adversarial machine learning classifies attacks on predictive and generative AI systems by the attacker's objectives, capabilities and lifecycle phase.
AI Act adversarial testing obligation applies
Under Art. 55(1)(a), providers of general-purpose AI models with systemic risk must conduct and document adversarial testing to identify and mitigate systemic risks.
The Essentials at a Glance
Six topic blocks — tap to expand.
Methodological anchors: OWASP and NIST
Two documents, one vocabulary — defining test scope and assessment criteria in a traceable way.
- A structured, vendor-neutral methodology published by the OWASP GenAI Security Project.
- Four testing layers: evaluation of the model itself, the implementation (including guardrails and system prompts), the surrounding infrastructure, and runtime behaviour in interaction with users and processes.
- A taxonomy and terminology for adversarial machine learning.
- Classifies attacks on predictive and generative AI systems by the attacker's objectives, capabilities and lifecycle phase.
- Together with the OWASP guide, it forms the vocabulary needed to define test scope and assessment criteria in a traceable way.
Standards & Sources
The content on this page is based on the following publicly available guides and studies.
GenAI Red Teaming Guide
Version 1.0 of 22 January 2025; structures AI red teaming across the four testing layers of model, implementation, infrastructure and runtime behaviour.
OWASP Top 10 for LLM Applications 2025
Current risk list for LLM applications with Prompt Injection (LLM01), Sensitive Information Disclosure (LLM02), Excessive Agency (LLM06) and System Prompt Leakage (LLM07) as typical test targets.
Verordnung (EU) 2024/1689 (KI-Verordnung / AI Act)
Art. 55(1)(a) obliges providers of GPAI models with systemic risk to conduct documented adversarial testing (applicable since 02.08.2025); Art. 15(5) requires high-risk AI systems to be resilient against AI-specific attacks.
Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2 E2025)
Taxonomy and terminology for attacks on predictive and generative AI systems (March 2025); provides the shared vocabulary for assessment criteria and reporting.
Lessons From Red Teaming 100 Generative AI Products
Eight practical lessons from more than 100 GenAI red team engagements, including the distinction from benchmarks and the interplay of automation and human expertise.
How resilient is your AI application?
In a no-obligation initial consultation, we clarify which attack surfaces your AI applications expose and which testing approach – one-off or continuous – fits your deployment scenario.