Book an Appointment

AI Red Teaming for generative AI systems

How to systematically test language models and AI applications for jailbreaks, prompt injection, data leakage and unsafe tool use – from scenario definition to a repeatable testing process.

Last updated: July 2026 · Valeri Milke, ISO 27001 & ISO 42001 Lead Auditor

4families of techniques in focus: jailbreaks, prompt injection, data extraction, unsafe tool use
1.0version of the OWASP GenAI Red Teaming Guide with four testing layers (January 2025)
100+GenAI products tested by the Microsoft AI Red Team (more than 100)
2025since 2 August: adversarial testing obligation for models with systemic risk (Art. 55(1)(a) AI Act)

As generative AI enters business processes, it creates an attack surface that conventional security testing does not cover: the behaviour of the model itself and its integration into applications. AI red teaming transfers the idea of the simulated attack to this layer – from jailbreaks and direct and indirect prompt injection to the abuse of connected tools. Since 2025, robust methodological anchors have been available, most notably the OWASP GenAI Red Teaming Guide and the NIST taxonomy for adversarial machine learning. At the same time, the AI Act makes adversarial testing an explicit legal obligation for certain models and requires high-risk AI systems to be resilient against AI-specific attacks.

2025: from methodological anchor to legal obligation

Three anchors from 2025 — tap a milestone for details.

The Essentials at a Glance

Six topic blocks — tap to expand.

Methodological anchors: OWASP and NIST

Two documents, one vocabulary — defining test scope and assessment criteria in a traceable way.

Version 1.0 · January 2025
  • A structured, vendor-neutral methodology published by the OWASP GenAI Security Project.
  • Four testing layers: evaluation of the model itself, the implementation (including guardrails and system prompts), the surrounding infrastructure, and runtime behaviour in interaction with users and processes.
modelimplementationguardrailssystem promptsinfrastructureruntime behaviour

Standards & Sources

The content on this page is based on the following publicly available guides and studies.

OWASP GenAI Security Project · 2025

GenAI Red Teaming Guide

Version 1.0 of 22 January 2025; structures AI red teaming across the four testing layers of model, implementation, infrastructure and runtime behaviour.

OWASP GenAI Security Project · 2025

OWASP Top 10 for LLM Applications 2025

Current risk list for LLM applications with Prompt Injection (LLM01), Sensitive Information Disclosure (LLM02), Excessive Agency (LLM06) and System Prompt Leakage (LLM07) as typical test targets.

Amtsblatt der EU / EUR-Lex · 2024

Verordnung (EU) 2024/1689 (KI-Verordnung / AI Act)

Art. 55(1)(a) obliges providers of GPAI models with systemic risk to conduct documented adversarial testing (applicable since 02.08.2025); Art. 15(5) requires high-risk AI systems to be resilient against AI-specific attacks.

NIST · 2025

Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2 E2025)

Taxonomy and terminology for attacks on predictive and generative AI systems (March 2025); provides the shared vocabulary for assessment criteria and reporting.

Microsoft AI Red Team · 2025

Lessons From Red Teaming 100 Generative AI Products

Eight practical lessons from more than 100 GenAI red team engagements, including the distinction from benchmarks and the interplay of automation and human expertise.

How resilient is your AI application?

In a no-obligation initial consultation, we clarify which attack surfaces your AI applications expose and which testing approach – one-off or continuous – fits your deployment scenario.