Red Teaming
Working knowledgeGovernance, Safety and Ethics
Also called: AI Red Teaming
Red teaming is adversarial security testing in which specialists deliberately attempt to breach an AI system's guardrails, bypass its alignment restrictions and trigger harmful outputs. By simulating real threat actors, it uncovers jailbreaks, prompt injection routes, data leakage paths and safety failures before the system is exposed to the public.
In practice
Treat a red team exercise as a release gate for anything customer-facing, and ask to see the findings rather than the certificate. The useful diligence question is not whether a vendor red teams, but what the last exercise found, what was fixed, and when it was run again.