Jailbreaking

Working knowledgeGovernance, Safety and Ethics

Also called: Jailbreak

Jailbreaking is the practice of using crafted prompts to bypass an AI model's built-in safety filters, ethical guardrails and operational alignment restrictions. A successful attack forces the model to generate prohibited, toxic, confidential or dangerous output. Defending against it takes persistent testing, system prompts and adversarial red-teaming, not a one-off configuration.

In practice

Any customer-facing conversational interface will be jailbroken by someone, so the useful question in a cyber risk assessment is what the model can reach when it happens, not whether the filters will hold. A jailbroken chatbot that can only talk is embarrassing; one wired into a payments or records system is a breach.

Not sure where your organisation stands?

Take the free AI-readiness diagnostic.

Start the diagnostic