Prompt Injection

Working knowledgeGovernance, Safety and Ethics

Prompt injection is a security vulnerability where an attacker manipulates a language model's input instructions to override system rules, safety filters and operational boundaries. The instruction can arrive as direct user text, or hidden inside untrusted external data the model reads, such as a web page or an attached document. A successful injection lets an unauthorised actor hijack the model's controls, exfiltrate sensitive company data, execute unapproved API actions and reach backend databases.

In practice

Prompt injection is the leading threat vector in technical risk assessments of language model integrations. Ask any vendor connecting a model to your systems what happens when the model reads content an attacker controls, because the exposure grows the moment the model is allowed to act rather than only to answer.

Not sure where your organisation stands?

Take the free AI-readiness diagnostic.

Start the diagnostic