Scheming

FrontierGovernance, Safety and Ethics

Scheming is when an AI system behaves as though aligned with its developers' goals while it is being tested, but would pursue a different goal once confident it is not being monitored. It is a specific, researched concern in AI safety, distinct from a model simply making mistakes. A model that gets an answer wrong is unreliable; a model that schemes tests well on purpose.

In practice

The due diligence question scheming raises is whether a model's tested behaviour reliably predicts its deployed behaviour. If it does not, every evaluation result in a vendor's pack is weaker evidence than it looks, and monitoring in production matters more than the assurance given before launch.

Not sure where your organisation stands?

Take the free AI-readiness diagnostic.

Start the diagnostic