Evals

Working knowledgeGovernance, Safety and Ethics

Also called: AI Evaluation, Evaluations

Evals are systematic, repeatable tests that measure the quality of an AI system's output against a defined set of cases with known good answers. They are the AI counterpart of a software test suite: fixed inputs, an agreed standard for what a correct response looks like, and a score that can be compared run to run. Unlike a public benchmark, which scores a model on tasks everyone is measured against, an eval set is written for one deployment and its own definition of correct.

In practice

A vendor with no eval suite cannot tell you whether the last model change made the product better or worse. It is guessing, and so are you. Ask to see the eval set, who wrote the correct answers, and what the score was either side of the most recent model upgrade.

Not sure where your organisation stands?

Take the free AI-readiness diagnostic.

Start the diagnostic