Benchmark
An AI benchmark is a standardised dataset, test suite or performance metric used to compare and rank artificial intelligence models on a specific capability, such as logical reasoning, language understanding or coding accuracy. Testing models against independent benchmarks lets a buyer compare vendors on identical terms, track performance over time and choose between architectures on evidence rather than assertion.
In practice
Benchmarks are the standard tool for comparing rival foundation models during vendor selection, but only where the benchmark resembles the work you need done. A high general reasoning score says little about extracting the right fields from your contracts. Build a small internal test set from real cases and score every vendor on that as well.