Distillation

FrontierTechniques and Architectures

Also called: Knowledge Distillation, Model Distillation

Distillation is a training technique where a smaller model learns to imitate a larger one, typically by training on the larger model's outputs rather than on raw data. It is a legitimate and common way to build efficient models cheaply. It became a geopolitical flashpoint when US labs accused Chinese developers of distilling frontier US models to build competitive open-weight alternatives, a practice export controls were designed to prevent but cannot technically stop.

In practice

A distilled model can look like a bargain on a pricing sheet, but it inherits the blind spots of the model it was trained on, and its provenance is rarely disclosed. Ask before adopting one: what was this trained on, and does anyone actually know?

Not sure where your organisation stands?

Take the free AI-readiness diagnostic.

Start the diagnostic