Training Data

FoundationsCore AI Concepts

Training data is the dataset a machine learning model learns from. It is usually historical text, images or numerical records, and the model works through it to find the patterns and relationships it then encodes in its internal parameters. What is missing from the training data is what the model will be worst at.

In practice

Training data is the raw material that sets the ceiling on accuracy, safety and defensibility, so in diligence treat it as an asset and check the title: where it came from, what rights were granted, and whether customer data was used with consent. A proprietary dataset a competitor cannot obtain is one of the few durable advantages in AI; a scraped one is a legal exposure.

Not sure where your organisation stands?

Take the free AI-readiness diagnostic.

Start the diagnostic