Multimodal AI

Working knowledgeCore AI Concepts

Also called: Multi-modal AI

Multimodal AI is a model built from the start to take in, cross-reference and generate several distinct data types at the same time, including text, images, audio, video and numerical tables. Handling more than one kind of input at once is closer to the way people take in a situation, and it lets software read a customer service video call alongside that customer's written account history rather than treating the two as separate problems.

In practice

Multimodal capability is what lets one platform replace separate visual, text and audio analytics tools, which is usually where the cost case sits. Before consolidating, check that the single model is genuinely good at each data type rather than strong at one and adequate at the rest.

Not sure where your organisation stands?

Take the free AI-readiness diagnostic.

Start the diagnostic