Fundamentals

Multimodal AI

AI that can process and generate multiple types of data — text, images, audio, video together. GPT-4V and Gemini Pro Vision are multimodal. Especially useful where voice, images, or mixed interfaces help people use AI without relying only on typed English.

Use this term as a quick reference, then jump back into the part of the site where it matters.

Explore next