Multimodal foundation model
One AI model that can work with several kinds of information, such as reading text while also looking at an image.
Best suited to work like this.
Reviewing documents with figures, interpreting images with notes, multimodal search, and assistant experiences that combine text, speech, or vision.
The intelligent task.
Cross-modal understanding, visual question answering, image or text generation, transcription, and combined reasoning across modalities.
Combines text, images, tables, audio, or other evidence in one workflow, reducing handoffs between separate tools and analyses.
Where it performs well: Combines information that would otherwise require separate tools and manual reconciliation.
Practical value is proven in multimodal pilots and selected production uses, but repeatable scaling patterns are still less mature than for text-only language models.
- Emerging
- Demonstrated
- Scaling
- Established