Multimodal AI
A model that handles more than one kind of input or output — text, images, audio, video.
Les modèles multimodaux projettent différents types de données dans une représentation commune : vous pouvez donc en donner un une capture d’écran accompagnée d’une question et obtenir une réponse en texte. Concrètement, tableaux blancs, PDF et voix deviennent des entrées valides, ce qui supprime l’étape de transcription dans quantité de flux de travail.
En pratique : Photographier un graphique bancal dans une présentation et demander ce qui cloche.
Where this comes up
- 18 Best AI Tools for Product Managers in 2026: Review and Comparison of Pros & Cons, Key Features and Pricing
- Best AI Tools for Business in 2026: The Complete Guide
- Best AI Tools for Ecommerce 2026: Shopify, Support & Growth
- DeepSeek V4.1 Flash Replaces V4 Pro: Pricing, Benchmarks, and What Changed Since V4 Flash
- Gemini Alternatives in 2026: How to Compare ChatGPT, Claude, Microsoft Copilot, and Perplexity
- Gemini in Google Slides: What It Can Do and How to Run a Governed Pilot