Multimodal AI
A model that handles more than one kind of input or output — text, images, audio, video.
Multimodale Modelle bilden verschiedene Datentypen auf eine gemeinsame Repräsentation ab, sodass du einem Modell einen Screenshot und eine Frage geben und eine Textantwort bekommen kannst. Praktisch heißt das: Whiteboards, PDFs und Sprache werden zu gültigen Eingaben, was in vielen Abläufen den Schritt des Transkribierens überflüssig macht.
In der Praxis: Ein kaputtes Diagramm in einem Foliendeck abfotografieren und fragen, was daran nicht stimmt.
Where this comes up
- 18 Best AI Tools for Product Managers in 2026: Review and Comparison of Pros & Cons, Key Features and Pricing
- Best AI Tools for Business in 2026: The Complete Guide
- Best AI Tools for Ecommerce 2026: Shopify, Support & Growth
- DeepSeek V4.1 Flash Replaces V4 Pro: Pricing, Benchmarks, and What Changed Since V4 Flash
- Gemini Alternatives in 2026: How to Compare ChatGPT, Claude, Microsoft Copilot, and Perplexity
- Gemini in Google Slides: What It Can Do and How to Run a Governed Pilot