Glossary · 1 · Foundations: how AI works
Multimodal model
Also known as: Multimodal AI
A multimodal model is an AI model that can process or generate more than one type of data — for example text and images, or text, audio and video — within the same model.
- Beginner
- Technical writers
- Technical marketers
- Technical project managers
In one sentence
Multimodal AI models defined: one model that reads and creates text, images, audio or video — useful for screenshots, diagrams and visual docs.
Example
A service technician photographs an error display; a multimodal assistant reads the photo, identifies the error code and points to the matching troubleshooting topic.
Why it matters on your learning path
- Technical writers: Multimodal models can describe screenshots, draft alt text and check whether images match the text — always with review.
- Technical marketers: Image and video generation (see Midjourney and Adobe Firefly) are multimodal use cases with their own rights questions.
- Technical project managers: Multimodal inputs increase token use and cost; check how the vendor bills images and audio.