Glossary
What is multimodal?
A model that accepts more than text, typically images, audio or video.
Multimodal models can read a screenshot, a chart or a scanned page. That makes them useful for material that has no selectable text.
Images are expensive in context terms compared with the same information as text, and detail in small type is often lost.
Where a page has real text, sending the text beats sending a screenshot on both accuracy and cost.
In practice
A chart image tells the model the shape of a trend; the underlying table tells it the numbers.
Common questions
- What does multimodal mean in one sentence?
- A model that accepts more than text, typically images, audio or video.
- Why does multimodal matter when you use Claude or ChatGPT?
- Where a page has real text, sending the text beats sending a screenshot on both accuracy and cost.
- Can you give an example of multimodal?
- A chart image tells the model the shape of a trend; the underlying table tells it the numbers.
Related terms
- TokenThe unit an AI model counts text in, usually a few characters long.
- Context windowThe amount of text an AI assistant can hold in mind during a single conversation.
- Readable content extractionPulling the main article text out of a page and discarding the furniture.
- Glossary indexEvery AI context term, defined in plain language.
Get LocalBridge
One click sends your open tabs into Claude or ChatGPT. One-time licence, no subscription.
See how it works