Multimodal

Working with more than one kind of input or output — text, images, audio, documents or video — in the same model. Multimodal input lets you hand over the actual artefact (a screenshot, a recording) instead of your description of it. Each modality has its own failure mode: small text inside images, accents and noise in audio, layout in scans.

Learn it in

← All terms