Home Glossary Multimodal AI

Multimodal AI - Page 9

Multimodal AI processes or generates more than one type of information, such as text, images, audio, video, sensor readings, or actions. A system might answer questions about a chart, create an image from a description, transcribe speech while recognizing speakers, or guide a robot using visual and language input. Models align different modalities in shared representations so information from one form can influence another. This creates flexible interfaces but also introduces new failure modes: visual content may contradict text, timing may be misaligned, and harmful information can cross between modalities. Evaluation must cover each input type, their combinations, accessibility needs, permissions, and the real conditions in which the system operates.

Andrew Tulloch Leaves $12B AI Startup to Join Meta After Turning Down $1.5B Offer
By • 3 mins read
AI & Machine Learning, Immersive Reality (AR, VR, MR, and XR), News, Startups & Investment

Andrew Tulloch Leaves $12B AI Startup to Join Meta After Turning Down $1.5B Offer

By • 3 mins read

Andrew Tulloch, co-founder of the $12 billion AI startup Thinking Machines Lab, has joined Meta after previously rejecting what reports described as a $1.5 billion offer — a figure Meta has since called ‘inaccurate and ridiculous.’