Combine text, images, audio, video, and sensor data through shared representations and cross-modal attention. Build systems for captioning, visual question answering, speech-language tasks, and multimodal retrieval.