Multimodal deep learning: Deep Learning course | Zoonk
46. Multimodal deep learning
Combine text, images, audio, video, and sensor data through modality-specific encoders, shared representations, cross-attention, and alignment objectives. Design and evaluate systems that reason or generate across more than one modality.