Combine text, images, audio, and video using shared representations, cross-modal attention, multimodal prompting, and modality-specific evaluation.