Combine text, images, audio, and video in multimodal models, and design applications that align or generate across those formats.