Multimodal foundation models: Generative AI course | Zoonk
39. Multimodal foundation models
Connect text, images, documents, and video frames in shared model architectures. Build visual question-answering and document-understanding tasks while addressing modality alignment and visual hallucination.