
Intermediate
Multimodal AI: Combining Vision, Language & More
Comprehensive course on Multimodal AI — from Vision-Language Models (CLIP, LLaVA), Visual Question Answering, Image Captioning, to Document AI, Video Understanding. Practice with Python, PyTorch, Hugging Face Transformers, and state-of-the-art models like GPT-4V, Gemini, LLaVA.
14 lessons42h