Chuyển đến nội dung chính

Multimodal AI: Combining Vision, Language & More

Comprehensive course on Multimodal AI — from Vision-Language Models (CLIP, LLaVA), Visual Question Answering, Image Captioning, to Document AI, Video Understanding. Practice with Python, PyTorch, Hugging Face Transformers, and state-of-the-art models like GPT-4V, Gemini, LLaVA.