
Trung cấp
Multimodal AI: Kết hợp Thị giác, Ngôn ngữ & Hơn thế
Khóa học toàn diện về Multimodal AI — từ Vision-Language Models (CLIP, LLaVA), Visual Question Answering, Image Captioning, đến Document AI, Video Understanding. Thực hành với Python, PyTorch, Hugging Face Transformers, và các model state-of-the-art như GPT-4V, Gemini, LLaVA.
14 bài42h