Chuyển đến nội dung chính

Computer Vision with Deep Learning: From CNN to Vision Transformer

Practical course on Computer Vision — from CNN, Object Detection (YOLO), Image Segmentation (SAM) to Vision Transformer and Multimodal AI. Hands-on with PyTorch, Ultralytics YOLO, Hugging Face. Deploy CV model to production with TensorRT and ONNX.

Introducing the Series

Computer Vision with Deep Learning is a hands-on course that helps you build image and video recognition systems — from basic image classification to object detection (YOLO), image segmentation (SAM), and Vision Transformer.

🎯 Why Computer Vision? CV is the AI segment with the most practical applications: self-driving cars, medical (X-ray, MRI), retail (customer counting), security (identification), agriculture (pest detection), manufacturing (error checking)...

What will you learn?

Part 1: CV foundation

  • Lesson 1: What is Computer Vision? Image processing with OpenCV
  • Lesson 2: CNN Deep Dive: ResNet, EfficientNet, MobileNet
  • Lesson 3: Transfer Learning — using the trained model

Part 2: Object Detection

  • Lesson 4: 🔥 YOLO from v3 to v11 — detect everything
  • Lesson 5: Custom YOLO training for a separate dataset
  • Lesson 6: Real-time detection — camera, video stream, tracking

Part 3: Segmentation & Modern CV

  • Lesson 7: Image Segmentation: semantic, instance, panoptic
  • Lesson 8: 🔥 SAM (Segment Anything) — segments do not need training
  • Lesson 9: Stable Diffusion & Image Generation
  • Lesson 10: Vision Transformer (ViT) & CLIP

Part 4: Application & Deployment

  • Lesson 11: OCR & Document Understanding for Vietnamese
  • Lesson 12: Multimodal AI: GPT-4o Vision, Gemini Vision
  • Lesson 13: Edge Deployment: TensorRT, ONNX, Mobile
  • Lesson 14: Capstone: end-to-end CV system

Input required

  • Intermediate Python (NumPy, matplotlib)
  • Basic understanding of Neural Networks (or complete "AI & LLM" series lessons 1-4)
  • Google Colab (free, offers T4 GPU)
  • Local GPU is advantageous but not required

Tools used

Python 3.11+       | Ngôn ngữ chính
PyTorch            | Deep Learning framework
Ultralytics        | YOLOv8/v11
OpenCV             | Image processing
Hugging Face       | ViT, SAM, SegFormer
Roboflow           | Data labeling & management
Google Colab       | Free GPU training
TensorRT / ONNX    | Model optimization
Streamlit          | Demo web app