Introducing the Series
Computer Vision with Deep Learning is a hands-on course that helps you build image and video recognition systems — from basic image classification to object detection (YOLO), image segmentation (SAM), and Vision Transformer.
🎯 Why Computer Vision? CV is the AI segment with the most practical applications: self-driving cars, medical (X-ray, MRI), retail (customer counting), security (identification), agriculture (pest detection), manufacturing (error checking)...
What will you learn?
Part 1: CV foundation
- Lesson 1: What is Computer Vision? Image processing with OpenCV
- Lesson 2: CNN Deep Dive: ResNet, EfficientNet, MobileNet
- Lesson 3: Transfer Learning — using the trained model
Part 2: Object Detection
- Lesson 4: 🔥 YOLO from v3 to v11 — detect everything
- Lesson 5: Custom YOLO training for a separate dataset
- Lesson 6: Real-time detection — camera, video stream, tracking
Part 3: Segmentation & Modern CV
- Lesson 7: Image Segmentation: semantic, instance, panoptic
- Lesson 8: 🔥 SAM (Segment Anything) — segments do not need training
- Lesson 9: Stable Diffusion & Image Generation
- Lesson 10: Vision Transformer (ViT) & CLIP
Part 4: Application & Deployment
- Lesson 11: OCR & Document Understanding for Vietnamese
- Lesson 12: Multimodal AI: GPT-4o Vision, Gemini Vision
- Lesson 13: Edge Deployment: TensorRT, ONNX, Mobile
- Lesson 14: Capstone: end-to-end CV system
Input required
- Intermediate Python (NumPy, matplotlib)
- Basic understanding of Neural Networks (or complete "AI & LLM" series lessons 1-4)
- Google Colab (free, offers T4 GPU)
- Local GPU is advantageous but not required
Tools used
Python 3.11+ | Ngôn ngữ chính
PyTorch | Deep Learning framework
Ultralytics | YOLOv8/v11
OpenCV | Image processing
Hugging Face | ViT, SAM, SegFormer
Roboflow | Data labeling & management
Google Colab | Free GPU training
TensorRT / ONNX | Model optimization
Streamlit | Demo web app