Introduction
YOLO pretrained detects 80 COCO classes. But do you need to detect defective products, Vietnamese license plates, or specific types of fish? Need to train YOLO custom on separate dataset.
🎯 This article: End-to-end workflow: collect images → label → config → train → evaluate → export.
1. Data Collection
1.1 How many photos is enough?
Rule of thumb cho YOLO custom:
───────────────────────────────────
Minimum: 100 ảnh/class (chỉ OK nếu dùng pretrained)
Tốt: 300-500 ảnh/class
Rất tốt: 1000+ ảnh/class
Production: 3000+ ảnh/class
Lưu ý:
- CHẤT LƯỢNG > Số lượng
- Đa dạng: nhiều góc chụp, ánh sáng, background
- Nếu ít data → dùng pretrained + augmentation mạnh
1.2 Data sources
1. Chụp ảnh thực tế (tốt nhất!)
2. Tải từ Roboflow Universe (roboflow.com/universe)
3. Google Images (cần lọc + verify)
4. Scraping (cẩn thận copyright)
5. Synthetic data (render 3D, paste objects)
1.3 Tips for collecting good photos
✅ ĐA DẠNG:
- Nhiều góc chụp (trước, sau, trên, nghiêng)
- Nhiều điều kiện ánh sáng (ngày, đêm, trong nhà, ngoài trời)
- Nhiều background
- Objects bị che 1 phần (occlusion)
- Objects nhỏ và lớn trong cùng ảnh
❌ TRÁNH:
- Tất cả ảnh cùng 1 góc
- Chỉ ảnh chất lượng studio
- Không có negative samples
- Objects luôn ở center
2. Label Data (Data Annotation)
2.1 Label Tool
| Tools | Type | Free? | Advantages |
|---|---|---|---|
| Roboflow | Web | Free tier | Easiest to use, export YOLO format |
| CVAT | Web/Self-hosted | Free | Strong, video support |
| Label Studio | Self-hosted | Free | Many task types |
| LabelImg | Desktop | Free | Simple, offline |
| Labelme | Desktop | Free | Polygon annotations |
2.2 Label with Roboflow (Recommended)
Workflow Roboflow:
1. Sign up: roboflow.com (free tier: 10K ảnh)
2. Create Project → Object Detection
3. Upload ảnh
4. Label:
- Vẽ bounding box quanh mỗi object
- Gán class name
5. Generate Dataset:
- Train/Val/Test split (70/20/10)
- Augmentation (optional)
6. Export → YOLO format → Download
2.3 YOLO Label Format
# Mỗi ảnh có 1 file .txt cùng tên
# Mỗi dòng: class_id x_center y_center width height (normalized 0-1)
# Ví dụ: labels/image001.txt
0 0.4523 0.3211 0.1250 0.2340
1 0.7100 0.6500 0.0890 0.1560
0 0.2300 0.8100 0.1100 0.1900
# class 0: person, class 1: car
# Tọa độ normalized (chia cho width/height ảnh)
2.4 Quality Label Tips
✅ TIPS:
- Tight bounding box (sát object, không quá rộng hoặc hẹp)
- Nhất quán: cùng 1 object → cùng 1 class
- Đừng quên objects nhỏ hoặc bị che 1 phần
- Label ALL objects of interest (đừng bỏ sót)
- Dùng "review" bước để double-check
❌ SAI LẦM:
- Box quá rộng (chứa nhiều background)
- Bỏ sót objects
- Label sai class
- Inconsistent: object giống nhau nhưng label khác tên
3. Configure Dataset
3.1 Directory structure
datasets/
my_project/
images/
train/
img001.jpg
img002.jpg
...
val/
img050.jpg
...
test/
img070.jpg
...
labels/
train/
img001.txt
img002.txt
...
val/
img050.txt
...
test/
img070.txt
...
3.2 YAML Dataset
# dataset.yaml — file cấu hình cho YOLO
path: /path/to/datasets/my_project # Root directory
train: images/train # Train images (relative to path)
val: images/val # Val images
test: images/test # Test images (optional)
# Classes
names:
0: product_ok # Sản phẩm tốt
1: product_defect # Sản phẩm lỗi
2: product_scratch # Sản phẩm xước
# Số classes
nc: 3
4. Training YOLO Custom
4.1 Basic training
"""Train YOLO custom model"""
from ultralytics import YOLO
# Bắt đầu từ pretrained model (TRANSFER LEARNING!)
model = YOLO("yolo11n.pt") # nano — bắt đầu từ đây
# Train
results = model.train(
data="dataset.yaml", # Path to dataset config
epochs=100, # Số epochs
imgsz=640, # Input image size
batch=16, # Batch size
device="0", # GPU 0 (hoặc "cpu")
project="runs/detect", # Output directory
name="my_product_model", # Run name
)
4.2 Advanced Training — Optimize hyperparameters
"""Training với cấu hình nâng cao"""
results = model.train(
data="dataset.yaml",
epochs=200,
imgsz=640,
batch=16,
# === Optimization ===
optimizer="AdamW", # AdamW thường tốt hơn SGD
lr0=0.01, # Initial learning rate
lrf=0.01, # Final learning rate factor
warmup_epochs=3, # Warmup epochs
weight_decay=0.0005, # Regularization
# === Augmentation ===
hsv_h=0.015, # Hue augmentation
hsv_s=0.7, # Saturation augmentation
hsv_v=0.4, # Value augmentation
degrees=10.0, # Rotation ±10°
translate=0.1, # Translation 10%
scale=0.5, # Scale ±50%
fliplr=0.5, # Horizontal flip 50%
flipud=0.0, # Vertical flip (tắt — không hợp lý cho hầu hết tasks)
mosaic=1.0, # Mosaic augmentation (ghép 4 ảnh)
mixup=0.1, # MixUp augmentation
# === Early Stopping ===
patience=50, # Stop nếu 50 epochs không improve
# === Resuming ===
# resume=True, # Resume from last checkpoint
)
4.3 Training on Google Colab
"""Google Colab training script"""
# Cell 1: Setup
!pip install ultralytics
# Cell 2: Check GPU
import torch
print(f"GPU: {torch.cuda.get_device_name(0)}")
print(f"Memory: {torch.cuda.get_device_properties(0).total_mem / 1e9:.1f} GB")
# Cell 3: Upload dataset
# Option A: Upload ZIP file
from google.colab import files
uploaded = files.upload() # Upload dataset.zip
!unzip dataset.zip -d /content/datasets/
# Option B: Từ Roboflow
!pip install roboflow
from roboflow import Roboflow
rf = Roboflow(api_key="YOUR_KEY")
project = rf.workspace("workspace").project("project")
dataset = project.version(1).download("yolov8")
# Cell 4: Train
from ultralytics import YOLO
model = YOLO("yolo11s.pt")
results = model.train(data="/content/datasets/data.yaml", epochs=100, imgsz=640)
# Cell 5: Evaluate
metrics = model.val()
print(f"mAP@50: {metrics.box.map50:.4f}")
print(f"mAP@50:95: {metrics.box.map:.4f}")
# Cell 6: Download model
from google.colab import files
files.download("runs/detect/train/weights/best.pt")
5. Evaluate & Debug
5.1 Model evaluation
"""Evaluate model chi tiết"""
from ultralytics import YOLO
# Load trained model
model = YOLO("runs/detect/my_product_model/weights/best.pt")
# Validate
metrics = model.val(data="dataset.yaml")
# Per-class results
print("\n📊 Per-class Performance:")
for i, name in enumerate(metrics.names.values()):
print(f" {name}:")
print(f" Precision: {metrics.box.p[i]:.4f}")
print(f" Recall: {metrics.box.r[i]:.4f}")
print(f" mAP@50: {metrics.box.ap50[i]:.4f}")
5.2 Confusion Matrix
"""Xem confusion matrix để hiểu errors"""
# Ultralytics tự tạo confusion matrix trong runs/detect/val/
# Kiểm tra file: confusion_matrix.png
# Manual analysis
results = model("test_images/", save=True, conf=0.5)
# Đếm detections theo class
from collections import Counter
class_counts = Counter()
for result in results:
for box in result.boxes:
class_name = result.names[int(box.cls)]
class_counts[class_name] += 1
print("\n📈 Detection Distribution:")
for cls, count in class_counts.most_common():
print(f" {cls}: {count}")
5.3 Common Training Issues
| Problem | Symptoms | Solution |
|---|---|---|
| Overfitting | Train loss ↓, Val loss ↑ | Add data, augmentation, dropout |
| Underfitting | Both losses are high | Increase model size, epochs, lr |
| Label errors | Abnormally low mAP | Review labels, check consistency |
| Small objects miss | Low recall for small objects | Increase imgsz=1280, add SAHI |
| Class imbalance | 1 class detects poorly | Oversample minority class |
6. Export Model
"""Export model sang các format khác nhau"""
model = YOLO("runs/detect/train/weights/best.pt")
# ONNX — cross-platform
model.export(format="onnx", imgsz=640, simplify=True)
# TensorRT — NVIDIA GPU (nhanh nhất)
model.export(format="engine", imgsz=640, half=True)
# CoreML — iOS
model.export(format="coreml", imgsz=640)
# TFLite — Android / Edge
model.export(format="tflite", imgsz=640)
# OpenVINO — Intel hardware
model.export(format="openvino", imgsz=640)
7. Use Case: Detecting defective products
"""Complete pipeline: phát hiện sản phẩm lỗi trong nhà máy"""
from ultralytics import YOLO
import cv2
from datetime import datetime
# Load custom trained model
model = YOLO("product_defect_model.pt")
def inspect_product(image_path):
"""Kiểm tra 1 sản phẩm"""
results = model(image_path, conf=0.5)
result = results[0]
defects_found = []
for box in result.boxes:
class_name = result.names[int(box.cls)]
confidence = box.conf[0].item()
x1, y1, x2, y2 = box.xyxy[0].tolist()
if "defect" in class_name or "scratch" in class_name:
defects_found.append({
"type": class_name,
"confidence": confidence,
"location": (x1, y1, x2, y2),
})
# Quyết định
status = "❌ REJECT" if defects_found else "✅ PASS"
print(f"{status} | {image_path}")
for d in defects_found:
print(f" ⚠️ {d['type']} ({d['confidence']:.1%})")
return {
"status": "reject" if defects_found else "pass",
"defects": defects_found,
"timestamp": datetime.now().isoformat(),
}
# Test
inspect_product("product_001.jpg")
inspect_product("product_002.jpg")
Summary
| Step | Work | Tools |
|---|---|---|
| 1. Data Collection | Collect 300+ photos/class | Camera, Roboflow Universe |
| 2. Labeling | Drawing bounding boxes | Roboflow, CVAT |
| 3. Config | Create dataset.yaml | YAML file |
| 4. Training | Train YOLO custom | Ultralytics, Google Colab |
| 5. Evaluation | Evaluation of mAP, P, R | model.val() |
| 6. Export | Convert to ONNX/TRT | model.export() |
General exercises
- Quick Project: Load dataset from Roboflow Universe (e.g. "hard hat detection"). Train YOLOv11n. Achieved mAP@50 > 80%.
- Custom Dataset: Take 100 photos of 1 object (for example: cup, pen, water bottle). Label on Roboflow. Train and test.
- Comparison: Train the same dataset as yolo11n, yolo11s, yolo11m. Compare accuracy, speed, model size.
- Export Test: Export to ONNX. Inference using ONNX Runtime. Compare speed with PyTorch.
Next article: Real-time Detection — detect on camera, video stream, object tracking with DeepSORT.