Chuyển đến nội dung chính

レッスン 4: YOLO v3 から v11 — 理論と実践

YOLO の歴史: YOLOv3 から YOLOv11 (Ultralytics) まで。 YOLO アーキテクチャ、アンカー ボックス、非最大抑制。ハンズオン: YOLOv8/v11 の事前トレーニング済みモデルを使用してオブジェクトを検出します。メトリック: mAP、IoU、精度、リコール。

🧠 AI と ML — レッスン 3 レッスン 4: YOLO v3 から v11 — 理論と 練習する

深層学習によるコンピューター ビジョン: CNN から Vision Transformer まで

パート 2: 物体の検出

xdev.asia

はじめに

オブジェクト検出 = 画像内のすべてのオブジェクトの 位置 (境界ボックス) + ラベル (クラス) を見つけます。そして、YOLO (You Only Look Once) が王です 👑 - 最速、最も正確、そして使いやすい。

🎯 YOLO を選択する理由 リアルタイム (>100 FPS)、高精度 (mAP 50%+)、検出するコードは 1 行で、Ultralytics エコシステムは非常に成熟しています。


1. 物体検出 101

1.1 分類、検出、セグメンテーション

Image Classification:   "This is a cat"          → 1 label
Object Detection:       "Cat at (x,y,w,h)"       → N bounding boxes + labels
Instance Segmentation:  "Cat occupies these pixels" → Pixel-level masks

1.2 境界ボックス — オブジェクトを囲むボックス

# Bounding box formats
# Format 1: (x_center, y_center, width, height) — YOLO format
bbox_yolo = [0.5, 0.5, 0.3, 0.4]  # Normalized (0-1)

# Format 2: (x_min, y_min, x_max, y_max) — Pascal VOC format
bbox_voc = [150, 100, 350, 300]    # Pixels

# Format 3: (x_min, y_min, width, height) — COCO format
bbox_coco = [150, 100, 200, 200]   # Pixels

# Convert YOLO → VOC
def yolo_to_voc(bbox, img_w, img_h):
    x_c, y_c, w, h = bbox
    x_min = int((x_c - w/2) * img_w)
    y_min = int((y_c - h/2) * img_h)
    x_max = int((x_c + w/2) * img_w)
    y_max = int((y_c + h/2) * img_h)
    return [x_min, y_min, x_max, y_max]

1.3 IoU — 和集合上の交差

"""IoU: đo mức overlap giữa 2 bounding boxes"""
def calculate_iou(box1, box2):
    """
    box1, box2: [x_min, y_min, x_max, y_max]
    """
    # Intersection
    x_inter_min = max(box1[0], box2[0])
    y_inter_min = max(box1[1], box2[1])
    x_inter_max = min(box1[2], box2[2])
    y_inter_max = min(box1[3], box2[3])

    inter_area = max(0, x_inter_max - x_inter_min) * \
                 max(0, y_inter_max - y_inter_min)

    # Union
    area1 = (box1[2] - box1[0]) * (box1[3] - box1[1])
    area2 = (box2[2] - box2[0]) * (box2[3] - box2[1])
    union_area = area1 + area2 - inter_area

    return inter_area / union_area if union_area > 0 else 0

# Ví dụ
pred_box = [100, 100, 300, 300]
gt_box = [120, 110, 310, 320]
iou = calculate_iou(pred_box, gt_box)
print(f"IoU: {iou:.4f}")  # ~0.73
IoU Thresholds:
IoU > 0.5  → True Positive (mAP@50)
IoU > 0.75 → True Positive (mAP@75 — strict)
IoU < 0.5  → False Positive (miss!)

2. YOLO — 一度しか見ない

2.1 中心的なアイデア

YOLO 以前は、2 段階の検出が使用されていました: 領域の提案 → 各領域の分類 (R-CNN、遅い)。

YOLO: 1 ステージ — 画像を 1 回だけ見て、すべてのボックスとクラスを出力します。

Input Image (640×640)
    ↓
YOLO Backbone (feature extraction)
    ↓
YOLO Neck (feature fusion — FPN/PAN)
    ↓
YOLO Head (predict boxes + classes)
    ↓
NMS (Non-Max Suppression — lọc box trùng)
    ↓
Final Detections: [(class, confidence, x, y, w, h), ...]

2.2 YOLOの歴史

バージョン年主な貢献mAP(ココ)
YOLOv12016年独自のアイデア — 1 段階検出63.4
ヨロv22017年アンカー ボックス、バッチ正規化78.6
ヨロv32018年マルチスケール検出、Darknet-5333.0 (mAP@50:95)
ヨロv42020年CSPDarknet、Mish アクティベーション、モザイク拡張43.5
ヨロv52020年PyTorch、Ultralytics エコシステム、使いやすい48.2
YOLOv62022年Meituan、BiC モジュール、SimOTA52.5
ヨロv72022年E-ELAN、モデルの再パラメータ化56.8
YOLOv82023アンカーフリー、分離ヘッド、Ultralytics53.9
ヨロv92024年PGI、GELAN アーキテクチャ55.6
YOLOv102024年NMS 不要、効率重視54.4
YOLOv112024C3k2 ブロック、注意、SOTA56.1

⭐ 実際には: YOLOv8 または YOLOv11 (Ultralytics) を使用します。最高のエコシステム、明確なドキュメント、最大のコミュニティ。

2.3 非最大抑制 (NMS)

"""NMS: loại bỏ bounding boxes trùng nhau"""
def nms(boxes, scores, iou_threshold=0.5):
    """
    boxes: [[x1,y1,x2,y2], ...] — tất cả predicted boxes
    scores: [0.9, 0.85, 0.7, ...] — confidence của mỗi box
    """
    indices = sorted(range(len(scores)), key=lambda i: scores[i], reverse=True)
    keep = []

    while indices:
        current = indices.pop(0)
        keep.append(current)

        remaining = []
        for idx in indices:
            iou = calculate_iou(boxes[current], boxes[idx])
            if iou < iou_threshold:  # Chỉ giữ box ít overlap
                remaining.append(idx)
        indices = remaining

    return keep

3. 実践: Ultralytics を使用した YOLO

3.1 インストール

pip install ultralytics

3.2 推論 — 即座に検出

"""YOLO Detection — chỉ 3 dòng code!"""
from ultralytics import YOLO

# Load pretrained model (tự download)
model = YOLO("yolo11n.pt")  # nano (nhanh nhất)
# model = YOLO("yolo11s.pt")  # small
# model = YOLO("yolo11m.pt")  # medium
# model = YOLO("yolo11l.pt")  # large
# model = YOLO("yolo11x.pt")  # extra large (chính xác nhất)

# Detect trên ảnh
results = model("street_photo.jpg")

# Hiển thị kết quả
results[0].show()  # Mở ảnh với bounding boxes
results[0].save("output.jpg")  # Lưu ảnh

3.3 詳細な結果を分析する

"""Phân tích detection results"""
results = model("street.jpg")

for result in results:
    boxes = result.boxes

    for box in boxes:
        # Bounding box coordinates
        x1, y1, x2, y2 = box.xyxy[0].tolist()
        # Confidence score
        confidence = box.conf[0].item()
        # Class
        class_id = int(box.cls[0].item())
        class_name = result.names[class_id]

        print(f"📦 {class_name}: {confidence:.2%} "
              f"at ({x1:.0f}, {y1:.0f}, {x2:.0f}, {y2:.0f})")

    # Summary
    print(f"\nTotal detections: {len(boxes)}")
    print(f"Classes found: {set(result.names[int(c)] for c in boxes.cls)}")

3.4 ビデオでの検出

"""Object detection trên video"""
import cv2
from ultralytics import YOLO

model = YOLO("yolo11n.pt")

# Mở video
cap = cv2.VideoCapture("traffic.mp4")
# Hoặc webcam: cap = cv2.VideoCapture(0)

# Video writer
width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
fps = cap.get(cv2.CAP_PROP_FPS)
writer = cv2.VideoWriter("output.mp4", cv2.VideoWriter_fourcc(*"mp4v"), fps, (width, height))

while cap.isOpened():
    ret, frame = cap.read()
    if not ret:
        break

    # Detect
    results = model(frame, verbose=False)

    # Vẽ results lên frame
    annotated = results[0].plot()

    # Lưu
    writer.write(annotated)

cap.release()
writer.release()
print("Done! Saved to output.mp4")

3.5 YOLO タスク: 検出だけではない

"""YOLO hỗ trợ nhiều tasks"""

# Object Detection
det_model = YOLO("yolo11n.pt")
det_results = det_model("street.jpg")

# Instance Segmentation
seg_model = YOLO("yolo11n-seg.pt")
seg_results = seg_model("street.jpg")

# Pose Estimation
pose_model = YOLO("yolo11n-pose.pt")
pose_results = pose_model("person.jpg")

# Classification
cls_model = YOLO("yolo11n-cls.pt")
cls_results = cls_model("cat.jpg")

# OBB (Oriented Bounding Boxes)
obb_model = YOLO("yolo11n-obb.pt")
obb_results = obb_model("aerial.jpg")

4. 評価指標

4.1 精度、リコール、mAP

                    Predicted Positive    Predicted Negative
Actual Positive     TP (True Positive)    FN (False Negative)
Actual Negative     FP (False Positive)   TN (True Negative)

Precision = TP / (TP + FP)  → "Trong tất cả detect, bao nhiêu đúng?"
Recall    = TP / (TP + FN)  → "Trong tất cả objects thật, bao nhiêu detect được?"

4.2 mAP (平均平均精度)

"""Đánh giá YOLO model"""
from ultralytics import YOLO

model = YOLO("yolo11n.pt")

# Evaluate trên COCO val set
metrics = model.val(data="coco.yaml")

print(f"mAP@50:      {metrics.box.map50:.4f}")     # mAP at IoU=0.5
print(f"mAP@50:95:   {metrics.box.map:.4f}")        # mAP at IoU=0.5:0.95
print(f"Precision:   {metrics.box.mp:.4f}")          # Mean Precision
print(f"Recall:      {metrics.box.mr:.4f}")          # Mean Recall

4.3 YOLO モデルの比較

モデルパラメータmAP@50mAP@50:95速度 (T4 GPU)
ヨロv11n2.6M70.339.51.5ミリ秒
YOLOv11s9.4M77.147.02.5ミリ秒
YOLOv11m20.1M80.451.54.7ミリ秒
YOLOv11l25.3M81.253.46.2ミリ秒
YOLOv11x56.9M82.054.711.3ミリ秒

5. 高度な YOLO 構成

"""Cấu hình detection chi tiết"""

results = model.predict(
    source="image.jpg",        # ảnh, video, folder, url, webcam
    conf=0.25,                 # Confidence threshold (default 0.25)
    iou=0.45,                  # NMS IoU threshold (default 0.7)
    classes=[0, 1, 2],         # Chỉ detect classes cụ thể (0=person, 1=bicycle, 2=car)
    max_det=300,               # Số detection tối đa
    imgsz=640,                 # Input size
    device="cuda",             # GPU
    save=True,                 # Lưu ảnh kết quả
    save_txt=True,             # Lưu labels txt
    save_conf=True,            # Lưu confidence trong txt
    show=False,                # Hiển thị real-time
    verbose=False,             # Tắt log
)

概要

コンセプト覚えておいてください
物体検出オブジェクトの位置 (bbox) + ラベル (クラス) を見つける
ヨロ1 ステージ検出器、リアルタイム、SOTA
IoU予測ボックスとグランド トゥルース ボックス間の重複を測定する
NMS重複したボックスを排除し、最適なボックスを維持
マップ検出の主な指標: 平均平均精度
ウルトラリティクスYOLO の Python ライブラリ — 3 行の検出コード

一般的な演習

  1. クイック スタート: Ultralytics をインストールし、10 種類の画像 (街頭、屋内、自然) で検出します。クラスごとにカウントします。
  2. ビデオ検出: ビデオ トラフィックを検出します。平均して、1 フレームあたり何台の車両ですか?
  3. モデル サイズの比較: yolo11n、yolo11m、yolo11x を比較します: 精度、速度、メモリ。
  4. IoU 計算ツール: IoU 関数を実装し、5 つの異なるボックスのペアでテストします。
  5. カスタム フィルター: 信頼度 > 80% で 人 (クラス 0) のみを検出するスクリプトを作成します。

次の記事: YOLO カスタム トレーニング — 独自のデータにラベルを付け、特定の問題に対して独自のモデルをトレーニングします。