はじめに
コンピューター ビジョン (CV) — または「コンピューター ビジョン」 — は、コンピューターに人間と同じように画像やビデオを見て理解できるように教える AI の分野です。自動運転車、欠陥製品のチェック、顔認識から店内の客数数えまで、すべてはコンピューター ビジョンの応用です。
🎯 履歴書を勉強する理由 コンピューター ビジョンは、最も実用的なアプリケーションを持つ AI の分野です。カメラや写真を扱うあらゆる業界には履歴書が必要です。
1. コンピュータービジョンとは何ですか?
1.1 定義
コンピューター ビジョンは、コンピューターが画像やビデオから 意味のある情報を抽出し、その情報に基づいて 意思決定を行うのを支援する AI の分野です。
Ảnh/Video (pixels) → CV Model → Thông tin (classification, detection, segmentation)
↓
Quyết định / Hành động
1.2 履歴書の主な問題点
| 数学の問題 | 説明 | 例 |
|---|---|---|
| 画像の分類 | 写真がどのグループに属しているかを分類する | 犬と猫、正常な X 線と異常な X 線 |
| 物体検出 | LOCATION + オブジェクト ラベルを検索 | YOLO は人、車両、標識を検出します |
| 画像のセグメンテーション | 各ピクセルを分類 | 前景/背景分離、医療 |
| 姿勢推定 | 体の関節を見つける | スポーツ分析、AR |
| 画像生成 | 新しい写真を作成 | 安定拡散、DALL・E |
| OCR | 画像内のテキストを認識する | ナンバープレートと請求書を読む |
1.3 実際の応用 — どのくらいの価値がありますか?
🚗 Xe tự lái (Tesla, Waymo) → Hàng tỷ USD
🏥 Y tế (X-ray, MRI, pathology) → Cứu mạng người
🏭 Sản xuất (kiểm tra lỗi) → Tiết kiệm triệu USD/năm
🛒 Bán lẻ (đếm khách, heatmap) → Tăng doanh thu 15-20%
🌾 Nông nghiệp (phát hiện sâu bệnh) → Bảo vệ mùa vụ
📱 Smartphone (Face ID, AR) → 3+ tỷ người dùng
🎮 Gaming / AR / VR → Ngành giải trí khổng lồ
2. デジタル画像 — コンピューターはどのように「見える」のでしょうか?
2.1 ピクセル — 基本単位
デジタル画像 = デジタル マトリックス (ピクセル)。各ピクセルは 1 つ以上の数値 (0 ~ 255) です。
import numpy as np
# Grayscale image: 1 giá trị/pixel (0=đen, 255=trắng)
gray_image = np.array([
[0, 50, 100],
[150, 200, 255],
])
# Shape: (2, 3) = 2 rows × 3 cols
# Color image (RGB): 3 giá trị/pixel
color_image = np.array([
[[255, 0, 0], [0, 255, 0]], # Red, Green
[[0, 0, 255], [255, 255, 0]], # Blue, Yellow
])
# Shape: (2, 2, 3) = 2 rows × 2 cols × 3 channels (R,G,B)
2.2 色空間 — RGB、BGR、HSV、グレースケール
import cv2
import matplotlib.pyplot as plt
# OpenCV đọc ảnh dạng BGR (không phải RGB!)
img_bgr = cv2.imread("photo.jpg") # Shape: (H, W, 3) — BGR
img_rgb = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2RGB) # Chuyển sang RGB
img_gray = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2GRAY) # Grayscale
img_hsv = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2HSV) # HSV
print(f"Kích thước ảnh: {img_bgr.shape}")
# Output: (480, 640, 3) → Height=480, Width=640, Channels=3
# HSV hữu ích cho detect màu cụ thể (lọc màu đỏ, xanh...)
# H: Hue (0-179), S: Saturation (0-255), V: Value/Brightness (0-255)
💡 重要な注意: OpenCV は RGB ではなく BGR (青-緑-赤) を使用します。 matplotlib (RGB を使用) で表示する場合は、次のように変換する必要があります。
cv2.cvtColor(img, cv2.COLOR_BGR2RGB)。
3. 基本的な OpenCV — 画像処理
3.1 画像の読み取り、表示、保存
import cv2
import matplotlib.pyplot as plt
# Đọc ảnh
img = cv2.imread("input.jpg")
print(f"Shape: {img.shape}") # (Height, Width, Channels)
print(f"Dtype: {img.dtype}") # uint8 (0-255)
print(f"Size: {img.size} bytes") # Total pixels × channels
# Hiển thị với matplotlib (cần BGR → RGB)
plt.figure(figsize=(10, 6))
plt.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))
plt.title("Original Image")
plt.axis("off")
plt.show()
# Lưu ảnh
cv2.imwrite("output.jpg", img)
3.2 サイズ変更、切り抜き、回転
# === RESIZE ===
# Resize về kích thước cố định
resized = cv2.resize(img, (640, 480)) # (width, height)
# Resize theo tỷ lệ
scaled = cv2.resize(img, None, fx=0.5, fy=0.5) # 50%
# Interpolation methods:
# - cv2.INTER_AREA: tốt khi thu nhỏ
# - cv2.INTER_LINEAR: default, tốt cho phóng to
# - cv2.INTER_CUBIC: chất lượng cao hơn, chậm hơn
resized_hq = cv2.resize(img, (1920, 1080), interpolation=cv2.INTER_CUBIC)
# === CROP ===
# Crop = slicing numpy array — đơn giản!
h, w = img.shape[:2]
cropped = img[100:400, 200:500] # img[y1:y2, x1:x2]
# Crop center
center_crop = img[h//4:3*h//4, w//4:3*w//4]
# === ROTATE ===
# Xoay 90°
rotated_90 = cv2.rotate(img, cv2.ROTATE_90_CLOCKWISE)
# Xoay góc bất kỳ
center = (w//2, h//2)
matrix = cv2.getRotationMatrix2D(center, angle=45, scale=1.0)
rotated_45 = cv2.warpAffine(img, matrix, (w, h))
3.3 写真に描画する — 描画する
# Vẽ lên bản copy (không thay đổi ảnh gốc)
canvas = img.copy()
# Rectangle (bounding box)
cv2.rectangle(canvas, (100, 50), (300, 250), color=(0, 255, 0), thickness=2)
# Circle
cv2.circle(canvas, center=(200, 150), radius=50, color=(255, 0, 0), thickness=3)
# Text
cv2.putText(canvas, "Hello CV!", (100, 50),
fontFace=cv2.FONT_HERSHEY_SIMPLEX,
fontScale=1.0, color=(255, 255, 255), thickness=2)
# Line
cv2.line(canvas, (0, 0), (w, h), color=(0, 0, 255), thickness=2)
4. 画像フィルタリングとエッジ検出
4.1 ブラー/スムージング — ノイズリダクション
# Gaussian Blur — phổ biến nhất
blurred_gauss = cv2.GaussianBlur(img, ksize=(5, 5), sigmaX=0)
# Median Blur — tốt cho salt-and-pepper noise
blurred_median = cv2.medianBlur(img, ksize=5)
# Bilateral Filter — giữ cạnh nét, mượt vùng phẳng
blurred_bilateral = cv2.bilateralFilter(img, d=9, sigmaColor=75, sigmaSpace=75)
# So sánh
fig, axes = plt.subplots(1, 4, figsize=(20, 5))
titles = ["Original", "Gaussian", "Median", "Bilateral"]
images = [img, blurred_gauss, blurred_median, blurred_bilateral]
for ax, title, image in zip(axes, titles, images):
ax.imshow(cv2.cvtColor(image, cv2.COLOR_BGR2RGB))
ax.set_title(title)
ax.axis("off")
plt.tight_layout()
plt.show()
4.2 エッジ検出 — エッジ検出
# Chuyển sang grayscale trước
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
# Canny Edge Detection — phổ biến nhất
edges_canny = cv2.Canny(gray, threshold1=50, threshold2=150)
# Sobel — detect cạnh theo hướng X hoặc Y
sobel_x = cv2.Sobel(gray, cv2.CV_64F, dx=1, dy=0, ksize=3)
sobel_y = cv2.Sobel(gray, cv2.CV_64F, dx=0, dy=1, ksize=3)
sobel_combined = cv2.magnitude(sobel_x, sobel_y)
# Laplacian
laplacian = cv2.Laplacian(gray, cv2.CV_64F)
# Visualize
fig, axes = plt.subplots(1, 4, figsize=(20, 5))
axes[0].imshow(gray, cmap='gray')
axes[0].set_title("Grayscale")
axes[1].imshow(edges_canny, cmap='gray')
axes[1].set_title("Canny")
axes[2].imshow(np.abs(sobel_combined), cmap='gray')
axes[2].set_title("Sobel")
axes[3].imshow(np.abs(laplacian), cmap='gray')
axes[3].set_title("Laplacian")
for ax in axes:
ax.axis("off")
plt.tight_layout()
plt.show()
4.3 しきい値処理 — しきい値処理
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
# Simple Threshold
_, thresh_binary = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY)
# Otsu's Threshold — tự tìm ngưỡng tối ưu
_, thresh_otsu = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
# Adaptive Threshold — thay đổi ngưỡng theo vùng (tốt khi ánh sáng không đều)
thresh_adaptive = cv2.adaptiveThreshold(
gray, 255,
cv2.ADAPTIVE_THRESH_GAUSSIAN_C,
cv2.THRESH_BINARY, blockSize=11, C=2
)
5. ヒストグラム — ピクセル分布を分析する
# Histogram cho ảnh grayscale
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
hist = cv2.calcHist([gray], [0], None, [256], [0, 256])
plt.figure(figsize=(10, 4))
plt.subplot(1, 2, 1)
plt.imshow(gray, cmap='gray')
plt.title("Image")
plt.axis("off")
plt.subplot(1, 2, 2)
plt.plot(hist, color='black')
plt.title("Histogram")
plt.xlabel("Pixel Value")
plt.ylabel("Frequency")
plt.tight_layout()
plt.show()
# Histogram Equalization — cải thiện contrast
equalized = cv2.equalizeHist(gray)
# CLAHE — Adaptive Histogram Equalization (tốt hơn!)
clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8, 8))
clahe_img = clahe.apply(gray)
6. 輪郭 — オブジェクトの輪郭を見つける
# Đọc và chuyển grayscale
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
# Threshold
_, thresh = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY)
# Tìm contours
contours, hierarchy = cv2.findContours(
thresh, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE
)
print(f"Tìm thấy {len(contours)} contours")
# Vẽ contours lên ảnh
result = img.copy()
cv2.drawContours(result, contours, -1, (0, 255, 0), 2)
# Phân tích từng contour
for i, contour in enumerate(contours):
area = cv2.contourArea(contour)
perimeter = cv2.arcLength(contour, closed=True)
x, y, w, h = cv2.boundingRect(contour)
# Lọc contour nhỏ (nhiễu)
if area > 500:
cv2.rectangle(result, (x, y), (x+w, y+h), (255, 0, 0), 2)
cv2.putText(result, f"Area: {area:.0f}", (x, y-10),
cv2.FONT_HERSHEY_SIMPLEX, 0.5, (255, 0, 0), 1)
7. デモ: 基本的な顔認識
"""Face Detection với Haar Cascade — phương pháp truyền thống"""
import cv2
# Load pre-trained face detector
face_cascade = cv2.CascadeClassifier(
cv2.data.haarcascades + "haarcascade_frontalface_default.xml"
)
# Đọc ảnh
img = cv2.imread("group_photo.jpg")
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
# Detect faces
faces = face_cascade.detectMultiScale(
gray,
scaleFactor=1.1, # Scale mỗi bước (1.1 = 10%)
minNeighbors=5, # Số neighbors tối thiểu
minSize=(30, 30) # Kích thước face tối thiểu
)
print(f"Tìm thấy {len(faces)} khuôn mặt")
# Vẽ bounding box
for (x, y, w, h) in faces:
cv2.rectangle(img, (x, y), (x+w, y+h), (0, 255, 0), 2)
cv2.putText(img, "Face", (x, y-10),
cv2.FONT_HERSHEY_SIMPLEX, 0.7, (0, 255, 0), 2)
# Hiển thị
plt.figure(figsize=(12, 8))
plt.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))
plt.title(f"Detected {len(faces)} faces")
plt.axis("off")
plt.show()
⚠️ 注: Haar Cascade は 古いメソッドです。レッスン 4 以降では、はるかに高速かつ正確な YOLO を使用します。
8. 基本的な CV パイプライン — まとめる
"""Pipeline hoàn chỉnh: đọc ảnh → xử lý → detect → hiển thị"""
def simple_cv_pipeline(image_path):
# 1. Đọc ảnh
img = cv2.imread(image_path)
if img is None:
raise FileNotFoundError(f"Không tìm thấy: {image_path}")
# 2. Tiền xử lý
# Resize nếu quá lớn
h, w = img.shape[:2]
if max(h, w) > 1000:
scale = 1000 / max(h, w)
img = cv2.resize(img, None, fx=scale, fy=scale)
# 3. Khử nhiễu
denoised = cv2.GaussianBlur(img, (3, 3), 0)
# 4. Chuyển grayscale
gray = cv2.cvtColor(denoised, cv2.COLOR_BGR2GRAY)
# 5. Edge detection
edges = cv2.Canny(gray, 50, 150)
# 6. Tìm contours
contours, _ = cv2.findContours(
edges, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE
)
# 7. Lọc và vẽ
result = img.copy()
significant = [c for c in contours if cv2.contourArea(c) > 1000]
cv2.drawContours(result, significant, -1, (0, 255, 0), 2)
return result, len(significant)
# Sử dụng
result, count = simple_cv_pipeline("test_image.jpg")
print(f"Detected {count} objects")
概要
| コンセプト | 覚えておいてください |
|---|---|
| コンピュータ ビジョン | AI がコンピューターに写真やビデオを見て理解できるように教える |
| ピクセル | 基本単位、値 0 ~ 255 |
| 色空間 | RGB、BGR (OpenCV)、HSV、グレースケール |
| フィルタリング | ガウスぼかし、中央値、両側 |
| エッジ検出 | キャニー (最も一般的)、ソーベル、ラプラシアン |
| しきい値 | バイナリ、大津(自己選択閾値)、アダプティブ |
| 輪郭 | 形状解析に使用されるオブジェクトのアウトライン |
| ハール カスケード | 従来の顔検出は YOLO に置き換えられます |
一般的な演習
- 環境のセットアップ: OpenCV をインストールします:
pip install opencv-python matplotlib numpy - イメージ エクスプローラー: 画像を読み取り、形状、dtype、最小/最大ピクセル値を出力するスクリプトを作成します。元の画像、グレースケール、ヒストグラムを表示します。
- エッジ アート: 5 つの異なる画像に Canny Edge Detection を適用します。しきい値1としきい値2を変更してみてください。どの写真が最良の結果をもたらしますか?
- 顔カウンター: Haar Cascade を使用して集合写真内の顔を検出します。あなたは顔をいくつ正確に数えることができますか?いつ間違っているのでしょうか?
- ミニ パイプライン: 輪郭 + 円形フィルターを使用して円形のオブジェクト (コイン/ボール) を検出するパイプラインを構築します。
次の記事: CNN Deep Dive — 最新の CV モデルすべての基盤である ResNet、EfficientNet、MobileNet アーキテクチャについて学びます。