簡介
電腦視覺 (CV) — 或「電腦視覺」 — 是人工智慧領域,教導電腦像人類一樣查看和理解圖像和影片。從自動駕駛汽車、檢查缺陷產品、臉部辨識到統計商店中的顧客數量——所有這些都是電腦視覺的應用。
🎯 為什麼要學履歷? ** 電腦視覺是人工智慧領域最實際應用**的領域——任何有相機和照片的行業都需要簡歷。
1.什麼是電腦視覺?
1.1 定義
電腦視覺是人工智慧領域,可協助電腦從圖像或影片中提取有意義的資訊,並根據該資訊做出決策。
Ảnh/Video (pixels) → CV Model → Thông tin (classification, detection, segmentation)
↓
Quyết định / Hành động
1.2 CV中的主要問題
| 數學問題 | 描述 | 範例 |
|---|---|---|
| 影像分類 | 分類照片屬於哪一組 | 狗與貓,正常與異常 X 光 |
| 物件偵測 | 尋找位置 + 物件標籤 | YOLO 偵測人、車輛、標誌 |
| 影像分割 | 對每個像素進行分類 | 前景/背景分離,醫療 |
| 姿勢估計 | 尋找身體關節 | 身體分析,AR |
| 影像產生 | 建立新照片 | 穩定擴散,DALL·E |
| OCR | 識別圖像中的文字 | 閱讀車牌和發票 |
1.3 實際應用-它值多少錢?
🚗 Xe tự lái (Tesla, Waymo) → Hàng tỷ USD
🏥 Y tế (X-ray, MRI, pathology) → Cứu mạng người
🏭 Sản xuất (kiểm tra lỗi) → Tiết kiệm triệu USD/năm
🛒 Bán lẻ (đếm khách, heatmap) → Tăng doanh thu 15-20%
🌾 Nông nghiệp (phát hiện sâu bệnh) → Bảo vệ mùa vụ
📱 Smartphone (Face ID, AR) → 3+ tỷ người dùng
🎮 Gaming / AR / VR → Ngành giải trí khổng lồ
2. 數位影像 — 電腦「看起來」如何?
2.1 像素-基本單位
數位影像=數位矩陣(像素)。每個像素是 1 個或多個數值 (0-255):
import numpy as np
# Grayscale image: 1 giá trị/pixel (0=đen, 255=trắng)
gray_image = np.array([
[0, 50, 100],
[150, 200, 255],
])
# Shape: (2, 3) = 2 rows × 3 cols
# Color image (RGB): 3 giá trị/pixel
color_image = np.array([
[[255, 0, 0], [0, 255, 0]], # Red, Green
[[0, 0, 255], [255, 255, 0]], # Blue, Yellow
])
# Shape: (2, 2, 3) = 2 rows × 2 cols × 3 channels (R,G,B)
2.2 色彩空間 — RGB、BGR、HSV、灰階
import cv2
import matplotlib.pyplot as plt
# OpenCV đọc ảnh dạng BGR (không phải RGB!)
img_bgr = cv2.imread("photo.jpg") # Shape: (H, W, 3) — BGR
img_rgb = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2RGB) # Chuyển sang RGB
img_gray = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2GRAY) # Grayscale
img_hsv = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2HSV) # HSV
print(f"Kích thước ảnh: {img_bgr.shape}")
# Output: (480, 640, 3) → Height=480, Width=640, Channels=3
# HSV hữu ích cho detect màu cụ thể (lọc màu đỏ, xanh...)
# H: Hue (0-179), S: Saturation (0-255), V: Value/Brightness (0-255)
💡重要提示: OpenCV 使用 BGR(藍色-綠色-紅色),而不是 RGB。使用 matplotlib(使用 RGB)顯示時,必須轉換:
cv2.cvtColor(img, cv2.COLOR_BGR2RGB)。
3. 基礎 OpenCV — 影像處理
3.1 讀取、顯示與儲存影像
import cv2
import matplotlib.pyplot as plt
# Đọc ảnh
img = cv2.imread("input.jpg")
print(f"Shape: {img.shape}") # (Height, Width, Channels)
print(f"Dtype: {img.dtype}") # uint8 (0-255)
print(f"Size: {img.size} bytes") # Total pixels × channels
# Hiển thị với matplotlib (cần BGR → RGB)
plt.figure(figsize=(10, 6))
plt.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))
plt.title("Original Image")
plt.axis("off")
plt.show()
# Lưu ảnh
cv2.imwrite("output.jpg", img)
3.2 調整大小、裁切、旋轉
# === RESIZE ===
# Resize về kích thước cố định
resized = cv2.resize(img, (640, 480)) # (width, height)
# Resize theo tỷ lệ
scaled = cv2.resize(img, None, fx=0.5, fy=0.5) # 50%
# Interpolation methods:
# - cv2.INTER_AREA: tốt khi thu nhỏ
# - cv2.INTER_LINEAR: default, tốt cho phóng to
# - cv2.INTER_CUBIC: chất lượng cao hơn, chậm hơn
resized_hq = cv2.resize(img, (1920, 1080), interpolation=cv2.INTER_CUBIC)
# === CROP ===
# Crop = slicing numpy array — đơn giản!
h, w = img.shape[:2]
cropped = img[100:400, 200:500] # img[y1:y2, x1:x2]
# Crop center
center_crop = img[h//4:3*h//4, w//4:3*w//4]
# === ROTATE ===
# Xoay 90°
rotated_90 = cv2.rotate(img, cv2.ROTATE_90_CLOCKWISE)
# Xoay góc bất kỳ
center = (w//2, h//2)
matrix = cv2.getRotationMatrix2D(center, angle=45, scale=1.0)
rotated_45 = cv2.warpAffine(img, matrix, (w, h))
3.3 在照片上繪圖 — 繪圖
# Vẽ lên bản copy (không thay đổi ảnh gốc)
canvas = img.copy()
# Rectangle (bounding box)
cv2.rectangle(canvas, (100, 50), (300, 250), color=(0, 255, 0), thickness=2)
# Circle
cv2.circle(canvas, center=(200, 150), radius=50, color=(255, 0, 0), thickness=3)
# Text
cv2.putText(canvas, "Hello CV!", (100, 50),
fontFace=cv2.FONT_HERSHEY_SIMPLEX,
fontScale=1.0, color=(255, 255, 255), thickness=2)
# Line
cv2.line(canvas, (0, 0), (w, h), color=(0, 0, 255), thickness=2)
4. 影像濾鏡與邊緣偵測
4.1 模糊/平滑-降噪
# Gaussian Blur — phổ biến nhất
blurred_gauss = cv2.GaussianBlur(img, ksize=(5, 5), sigmaX=0)
# Median Blur — tốt cho salt-and-pepper noise
blurred_median = cv2.medianBlur(img, ksize=5)
# Bilateral Filter — giữ cạnh nét, mượt vùng phẳng
blurred_bilateral = cv2.bilateralFilter(img, d=9, sigmaColor=75, sigmaSpace=75)
# So sánh
fig, axes = plt.subplots(1, 4, figsize=(20, 5))
titles = ["Original", "Gaussian", "Median", "Bilateral"]
images = [img, blurred_gauss, blurred_median, blurred_bilateral]
for ax, title, image in zip(axes, titles, images):
ax.imshow(cv2.cvtColor(image, cv2.COLOR_BGR2RGB))
ax.set_title(title)
ax.axis("off")
plt.tight_layout()
plt.show()
4.2 邊緣偵測-邊緣偵測
# Chuyển sang grayscale trước
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
# Canny Edge Detection — phổ biến nhất
edges_canny = cv2.Canny(gray, threshold1=50, threshold2=150)
# Sobel — detect cạnh theo hướng X hoặc Y
sobel_x = cv2.Sobel(gray, cv2.CV_64F, dx=1, dy=0, ksize=3)
sobel_y = cv2.Sobel(gray, cv2.CV_64F, dx=0, dy=1, ksize=3)
sobel_combined = cv2.magnitude(sobel_x, sobel_y)
# Laplacian
laplacian = cv2.Laplacian(gray, cv2.CV_64F)
# Visualize
fig, axes = plt.subplots(1, 4, figsize=(20, 5))
axes[0].imshow(gray, cmap='gray')
axes[0].set_title("Grayscale")
axes[1].imshow(edges_canny, cmap='gray')
axes[1].set_title("Canny")
axes[2].imshow(np.abs(sobel_combined), cmap='gray')
axes[2].set_title("Sobel")
axes[3].imshow(np.abs(laplacian), cmap='gray')
axes[3].set_title("Laplacian")
for ax in axes:
ax.axis("off")
plt.tight_layout()
plt.show()
4.3 閾值處理-閾值處理
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
# Simple Threshold
_, thresh_binary = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY)
# Otsu's Threshold — tự tìm ngưỡng tối ưu
_, thresh_otsu = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
# Adaptive Threshold — thay đổi ngưỡng theo vùng (tốt khi ánh sáng không đều)
thresh_adaptive = cv2.adaptiveThreshold(
gray, 255,
cv2.ADAPTIVE_THRESH_GAUSSIAN_C,
cv2.THRESH_BINARY, blockSize=11, C=2
)
5. 直方圖 — 分析像素分佈
# Histogram cho ảnh grayscale
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
hist = cv2.calcHist([gray], [0], None, [256], [0, 256])
plt.figure(figsize=(10, 4))
plt.subplot(1, 2, 1)
plt.imshow(gray, cmap='gray')
plt.title("Image")
plt.axis("off")
plt.subplot(1, 2, 2)
plt.plot(hist, color='black')
plt.title("Histogram")
plt.xlabel("Pixel Value")
plt.ylabel("Frequency")
plt.tight_layout()
plt.show()
# Histogram Equalization — cải thiện contrast
equalized = cv2.equalizeHist(gray)
# CLAHE — Adaptive Histogram Equalization (tốt hơn!)
clahe = cv2.createCLAHE(clipLimit=2.0, tileGridSize=(8, 8))
clahe_img = clahe.apply(gray)
6. 輪廓 — 尋找物件輪廓
# Đọc và chuyển grayscale
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
# Threshold
_, thresh = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY)
# Tìm contours
contours, hierarchy = cv2.findContours(
thresh, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE
)
print(f"Tìm thấy {len(contours)} contours")
# Vẽ contours lên ảnh
result = img.copy()
cv2.drawContours(result, contours, -1, (0, 255, 0), 2)
# Phân tích từng contour
for i, contour in enumerate(contours):
area = cv2.contourArea(contour)
perimeter = cv2.arcLength(contour, closed=True)
x, y, w, h = cv2.boundingRect(contour)
# Lọc contour nhỏ (nhiễu)
if area > 500:
cv2.rectangle(result, (x, y), (x+w, y+h), (255, 0, 0), 2)
cv2.putText(result, f"Area: {area:.0f}", (x, y-10),
cv2.FONT_HERSHEY_SIMPLEX, 0.5, (255, 0, 0), 1)
7. 示範:基本臉部辨識
"""Face Detection với Haar Cascade — phương pháp truyền thống"""
import cv2
# Load pre-trained face detector
face_cascade = cv2.CascadeClassifier(
cv2.data.haarcascades + "haarcascade_frontalface_default.xml"
)
# Đọc ảnh
img = cv2.imread("group_photo.jpg")
gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
# Detect faces
faces = face_cascade.detectMultiScale(
gray,
scaleFactor=1.1, # Scale mỗi bước (1.1 = 10%)
minNeighbors=5, # Số neighbors tối thiểu
minSize=(30, 30) # Kích thước face tối thiểu
)
print(f"Tìm thấy {len(faces)} khuôn mặt")
# Vẽ bounding box
for (x, y, w, h) in faces:
cv2.rectangle(img, (x, y), (x+w, y+h), (0, 255, 0), 2)
cv2.putText(img, "Face", (x, y-10),
cv2.FONT_HERSHEY_SIMPLEX, 0.7, (0, 255, 0), 2)
# Hiển thị
plt.figure(figsize=(12, 8))
plt.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))
plt.title(f"Detected {len(faces)} faces")
plt.axis("off")
plt.show()
⚠️ 注意: Haar Cascade 是一種 舊 方法。從第 4 課開始,我們將使用 YOLO — 更快、更準確。
8. 基本履歷管道 - 將其放在一起
"""Pipeline hoàn chỉnh: đọc ảnh → xử lý → detect → hiển thị"""
def simple_cv_pipeline(image_path):
# 1. Đọc ảnh
img = cv2.imread(image_path)
if img is None:
raise FileNotFoundError(f"Không tìm thấy: {image_path}")
# 2. Tiền xử lý
# Resize nếu quá lớn
h, w = img.shape[:2]
if max(h, w) > 1000:
scale = 1000 / max(h, w)
img = cv2.resize(img, None, fx=scale, fy=scale)
# 3. Khử nhiễu
denoised = cv2.GaussianBlur(img, (3, 3), 0)
# 4. Chuyển grayscale
gray = cv2.cvtColor(denoised, cv2.COLOR_BGR2GRAY)
# 5. Edge detection
edges = cv2.Canny(gray, 50, 150)
# 6. Tìm contours
contours, _ = cv2.findContours(
edges, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE
)
# 7. Lọc và vẽ
result = img.copy()
significant = [c for c in contours if cv2.contourArea(c) > 1000]
cv2.drawContours(result, significant, -1, (0, 255, 0), 2)
return result, len(significant)
# Sử dụng
result, count = simple_cv_pipeline("test_image.jpg")
print(f"Detected {count} objects")
總結
| 概念 | 記住 |
|---|---|
| 電腦視覺 | 人工智慧教會電腦檢視和理解照片/影片 |
| 像素 | 基本單位,值0-255 |
| 色彩空間 | RGB、BGR (OpenCV)、HSV、灰階 |
| 過濾 | 高斯模糊、中位數、雙邊 |
| 邊緣偵測 | Canny(最常見)、Sobel、拉普拉斯 |
| 閾值 | 二元、Otsu(自選閾值)、自適應 |
| 輪廓 | 物體輪廓,用於形狀分析 |
| 哈爾級聯 | 經典人臉偵測,將被YOLO取代 |
一般練習
- 設定環境: 安裝OpenCV:
pip install opencv-python matplotlib numpy - 圖像瀏覽器: 編寫腳本來讀取圖像,列印出形狀、資料類型、最小/最大像素值。顯示原始影像、灰階圖和直方圖。
- 邊緣藝術: 將 Canny 邊緣偵測應用於 5 個不同的影像。嘗試更改閾值1和閾值2。哪張照片效果最好?
- 人臉計數器: 使用 Haar Cascade 偵測合照中的人臉。你能正確數出多少張臉?什麼時候錯了?
- 迷你管道: 使用輪廓+圓形過濾器建立管道來偵測圓形物體(硬幣/球)。
下一篇文章: CNN 深度探究 — 了解 ResNet、EfficientNet、MobileNet 架構 — 所有現代 CV 模型的基礎。