簡介
從DALL·E、中途到穩定擴散——從文字創建圖像的人工智慧改變了創意產業。本文深入探討擴散模型:它們如何運作、穩定擴散架構以及文字到圖像、圖像到圖像、ControlNet 的實踐。
🎯 **為什麼要學習? ** 了解擴散模型可以幫助您建立影像、編輯影像並將其應用到 CV 管道中(合成資料、增強、修復)。
1. 擴散模型—理論
1.1 核心思想
Forward Process (Thêm noise):
Ảnh gốc → Ảnh hơi noise → ... → Noise thuần túy (Gaussian)
x_0 x_1 x_T
Reverse Process (Khử noise):
Noise → Bớt noise → ... → Ảnh sạch!
x_T x_{T-1} x_0
Model học: "Cho 1 ảnh có noise, predict noise đã thêm vào"
→ Trừ noise đi = ảnh sạch hơn
→ Lặp lại T bước = ảnh hoàn chỉnh
1.2 為什麼 Diffusion 比 GAN 更好?
| 甘 | 擴散 | |
|---|---|---|
| 訓練 | 困難(模式崩潰、不穩定) | 穩定 |
| 多元化 | 低(模式崩潰) | 高(不同的輸出) |
| 品質 | 不錯不過神器 | 優 |
| 控制 | 難以控制 | 文字提示,ControlNet |
| 速度 | 快速(1 次前向傳球) | 較慢(迭代) |
2. 穩定擴散-架構
2.1 三個主要組成部分
Text Prompt: "A cat sitting on a beach at sunset"
↓
┌──────────────────┐
│ CLIP Text Encoder│ Biến text → embedding vector
│ (Frozen) │
└──────────────────┘
↓ text embedding
┌──────────────────┐
│ U-Net │ Predict noise trong LATENT SPACE
│ (Trained) │ (không phải pixel space!)
│ + Cross-Attention│ (text embedding guide denoising)
└──────────────────┘
↓ denoised latent
┌──────────────────┐
│ VAE Decoder │ Latent → Pixel image (upscale 8×)
│ (Frozen) │
└──────────────────┘
↓
Output: 512×512 image
2.2 潛在空間-速度的秘密
Pixel Space: 512×512×3 = 786,432 numbers → Chậm!
Latent Space: 64×64×4 = 16,384 numbers → Nhanh 48×!
VAE Encoder: Image (512×512) → Latent (64×64)
U-Net: Denoise in Latent Space
VAE Decoder: Latent (64×64) → Image (512×512)
3. 實踐:文字到圖像
3.1 設置
pip install diffusers transformers accelerate torch
3.2 基本文字轉圖像
"""Text-to-Image với Stable Diffusion"""
import torch
from diffusers import StableDiffusionPipeline
# Load model (tải lần đầu ~5GB)
pipe = StableDiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-2-1",
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
# Generate
prompt = "A majestic lion wearing a crown, digital art, 4k, highly detailed"
negative_prompt = "blurry, low quality, deformed"
image = pipe(
prompt=prompt,
negative_prompt=negative_prompt,
num_inference_steps=30, # Số bước denoise (20-50)
guidance_scale=7.5, # CFG: prompt adherence (5-15)
width=768,
height=768,
).images[0]
image.save("lion_king.png")
image.show()
3.3 SDXL — 穩定擴散 XL
"""SDXL: chất lượng cao hơn, 1024×1024"""
from diffusers import StableDiffusionXLPipeline
pipe = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
torch_dtype=torch.float16,
variant="fp16",
)
pipe = pipe.to("cuda")
# Tối ưu memory
pipe.enable_model_cpu_offload()
image = pipe(
prompt="A Vietnamese phở restaurant in cyberpunk style, neon lights, rain",
negative_prompt="ugly, blurry, low quality",
num_inference_steps=40,
guidance_scale=8.0,
).images[0]
image.save("cyberpunk_pho.png")
3.4 重要參數
| 參數 | 價值 | 影響力 |
|---|---|---|
num_inference_steps | 20-50 | 20-50更多 = 更多細節,更慢 |
guidance_scale (CFG) | 1-20 | 1-20高=較黏性,但較不自然 |
seed | 整數 | 相同的種子=相同的結果(可重現) |
negative_prompt | 文字 | 照片中不想要什麼? |
# Reproducible generation
generator = torch.Generator(device="cuda").manual_seed(42)
image = pipe(prompt="...", generator=generator).images[0]
4. 影像到影像
"""Image-to-Image: biến đổi ảnh có sẵn"""
from diffusers import StableDiffusionImg2ImgPipeline
from PIL import Image
pipe = StableDiffusionImg2ImgPipeline.from_pretrained(
"stabilityai/stable-diffusion-2-1",
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
# Load ảnh gốc
init_image = Image.open("sketch.png").resize((768, 768))
# Transform
result = pipe(
prompt="A detailed watercolor painting of a village",
image=init_image,
strength=0.75, # 0 = giữ nguyên, 1 = tạo mới hoàn toàn
guidance_scale=7.5,
num_inference_steps=30,
).images[0]
result.save("watercolor_village.png")
5. ControlNet — 精確控制
"""ControlNet: điều khiển layout ảnh bằng edge/pose/depth"""
from diffusers import StableDiffusionControlNetPipeline, ControlNetModel
import cv2
import numpy as np
# Load ControlNet (Canny edge)
controlnet = ControlNetModel.from_pretrained(
"lllyasviel/control_v11p_sd15_canny",
torch_dtype=torch.float16,
)
pipe = StableDiffusionControlNetPipeline.from_pretrained(
"runwayml/stable-diffusion-v1-5",
controlnet=controlnet,
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
# Tạo Canny edge map từ ảnh gốc
image = np.array(Image.open("room_photo.jpg"))
edges = cv2.Canny(image, 100, 200)
canny_image = Image.fromarray(edges)
# Generate với cùng layout nhưng style khác
result = pipe(
prompt="A modern minimalist living room, interior design magazine",
image=canny_image,
num_inference_steps=30,
).images[0]
result.save("modern_room.png")
6. LoRA-輕微調風格
"""LoRA: thêm style mới cho Stable Diffusion"""
from diffusers import StableDiffusionPipeline
pipe = StableDiffusionPipeline.from_pretrained(
"runwayml/stable-diffusion-v1-5",
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
# Load LoRA weights (ví dụ: anime style)
pipe.load_lora_weights("path/to/anime_lora.safetensors")
# Generate với style mới
image = pipe(
prompt="1girl, cherry blossoms, anime style",
num_inference_steps=30,
).images[0]
# Unload LoRA
pipe.unload_lora_weights()
7. 修復 — 編輯圖片的一部分
"""Inpainting: sửa/thay thế 1 vùng trong ảnh"""
from diffusers import StableDiffusionInpaintPipeline
from PIL import Image
pipe = StableDiffusionInpaintPipeline.from_pretrained(
"stabilityai/stable-diffusion-2-inpainting",
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
# Ảnh gốc + mask (vùng trắng = vùng cần sửa)
image = Image.open("photo.jpg").resize((512, 512))
mask = Image.open("mask.png").resize((512, 512)) # Trắng = replace
result = pipe(
prompt="A beautiful garden with flowers",
image=image,
mask_image=mask,
num_inference_steps=30,
).images[0]
8. CV Pipeline 中的應用
🎨 Synthetic Data: Tạo training data cho rare cases
🔄 Data Augmentation: Stable Diffusion biến thể ảnh
🖼️ Inpainting: Xóa watermark, sửa defects
📐 ControlNet: Maintain layout, thay đổi style
🎭 Style Transfer: Chuyển đổi phong cách ảnh
📸 Super Resolution: Upscale ảnh chất lượng thấp
總結
| 概念 | 記住 |
|---|---|
| 擴散 | 增加雜訊→學習消除雜訊→生成 |
| 潛在空間 | 在壓縮空間 (64×64) 中處理 → 更快 |
| CFG 規模 | 指導意見:高=更及時 |
| 控制網 | 使用邊緣/姿態/深度控制佈局 |
| 洛拉 | 燈光微調新風格 |
| 修復 | 編輯部分照片,保留其餘部分 |
一般練習
- 文字轉圖像: 產生 10 張具有相同提示但不同種子的圖像。哪張照片最漂亮?
- 風格轉換: 使用img2img將照片變成油畫、水彩、動漫。
- ControlNet: 從您的房間使用 Canny Edge → 產生不同風格的房間。
- 合成資料: 使用 SD 為訓練資料集建立 50 個「產品缺陷」影像。
下一篇文章: Vision Transformer (ViT) 和 CLIP — 圖像轉換器,連接文字和圖像。