Introduction
From DALL·E, Midjourney to Stable Diffusion — AI that creates images from text has changed the creative industry. This article dives deep into Diffusion Models: how they work, the Stable Diffusion architecture, and hands-on text-to-image, image-to-image, ControlNet.
🎯 Why learn? Understanding Diffusion Models helps you create images, edit images, and apply them in the CV pipeline (synthetic data, augmentation, inpainting).
1. Diffusion Models — Theory
1.1 Core ideas
Forward Process (Thêm noise):
Ảnh gốc → Ảnh hơi noise → ... → Noise thuần túy (Gaussian)
x_0 x_1 x_T
Reverse Process (Khử noise):
Noise → Bớt noise → ... → Ảnh sạch!
x_T x_{T-1} x_0
Model học: "Cho 1 ảnh có noise, predict noise đã thêm vào"
→ Trừ noise đi = ảnh sạch hơn
→ Lặp lại T bước = ảnh hoàn chỉnh
1.2 Why is Diffusion better than GAN?
| GAN | Diffusion | |
|---|---|---|
| Training | Difficult (mode collapse, unstable) | Stable |
| Diversity | Low (mode collapse) | High (varied output) |
| Quality | Good but artifacts | Excellent |
| Control | Difficult to control | Text prompt, ControlNet |
| Speed | Fast (1 forward pass) | Slower (iterative) |
2. Stable Diffusion — Architecture
2.1 Three main components
Text Prompt: "A cat sitting on a beach at sunset"
↓
┌──────────────────┐
│ CLIP Text Encoder│ Biến text → embedding vector
│ (Frozen) │
└──────────────────┘
↓ text embedding
┌──────────────────┐
│ U-Net │ Predict noise trong LATENT SPACE
│ (Trained) │ (không phải pixel space!)
│ + Cross-Attention│ (text embedding guide denoising)
└──────────────────┘
↓ denoised latent
┌──────────────────┐
│ VAE Decoder │ Latent → Pixel image (upscale 8×)
│ (Frozen) │
└──────────────────┘
↓
Output: 512×512 image
2.2 Latent Space — The secret to speed
Pixel Space: 512×512×3 = 786,432 numbers → Chậm!
Latent Space: 64×64×4 = 16,384 numbers → Nhanh 48×!
VAE Encoder: Image (512×512) → Latent (64×64)
U-Net: Denoise in Latent Space
VAE Decoder: Latent (64×64) → Image (512×512)
3. Hands-on: Text-to-Image
3.1 Setup
pip install diffusers transformers accelerate torch
3.2 Basic Text-to-Image
"""Text-to-Image với Stable Diffusion"""
import torch
from diffusers import StableDiffusionPipeline
# Load model (tải lần đầu ~5GB)
pipe = StableDiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-2-1",
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
# Generate
prompt = "A majestic lion wearing a crown, digital art, 4k, highly detailed"
negative_prompt = "blurry, low quality, deformed"
image = pipe(
prompt=prompt,
negative_prompt=negative_prompt,
num_inference_steps=30, # Số bước denoise (20-50)
guidance_scale=7.5, # CFG: prompt adherence (5-15)
width=768,
height=768,
).images[0]
image.save("lion_king.png")
image.show()
3.3 SDXL — Stable Diffusion XL
"""SDXL: chất lượng cao hơn, 1024×1024"""
from diffusers import StableDiffusionXLPipeline
pipe = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
torch_dtype=torch.float16,
variant="fp16",
)
pipe = pipe.to("cuda")
# Tối ưu memory
pipe.enable_model_cpu_offload()
image = pipe(
prompt="A Vietnamese phở restaurant in cyberpunk style, neon lights, rain",
negative_prompt="ugly, blurry, low quality",
num_inference_steps=40,
guidance_scale=8.0,
).images[0]
image.save("cyberpunk_pho.png")
3.4 Important Parameters
| Parameters | Value | Influence |
|---|---|---|
num_inference_steps | 20-50 | More = more detail, slower |
guidance_scale (CFG) | 1-20 | High = more sticky, but less natural |
seed | int | Same seed = same result (reproducible) |
negative_prompt | text | What NOT to want in photos |
# Reproducible generation
generator = torch.Generator(device="cuda").manual_seed(42)
image = pipe(prompt="...", generator=generator).images[0]
4. Image-to-Image
"""Image-to-Image: biến đổi ảnh có sẵn"""
from diffusers import StableDiffusionImg2ImgPipeline
from PIL import Image
pipe = StableDiffusionImg2ImgPipeline.from_pretrained(
"stabilityai/stable-diffusion-2-1",
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
# Load ảnh gốc
init_image = Image.open("sketch.png").resize((768, 768))
# Transform
result = pipe(
prompt="A detailed watercolor painting of a village",
image=init_image,
strength=0.75, # 0 = giữ nguyên, 1 = tạo mới hoàn toàn
guidance_scale=7.5,
num_inference_steps=30,
).images[0]
result.save("watercolor_village.png")
5. ControlNet — Precise control
"""ControlNet: điều khiển layout ảnh bằng edge/pose/depth"""
from diffusers import StableDiffusionControlNetPipeline, ControlNetModel
import cv2
import numpy as np
# Load ControlNet (Canny edge)
controlnet = ControlNetModel.from_pretrained(
"lllyasviel/control_v11p_sd15_canny",
torch_dtype=torch.float16,
)
pipe = StableDiffusionControlNetPipeline.from_pretrained(
"runwayml/stable-diffusion-v1-5",
controlnet=controlnet,
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
# Tạo Canny edge map từ ảnh gốc
image = np.array(Image.open("room_photo.jpg"))
edges = cv2.Canny(image, 100, 200)
canny_image = Image.fromarray(edges)
# Generate với cùng layout nhưng style khác
result = pipe(
prompt="A modern minimalist living room, interior design magazine",
image=canny_image,
num_inference_steps=30,
).images[0]
result.save("modern_room.png")
6. LoRA — Light Fine-tune Style
"""LoRA: thêm style mới cho Stable Diffusion"""
from diffusers import StableDiffusionPipeline
pipe = StableDiffusionPipeline.from_pretrained(
"runwayml/stable-diffusion-v1-5",
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
# Load LoRA weights (ví dụ: anime style)
pipe.load_lora_weights("path/to/anime_lora.safetensors")
# Generate với style mới
image = pipe(
prompt="1girl, cherry blossoms, anime style",
num_inference_steps=30,
).images[0]
# Unload LoRA
pipe.unload_lora_weights()
7. Inpainting — Edit part of an image
"""Inpainting: sửa/thay thế 1 vùng trong ảnh"""
from diffusers import StableDiffusionInpaintPipeline
from PIL import Image
pipe = StableDiffusionInpaintPipeline.from_pretrained(
"stabilityai/stable-diffusion-2-inpainting",
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
# Ảnh gốc + mask (vùng trắng = vùng cần sửa)
image = Image.open("photo.jpg").resize((512, 512))
mask = Image.open("mask.png").resize((512, 512)) # Trắng = replace
result = pipe(
prompt="A beautiful garden with flowers",
image=image,
mask_image=mask,
num_inference_steps=30,
).images[0]
8. Application in CV Pipeline
🎨 Synthetic Data: Tạo training data cho rare cases
🔄 Data Augmentation: Stable Diffusion biến thể ảnh
🖼️ Inpainting: Xóa watermark, sửa defects
📐 ControlNet: Maintain layout, thay đổi style
🎭 Style Transfer: Chuyển đổi phong cách ảnh
📸 Super Resolution: Upscale ảnh chất lượng thấp
Summary
| Concepts | Remember |
|---|---|
| Diffusion | Add noise → learn to remove noise → generate |
| Latent Space | Processing in compressed space (64×64) → faster |
| CFG Scale | Guidance: high = more prompt |
| ControlNet | Control layout using edge/pose/depth |
| LoRA | Light fine-tune for new style |
| Inpainting | Edit part of the photo, keep the rest |
General exercises
- Text-to-Image: Generate 10 images with the same prompt but different seeds. Which photo is the most beautiful?
- Style Transfer: Use img2img to turn a photo into oil painting, watercolor, anime.
- ControlNet: Use Canny edge from your room → generate room in a different style.
- Synthetic Data: Use SD to create 50 "product defect" images for the training dataset.
Next article: Vision Transformer (ViT) & CLIP — Transformer for images, connecting text and images.