はじめに
安定拡散は 潜在拡散モデル (LDM) です。ピクセル空間 (512×512×3) での拡散ではなく、48 倍小さい 潜在空間 (64×64×4) で動作します。これは、コンシューマー GPU での拡散の実行を可能にする画期的な進歩です。
1. 安定した拡散アーキテクチャ
┌─────────────────────────────────────────────────────────┐
│ STABLE DIFFUSION PIPELINE │
│ │
│ "a cat in space" │
│ ↓ │
│ ┌──────────┐ ┌─────────────────────────┐ │
│ │ CLIP │────→│ UNet + Scheduler │ │
│ │ Text │ │ (Denoise in latent) │ │
│ │ Encoder │ │ T steps: 20-50 │ │
│ └──────────┘ └───────────┬─────────────┘ │
│ ↓ │
│ ┌──────────────┐ │
│ Noise z (64x64x4) → │ VAE Decoder │ → Image 512x512│
│ └──────────────┘ │
└─────────────────────────────────────────────────────────┘
コンポーネントの詳細
from diffusers import StableDiffusionPipeline
pipe = StableDiffusionPipeline.from_pretrained("stabilityai/stable-diffusion-xl-base-1.0")
# 4 components chính:
# 1. Text Encoder (CLIP): text → embeddings
pipe.text_encoder # CLIPTextModel
# 2. UNet: noise predictor, conditioned on text
pipe.unet # UNet2DConditionModel
# 3. VAE: latent ↔ pixel space
pipe.vae # AutoencoderKL
# 4. Scheduler: noise schedule algorithm
pipe.scheduler # EulerDiscreteScheduler
2. パイプラインのステップバイステップ
import torch
from diffusers import AutoencoderKL, UNet2DConditionModel, EulerDiscreteScheduler
from transformers import CLIPTextModel, CLIPTokenizer
# Step 1: Tokenize & encode text
tokenizer = CLIPTokenizer.from_pretrained("openai/clip-vit-large-patch14")
text_encoder = CLIPTextModel.from_pretrained("openai/clip-vit-large-patch14")
prompt = "a cat wearing sunglasses, digital art"
tokens = tokenizer(prompt, return_tensors="pt", padding="max_length", max_length=77)
text_embeddings = text_encoder(tokens.input_ids)[0] # [1, 77, 768]
# Step 2: Initialize random latent
latent = torch.randn(1, 4, 64, 64) # latent space
# Step 3: Denoise loop
scheduler = EulerDiscreteScheduler(num_train_timesteps=1000)
scheduler.set_timesteps(30) # 30 denoising steps
for t in scheduler.timesteps:
# Predict noise conditioned on text
noise_pred = unet(latent, t, encoder_hidden_states=text_embeddings).sample
# Classifier-free guidance
noise_uncond = unet(latent, t, encoder_hidden_states=uncond_embeddings).sample
noise_pred = noise_uncond + guidance_scale * (noise_pred - noise_uncond)
# Update latent
latent = scheduler.step(noise_pred, t, latent).prev_sample
# Step 4: Decode latent → pixel image
image = vae.decode(latent / 0.18215).sample
3. 分類子を使用しないガイダンス (CFG)
$$\hat{\epsilon} = \epsilon_{uncond} + s \cdot (\epsilon_{cond} - \epsilon_{uncond})$$
# guidance_scale s controls text adherence
# s = 1: no guidance (ignore prompt)
# s = 7-8: balanced (default)
# s = 15-20: strong adherence (can be over-saturated)
guidance_scale = 7.5
# During inference: run UNet twice
noise_cond = unet(latent, t, text_embeddings) # conditioned
noise_uncond = unet(latent, t, empty_embeddings) # unconditioned
# Interpolate
noise_pred = noise_uncond + guidance_scale * (noise_cond - noise_uncond)
4. VAE — 潜在空間圧縮
# Encode: 512x512x3 → 64x64x4 (compression ratio ~48x)
with torch.no_grad():
latent = vae.encode(image).latent_dist.sample()
latent = latent * 0.18215 # scaling factor
# Decode: 64x64x4 → 512x512x3
with torch.no_grad():
image = vae.decode(latent / 0.18215).sample
なぜ潜在空間なのでしょうか?
- メモリ: 512×512×3 = 786K ピクセル → 64×64×4 = 16K 値
- 速度: UNet は 48x より小さいテンソルを処理します
- 品質: VAE はスマート圧縮を学習しました
5. スケジューラ — ノイズ除去アルゴリズム
from diffusers import (
DDPMScheduler, # Original, 1000 steps
DDIMScheduler, # Deterministic, 50 steps
EulerDiscreteScheduler, # Fast, 20-30 steps
DPMSolverMultistepScheduler, # DPM++, 20 steps, high quality
UniPCMultistepScheduler, # 10-20 steps
)
# Đổi scheduler dễ dàng
pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)
| スケジューラー | ステップ | スピード | 品質 | 使用例 |
|---|---|---|---|---|
| DDPM | 1000 | 非常に遅い | 参考資料 | トレーニング |
| DDIM | 50 | 中 | 良い | 一般 |
| オイラー | 20-30 | 速い | すばらしい | デフォルトの SDXL |
| DPM++ 2M | 20-25 | 速い | 素晴らしい | おすすめ |
| ユニPC | 10-15 | 非常に速い | 良い | リアルタイム |
6. 安定した拡散バージョン
| バージョン | 解像度 | テキストエンコーダ | リリース |
|---|---|---|---|
| SD1.5 | 512×512 | クリップ ViT-L/14 | 2022年 |
| SD2.1 | 768×768 | OpenCLIP ViT-H | 2022年 |
| SDXL | 1024×1024 | CLIP ViT-L + OpenCLIP ViT-bigG | 2023年 |
| SD3 | 1024×1024 | トリプルテキストエンコーダー(CLIP×2 + T5) | 2024年 |
| フラックス | 1024×1024 | T5XXL | 2024年 |
# SDXL — recommended cho production
pipe = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
torch_dtype=torch.float16,
variant="fp16",
)
pipe.to("cuda")
image = pipe(
prompt="a majestic lion in a forest, photorealistic",
negative_prompt="blurry, low quality, distorted",
num_inference_steps=30,
guidance_scale=7.5,
width=1024,
height=1024,
).images[0]
7. 否定的なプロンプトとパラメータ
# Negative prompt: things to avoid
negative_prompt = "blurry, low quality, deformed, ugly, bad anatomy"
# Key parameters
image = pipe(
prompt="...",
negative_prompt=negative_prompt,
num_inference_steps=30, # More = better quality, slower
guidance_scale=7.5, # Text adherence (7-9 optimal)
width=1024,
height=1024,
seed=42, # Reproducibility
).images[0]
概要
| コンポーネント | 役割 |
|---|---|
| CLIP テキスト エンコーダ | テキスト → 埋め込み (意味論的な意味) |
| Uネット | テキストに基づいたノイズ予測器 |
| VAE | ピクセル ↔ 潜在空間を圧縮 |
| スケジューラー | ノイズ除去ステップのアルゴリズム |
| CFG | ガイダンス スケールでテキストの遵守度を調整 |
📌 次の記事: 画像生成のためのプロンプト エンジニアリング — 効果的なプロンプト作成テクニック。