Chuyển đến nội dung chính

レッスン 5: 安定拡散の詳細 - アーキテクチャとパイプライン

潜在拡散モデル: なぜ潜在空間で機能するのでしょうか? UNet アーキテクチャ。 CLIP を使用したテキストの調整。 VAEエンコーダ/デコーダ。スケジューラ: DDIM、オイラー、DPM++。プロンプトから画像までの詳細なパイプライン。

🧠 AI と ML — レッスン 4 レッスン 5: 安定拡散の詳細 - アリ 構造とパイプライン

生成 AI: AI を使用して画像とビデオを作成する

パート 2: 拡散モデル — 革新的なイメージの作成

xdev.asia

はじめに

安定拡散は 潜在拡散モデル (LDM) です。ピクセル空間 (512×512×3) での拡散ではなく、48 倍小さい 潜在空間 (64×64×4) で動作します。これは、コンシューマー GPU での拡散の実行を可能にする画期的な進歩です。


1. 安定した拡散アーキテクチャ

┌─────────────────────────────────────────────────────────┐
│              STABLE DIFFUSION PIPELINE                   │
│                                                         │
│  "a cat in space"                                       │
│        ↓                                                │
│  ┌──────────┐     ┌─────────────────────────┐          │
│  │   CLIP   │────→│    UNet + Scheduler      │          │
│  │  Text    │     │  (Denoise in latent)     │          │
│  │ Encoder  │     │  T steps: 20-50          │          │
│  └──────────┘     └───────────┬─────────────┘          │
│                               ↓                         │
│                       ┌──────────────┐                  │
│  Noise z (64x64x4) → │  VAE Decoder  │ → Image 512x512│
│                       └──────────────┘                  │
└─────────────────────────────────────────────────────────┘

コンポーネントの詳細

from diffusers import StableDiffusionPipeline

pipe = StableDiffusionPipeline.from_pretrained("stabilityai/stable-diffusion-xl-base-1.0")

# 4 components chính:
# 1. Text Encoder (CLIP): text → embeddings
pipe.text_encoder  # CLIPTextModel

# 2. UNet: noise predictor, conditioned on text
pipe.unet  # UNet2DConditionModel

# 3. VAE: latent ↔ pixel space
pipe.vae  # AutoencoderKL

# 4. Scheduler: noise schedule algorithm
pipe.scheduler  # EulerDiscreteScheduler

2. パイプラインのステップバイステップ

import torch
from diffusers import AutoencoderKL, UNet2DConditionModel, EulerDiscreteScheduler
from transformers import CLIPTextModel, CLIPTokenizer

# Step 1: Tokenize & encode text
tokenizer = CLIPTokenizer.from_pretrained("openai/clip-vit-large-patch14")
text_encoder = CLIPTextModel.from_pretrained("openai/clip-vit-large-patch14")

prompt = "a cat wearing sunglasses, digital art"
tokens = tokenizer(prompt, return_tensors="pt", padding="max_length", max_length=77)
text_embeddings = text_encoder(tokens.input_ids)[0]  # [1, 77, 768]

# Step 2: Initialize random latent
latent = torch.randn(1, 4, 64, 64)  # latent space

# Step 3: Denoise loop
scheduler = EulerDiscreteScheduler(num_train_timesteps=1000)
scheduler.set_timesteps(30)  # 30 denoising steps

for t in scheduler.timesteps:
    # Predict noise conditioned on text
    noise_pred = unet(latent, t, encoder_hidden_states=text_embeddings).sample

    # Classifier-free guidance
    noise_uncond = unet(latent, t, encoder_hidden_states=uncond_embeddings).sample
    noise_pred = noise_uncond + guidance_scale * (noise_pred - noise_uncond)

    # Update latent
    latent = scheduler.step(noise_pred, t, latent).prev_sample

# Step 4: Decode latent → pixel image
image = vae.decode(latent / 0.18215).sample

3. 分類子を使用しないガイダンス (CFG)

$$\hat{\epsilon} = \epsilon_{uncond} + s \cdot (\epsilon_{cond} - \epsilon_{uncond})$$

# guidance_scale s controls text adherence
# s = 1: no guidance (ignore prompt)
# s = 7-8: balanced (default)
# s = 15-20: strong adherence (can be over-saturated)

guidance_scale = 7.5

# During inference: run UNet twice
noise_cond = unet(latent, t, text_embeddings)    # conditioned
noise_uncond = unet(latent, t, empty_embeddings)  # unconditioned

# Interpolate
noise_pred = noise_uncond + guidance_scale * (noise_cond - noise_uncond)

4. VAE — 潜在空間圧縮

# Encode: 512x512x3 → 64x64x4 (compression ratio ~48x)
with torch.no_grad():
    latent = vae.encode(image).latent_dist.sample()
    latent = latent * 0.18215  # scaling factor

# Decode: 64x64x4 → 512x512x3
with torch.no_grad():
    image = vae.decode(latent / 0.18215).sample

なぜ潜在空間なのでしょうか?

  • メモリ: 512×512×3 = 786K ピクセル → 64×64×4 = 16K 値
  • 速度: UNet は 48x より小さいテンソルを処理します
  • 品質: VAE はスマート圧縮を学習しました

5. スケジューラ — ノイズ除去アルゴリズム

from diffusers import (
    DDPMScheduler,        # Original, 1000 steps
    DDIMScheduler,        # Deterministic, 50 steps
    EulerDiscreteScheduler,       # Fast, 20-30 steps
    DPMSolverMultistepScheduler,  # DPM++, 20 steps, high quality
    UniPCMultistepScheduler,      # 10-20 steps
)

# Đổi scheduler dễ dàng
pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)
スケジューラーステップスピード品質使用例
DDPM1000非常に遅い参考資料トレーニング
DDIM50中良い一般
オイラー20-30速いすばらしいデフォルトの SDXL
DPM++ 2M20-25速い素晴らしいおすすめ
ユニPC10-15非常に速い良いリアルタイム

6. 安定した拡散バージョン

バージョン解像度テキストエンコーダリリース
SD1.5512×512クリップ ViT-L/142022年
SD2.1768×768OpenCLIP ViT-H2022年
SDXL1024×1024CLIP ViT-L + OpenCLIP ViT-bigG2023年
SD31024×1024トリプルテキストエンコーダー(CLIP×2 + T5)2024年
フラックス1024×1024T5XXL2024年
# SDXL — recommended cho production
pipe = StableDiffusionXLPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16,
    variant="fp16",
)
pipe.to("cuda")

image = pipe(
    prompt="a majestic lion in a forest, photorealistic",
    negative_prompt="blurry, low quality, distorted",
    num_inference_steps=30,
    guidance_scale=7.5,
    width=1024,
    height=1024,
).images[0]

7. 否定的なプロンプトとパラメータ

# Negative prompt: things to avoid
negative_prompt = "blurry, low quality, deformed, ugly, bad anatomy"

# Key parameters
image = pipe(
    prompt="...",
    negative_prompt=negative_prompt,
    num_inference_steps=30,    # More = better quality, slower
    guidance_scale=7.5,        # Text adherence (7-9 optimal)
    width=1024,
    height=1024,
    seed=42,                   # Reproducibility
).images[0]

概要

コンポーネント役割
CLIP テキスト エンコーダテキスト → 埋め込み (意味論的な意味)
Uネットテキストに基づいたノイズ予測器
VAEピクセル ↔ 潜在空間を圧縮
スケジューラーノイズ除去ステップのアルゴリズム
CFGガイダンス スケールでテキストの遵守度を調整

📌 次の記事: 画像生成のためのプロンプト エンジニアリング — 効果的なプロンプト作成テクニック。