Chuyển đến nội dung chính

レッスン 9: 画像の生成と安定した拡散

理論的拡散モデル: 順方向/逆方向プロセス。安定した拡散アーキテクチャ: VAE、U-Net、CLIP テキスト エンコーダー。ハンズオン: テキストから画像へ、画像から画像へ、ControlNet、スタイル用の LoRA。

🧠 AI と ML — レッスン 8 レッスン 9: 画像の生成と安定した拡散

深層学習によるコンピューター ビジョン: CNN から Vision Transformer まで

パート 3: セグメンテーションと最新の履歴書

xdev.asia

はじめに

DALL·E、Midjourney から 安定した普及まで — テキストから画像を作成する AI はクリエイティブ業界を変えました。この記事では、拡散モデル: その仕組み、安定した拡散アーキテクチャ、およびテキストから画像への変換、画像から画像への実践、ControlNet について詳しく説明します。

🎯 学ぶ理由 拡散モデルを理解すると、イメージの作成、編集、CV パイプライン (合成データ、拡張、修復) への適用に役立ちます。


1. 拡散モデル — 理論

1.1 中心となるアイデア

Forward Process (Thêm noise):
Ảnh gốc → Ảnh hơi noise → ... → Noise thuần túy (Gaussian)
  x_0        x_1                      x_T

Reverse Process (Khử noise):
Noise → Bớt noise → ... → Ảnh sạch!
 x_T      x_{T-1}           x_0

Model học: "Cho 1 ảnh có noise, predict noise đã thêm vào"
→ Trừ noise đi = ảnh sạch hơn
→ Lặp lại T bước = ảnh hoàn chỉnh

1.2 拡散が GAN より優れているのはなぜですか?

ガン拡散
トレーニング難しい(モード崩壊、不安定)安定
多様性低 (モード崩壊)高 (可変出力)
品質良いですが、人工物です。素晴らしい
コントロールコントロールが難しいテキスト プロンプト、ControlNet
速度高速 (前方パス 1 回)遅い (反復)

2. 安定した普及 — アーキテクチャ

2.1 3 つの主要コンポーネント

Text Prompt: "A cat sitting on a beach at sunset"
    ↓
┌──────────────────┐
│ CLIP Text Encoder│  Biến text → embedding vector
│ (Frozen)         │
└──────────────────┘
    ↓ text embedding
┌──────────────────┐
│ U-Net            │  Predict noise trong LATENT SPACE
│ (Trained)        │  (không phải pixel space!)
│ + Cross-Attention│  (text embedding guide denoising)
└──────────────────┘
    ↓ denoised latent
┌──────────────────┐
│ VAE Decoder      │  Latent → Pixel image (upscale 8×)
│ (Frozen)         │
└──────────────────┘
    ↓
Output: 512×512 image

2.2 潜在空間 — スピードの秘密

Pixel Space:  512×512×3  = 786,432 numbers → Chậm!
Latent Space: 64×64×4    = 16,384 numbers  → Nhanh 48×!

VAE Encoder: Image (512×512) → Latent (64×64)
U-Net: Denoise in Latent Space
VAE Decoder: Latent (64×64) → Image (512×512)

3. 実践: テキストから画像への変換

3.1 セットアップ

pip install diffusers transformers accelerate torch

3.2 基本的なテキストから画像への変換

"""Text-to-Image với Stable Diffusion"""
import torch
from diffusers import StableDiffusionPipeline

# Load model (tải lần đầu ~5GB)
pipe = StableDiffusionPipeline.from_pretrained(
    "stabilityai/stable-diffusion-2-1",
    torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")

# Generate
prompt = "A majestic lion wearing a crown, digital art, 4k, highly detailed"
negative_prompt = "blurry, low quality, deformed"

image = pipe(
    prompt=prompt,
    negative_prompt=negative_prompt,
    num_inference_steps=30,    # Số bước denoise (20-50)
    guidance_scale=7.5,        # CFG: prompt adherence (5-15)
    width=768,
    height=768,
).images[0]

image.save("lion_king.png")
image.show()

3.3 SDXL — 安定した拡散 XL

"""SDXL: chất lượng cao hơn, 1024×1024"""
from diffusers import StableDiffusionXLPipeline

pipe = StableDiffusionXLPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16,
    variant="fp16",
)
pipe = pipe.to("cuda")

# Tối ưu memory
pipe.enable_model_cpu_offload()

image = pipe(
    prompt="A Vietnamese phở restaurant in cyberpunk style, neon lights, rain",
    negative_prompt="ugly, blurry, low quality",
    num_inference_steps=40,
    guidance_scale=8.0,
).images[0]

image.save("cyberpunk_pho.png")

3.4 重要なパラメータ

パラメータ値影響
num_inference_steps20-50より多く = より詳細、より遅く
guidance_scale (CFG)1-20高 = 粘着性が高くなりますが、自然さは劣ります
seed整数同じシード = 同じ結果 (再現可能)
negative_promptテキスト写真に写りたくないもの
# Reproducible generation
generator = torch.Generator(device="cuda").manual_seed(42)
image = pipe(prompt="...", generator=generator).images[0]

4. 画像から画像へ

"""Image-to-Image: biến đổi ảnh có sẵn"""
from diffusers import StableDiffusionImg2ImgPipeline
from PIL import Image

pipe = StableDiffusionImg2ImgPipeline.from_pretrained(
    "stabilityai/stable-diffusion-2-1",
    torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")

# Load ảnh gốc
init_image = Image.open("sketch.png").resize((768, 768))

# Transform
result = pipe(
    prompt="A detailed watercolor painting of a village",
    image=init_image,
    strength=0.75,     # 0 = giữ nguyên, 1 = tạo mới hoàn toàn
    guidance_scale=7.5,
    num_inference_steps=30,
).images[0]

result.save("watercolor_village.png")

5. ControlNet — 正確な制御

"""ControlNet: điều khiển layout ảnh bằng edge/pose/depth"""
from diffusers import StableDiffusionControlNetPipeline, ControlNetModel
import cv2
import numpy as np

# Load ControlNet (Canny edge)
controlnet = ControlNetModel.from_pretrained(
    "lllyasviel/control_v11p_sd15_canny",
    torch_dtype=torch.float16,
)

pipe = StableDiffusionControlNetPipeline.from_pretrained(
    "runwayml/stable-diffusion-v1-5",
    controlnet=controlnet,
    torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")

# Tạo Canny edge map từ ảnh gốc
image = np.array(Image.open("room_photo.jpg"))
edges = cv2.Canny(image, 100, 200)
canny_image = Image.fromarray(edges)

# Generate với cùng layout nhưng style khác
result = pipe(
    prompt="A modern minimalist living room, interior design magazine",
    image=canny_image,
    num_inference_steps=30,
).images[0]

result.save("modern_room.png")

6. LoRA — ライト微調整スタイル

"""LoRA: thêm style mới cho Stable Diffusion"""
from diffusers import StableDiffusionPipeline

pipe = StableDiffusionPipeline.from_pretrained(
    "runwayml/stable-diffusion-v1-5",
    torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")

# Load LoRA weights (ví dụ: anime style)
pipe.load_lora_weights("path/to/anime_lora.safetensors")

# Generate với style mới
image = pipe(
    prompt="1girl, cherry blossoms, anime style",
    num_inference_steps=30,
).images[0]

# Unload LoRA
pipe.unload_lora_weights()

7. 修復 — 画像の一部を編集する

"""Inpainting: sửa/thay thế 1 vùng trong ảnh"""
from diffusers import StableDiffusionInpaintPipeline
from PIL import Image

pipe = StableDiffusionInpaintPipeline.from_pretrained(
    "stabilityai/stable-diffusion-2-inpainting",
    torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")

# Ảnh gốc + mask (vùng trắng = vùng cần sửa)
image = Image.open("photo.jpg").resize((512, 512))
mask = Image.open("mask.png").resize((512, 512))  # Trắng = replace

result = pipe(
    prompt="A beautiful garden with flowers",
    image=image,
    mask_image=mask,
    num_inference_steps=30,
).images[0]

8. CV パイプラインでのアプリケーション

🎨 Synthetic Data:     Tạo training data cho rare cases
🔄 Data Augmentation:  Stable Diffusion biến thể ảnh
🖼️ Inpainting:        Xóa watermark, sửa defects
📐 ControlNet:         Maintain layout, thay đổi style
🎭 Style Transfer:     Chuyển đổi phong cách ảnh
📸 Super Resolution:   Upscale ảnh chất lượng thấp

概要

コンセプト覚えておいてください
拡散ノイズを追加する → ノイズを除去する方法を学ぶ → 生成
潜在空間圧縮空間(64×64)での処理 → 高速化
CFG スケールガイダンス: 高 = より迅速
コントロールネットエッジ/ポーズ/深度を使用してレイアウトを制御
LoRA新しいスタイルのための光の微調整
修復写真の一部を編集し、残りはそのまま残します

一般的な演習

  1. Text-to-Image: 同じプロンプトでシードが異なる 10 個の画像を生成します。どの写真が一番美しいですか?
  2. スタイル転送: img2img を使用して、写真を油絵、水彩、アニメに変換します。
  3. ControlNet: 部屋から Canny Edge を使用 → 別のスタイルで部屋を生成します。
  4. 合成データ: SD を使用して、トレーニング データセット用に 50 個の「製品欠陥」画像を作成します。

次の記事: Vision Transformer (ViT) & CLIP — テキストと画像を接続する、画像用のトランスフォーマー。