Chuyển đến nội dung chính

レッスン 9: LoRA とカスタム モデルのトレーニング — 独自のスタイルの作成

LoRA による安定拡散の微調整: 概念、計算、実装。 DreamBooth: パーソナライズされた世代。テキストの反転。トレーニング データセットの準備とベスト プラクティス。 LoRA モデルをマージします。

🧠 AI と ML — レッスン 8 レッスン 9: LoRA とカスタム モデル トレーニング — 作成 パーソナルスタイル

生成 AI: AI を使用して画像とビデオを作成する

パート 3: 高度な画像生成の実践

xdev.asia

はじめに

事前トレーニングされた安定拡散は優れた画像を作成しますが、ユニークなスタイル や 特定のコンセプト (ブランド、キャラクター、製品) が欠けています。 LoRA (低ランク適応) では、コンシューマ GPU でトレーニングされたわずか 10 ~ 20 枚の画像を使用してモデルを微調整でき、出力ファイルはわずか 10 ~ 100 MB です。


1. LoRA — 低ランクの適応

Full fine-tuning: update ALL parameters (hàng tỷ) → đắt, cần nhiều data
LoRA: chỉ thêm low-rank matrices nhỏ → rẻ, ít data, kết quả tốt

Toán: W' = W + ΔW = W + BA
- W: original weight matrix (frozen)
- B: low-rank matrix (r × d), trainable
- A: low-rank matrix (d × r), trainable
- r << d (rank 4-128, thường 8-32)

$$W' = W + \alpha \cdot BA$$

B と A のみをトレーニング → パラメーターが $d^2$ から $2dr$ に減少しました (99% 以上減少)。


2. データセットの準備

Yêu cầu:
- 10-30 ảnh high quality cho subject/style
- Consistent quality và resolution
- Đa dạng góc, lighting, background (cho subject)
- Uniform style (cho style LoRA)

Cấu trúc folder:
dataset/
├── image_001.png    # 768x768 hoặc 1024x1024
├── image_001.txt    # caption: "a photo of sks person, smiling"
├── image_002.png
├── image_002.txt    # caption: "a photo of sks person, side view"
└── ...

自動キャプション

from transformers import BlipForConditionalGeneration, BlipProcessor

processor = BlipProcessor.from_pretrained("Salesforce/blip-image-captioning-large")
model = BlipForConditionalGeneration.from_pretrained("Salesforce/blip-image-captioning-large")

def caption_image(image_path, trigger_word="sks"):
    image = Image.open(image_path)
    inputs = processor(image, return_tensors="pt")
    output = model.generate(**inputs, max_length=50)
    caption = processor.decode(output[0], skip_special_tokens=True)
    # Prepend trigger word
    return f"a photo of {trigger_word}, {caption}"

3. ディフューザーを使用した LoRA のトレーニング

# Install training dependencies
pip install peft accelerate bitsandbytes

# Training script
accelerate launch train_dreambooth_lora_sdxl.py \
  --pretrained_model_name_or_path="stabilityai/stable-diffusion-xl-base-1.0" \
  --instance_data_dir="./dataset" \
  --instance_prompt="a photo of sks dog" \
  --output_dir="./lora_output" \
  --resolution=1024 \
  --train_batch_size=1 \
  --gradient_accumulation_steps=4 \
  --learning_rate=1e-4 \
  --lr_scheduler="cosine" \
  --lr_warmup_steps=100 \
  --max_train_steps=1000 \
  --rank=32 \
  --mixed_precision="fp16" \
  --seed=42

トレーニングパラメータガイド

パラメータおすすめメモ
ランク8-32高い = より多くの容量、より多くの VRAM
学習率1e-5 から 1e-4低く開始し、適合が不十分な場合は増加します
最大トレインステップ数500-2000より多くのデータ → より多くのステップ
解像度1024基本モデルの解像度を一致させる
トレイン_バッチ_サイズ1-4VRAM に依存

4. DreamBooth — パーソナライズされた生成

# DreamBooth concept: fine-tune the model to learn a specific subject
# Uses a rare trigger word (e.g., "sks") to represent the concept

# Training data:
# - 5-30 images of YOUR specific subject
# - Caption: "a photo of sks [class]" (e.g., "a photo of sks dog")

# Class images (regularization):
# - Generated images of the general class
# - Prevents model from forgetting the class concept
# - "a photo of dog" (without sks)
# DreamBooth + LoRA training
accelerate launch train_dreambooth_lora_sdxl.py \
  --instance_data_dir="./my_dog_photos" \
  --instance_prompt="a photo of sks dog" \
  --class_data_dir="./dog_class_images" \
  --class_prompt="a photo of dog" \
  --num_class_images=200 \
  --with_prior_preservation \
  --prior_loss_weight=1.0 \
  --max_train_steps=800

5. テキストの反転

# Concept: Học một embedding vector mới cho concept
# Không thay đổi model weights → chỉ thêm 1 token

# Train: 3-10 images → learn embedding cho <my-concept>
# Use: "a painting in the style of <my-concept>"

# Ưu điểm: rất nhỏ (vài KB), không ảnh hưởng model
# Nhược điểm: ít expressive hơn LoRA

from diffusers import StableDiffusionPipeline

pipe = StableDiffusionPipeline.from_pretrained("runwayml/stable-diffusion-v1-5")
pipe.load_textual_inversion("path/to/embedding.safetensors", token="<my-style>")

image = pipe("a landscape in the style of <my-style>").images[0]

6. LoRA をロードして使用する

from diffusers import StableDiffusionXLPipeline
import torch

pipe = StableDiffusionXLPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16,
)
pipe.to("cuda")

# Load LoRA weights
pipe.load_lora_weights("./lora_output", weight_name="pytorch_lora_weights.safetensors")

# Adjust LoRA strength
pipe.fuse_lora(lora_scale=0.8)  # 0.0-1.0

# Generate with LoRA
image = pipe(
    prompt="a photo of sks dog wearing a crown, royal portrait",
    num_inference_steps=30,
    guidance_scale=7.5,
).images[0]

# Unload LoRA
pipe.unfuse_lora()
pipe.unload_lora_weights()

7. 複数の LoRA をマージする

# Combine style LoRA + character LoRA
pipe.load_lora_weights("style_lora.safetensors", adapter_name="style")
pipe.load_lora_weights("character_lora.safetensors", adapter_name="character")

pipe.set_adapters(["style", "character"], adapter_weights=[0.7, 0.9])

image = pipe(
    prompt="sks person in anime style, bright colors",
    num_inference_steps=30,
).images[0]

8. ベストプラクティス

Dataset:
✅ High quality, consistent resolution
✅ Variety in poses/angles (cho subject)
✅ Clear, noise-free images
✅ Good captions with trigger word
❌ Blurry, low-res images
❌ Watermarked images
❌ Too few images (< 5)

Training:
✅ Start with low learning rate
✅ Use cosine scheduler
✅ Save checkpoints frequently
✅ Compare different ranks (8, 16, 32)
✅ Train 500-1500 steps cho LoRA
❌ Overtrain (> 3000 steps usually)
❌ Too high rank (> 64) without enough data

概要

方法パラメータファイルサイズ品質必要なデータ
フル微調整~1B~6GBベスト1000 枚以上の画像
ロラ~1-50M10~100MBすばらしい画像 10 ~ 30 枚
ドリームブース~1B~6GBすばらしい画像5~30枚
ドリームブース+LoRA~1-50M10~100MBすばらしい画像5~30枚
テキストの反転1トークン~4KB良い画像 3 ~ 10 枚

📌 次の記事: DALL-E 3 API — OpenAI Image Generation をアプリケーションに統合します。