はじめに
テキストから画像への変換は強力ですが、レイアウト、ポーズ、スタイルを正確に制御するのは困難です。 ControlNet はこの問題を解決します。創造性を維持しながら出力を制御するために、空間調整 (鋭いエッジ、深さ、ポーズ) を追加します。
1. イメージからイメージへのパイプライン
from diffusers import StableDiffusionImg2ImgPipeline
from PIL import Image
pipe = StableDiffusionImg2ImgPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
torch_dtype=torch.float16,
)
pipe.to("cuda")
init_image = Image.open("sketch.png").resize((1024, 1024))
image = pipe(
prompt="a detailed digital painting of a castle",
image=init_image,
strength=0.75, # 0.0 = no change, 1.0 = full generation
guidance_scale=7.5,
num_inference_steps=30,
).images[0]
強度パラメータ
strength=0.3 → Giữ nguyên phần lớn input, chỉ enhance
strength=0.5 → Balance giữa input và generation
strength=0.75 → Generate nhiều, giữ composition
strength=1.0 → Hoàn toàn mới (giống txt2img)
2. ControlNet — 空間制御
from diffusers import StableDiffusionXLControlNetPipeline, ControlNetModel
import cv2
import numpy as np
# Load ControlNet model (canny edge)
controlnet = ControlNetModel.from_pretrained(
"diffusers/controlnet-canny-sdxl-1.0",
torch_dtype=torch.float16,
)
pipe = StableDiffusionXLControlNetPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
controlnet=controlnet,
torch_dtype=torch.float16,
)
pipe.to("cuda")
# Extract canny edges from reference image
image = cv2.imread("reference.jpg")
edges = cv2.Canny(image, 100, 200)
control_image = Image.fromarray(edges)
# Generate with ControlNet
result = pipe(
prompt="a futuristic city, cyberpunk style",
image=control_image,
controlnet_conditioning_scale=0.7, # strength of control
num_inference_steps=30,
).images[0]
3. ControlNet の種類
キャニーエッジ
# Detect edges → preserve structure/outlines
edges = cv2.Canny(image, 100, 200)
# Use case: maintain composition, change style
深度マップ
from transformers import pipeline
depth_estimator = pipeline("depth-estimation", model="Intel/dpt-hybrid-midas")
depth = depth_estimator(image)["depth"]
# Use case: maintain 3D structure, change scene
OpenPose (人間のポーズ)
# Detect skeleton → control body pose
from controlnet_aux import OpenposeDetector
openpose = OpenposeDetector.from_pretrained("lllyasviel/ControlNet")
pose = openpose(image)
# Use case: maintain pose, change character appearance
セグメンテーション マップ
# Semantic regions → control spatial layout
from controlnet_aux import SamDetector
sam = SamDetector.from_pretrained("ybelkada/segment-anything")
seg_map = sam(image)
# Use case: control where each element goes
落書き/線画
# Hand-drawn sketches → detailed images
from controlnet_aux import PidiNetDetector
pidi = PidiNetDetector.from_pretrained("lllyasviel/Annotators")
scribble = pidi(image, safe=True)
4. マルチコントロールネット
from diffusers import StableDiffusionXLControlNetPipeline, ControlNetModel
# Combine multiple controlnets
controlnets = [
ControlNetModel.from_pretrained("controlnet-canny-sdxl-1.0"),
ControlNetModel.from_pretrained("controlnet-depth-sdxl-1.0"),
]
pipe = StableDiffusionXLControlNetPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
controlnet=controlnets,
torch_dtype=torch.float16,
)
result = pipe(
prompt="...",
image=[canny_image, depth_image],
controlnet_conditioning_scale=[0.7, 0.5], # per-controlnet weight
).images[0]
5. IP アダプター — スタイル転送
from diffusers import StableDiffusionXLPipeline
from diffusers.utils import load_image
pipe = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
torch_dtype=torch.float16,
)
pipe.load_ip_adapter(
"h94/IP-Adapter",
subfolder="sdxl_models",
weight_name="ip-adapter_sdxl.bin"
)
# Style image determines the aesthetic
style_image = load_image("style_reference.jpg")
pipe.set_ip_adapter_scale(0.6) # strength of style transfer
result = pipe(
prompt="a cat sitting on a windowsill",
ip_adapter_image=style_image,
num_inference_steps=30,
).images[0]
6. ControlNet コンディショニング_スケール ガイド
scale=0.0 → Ignore control (pure text-to-image)
scale=0.3 → Subtle guidance, high creativity
scale=0.5 → Balanced control and creativity
scale=0.7 → Strong control (recommended)
scale=1.0 → Strict adherence to control image
scale=1.5+ → Over-constrained (artifacts)
概要
| エンジニアリング | 入力→出力 |
|---|---|
| img2img | 参考画像→スタイルバリエーション |
| コントロールネットキャニー | エッジライン → 構造を保持 |
| コントロールネットの深さ | 深度マップ → 3D レイアウトを維持 |
| コントロールネットのポーズ | スケルトン → コントロールボディポーズ |
| IPアダプター | スタイルイメージ→トランスファー美学 |
| マルチコントロールネット | 複数のコントロール → 結合 |
📌 次の記事: AI によるインペイント、アウトペイント、画像編集。