Chuyển đến nội dung chính

レッスン 7: ControlNet と Image-to-Image — 出力の制御

img2img パイプライン: 参照イメージを使用します。 ControlNet: キャニー エッジ、深度マップ、姿勢推定、セグメンテーション。スタイル転送用の IP アダプター。 ComfyUI とディフューザーを使用したレイアウト制御による実践的な画像作成。

🧠 AI と ML — レッスン 6 レッスン 7: ControlNet と Image-to-Image — チェック 出力制御

生成 AI: AI を使用して画像とビデオを作成する

パート 3: 高度な画像生成の実践

xdev.asia

はじめに

テキストから画像への変換は強力ですが、レイアウト、ポーズ、スタイルを正確に制御するのは困難です。 ControlNet はこの問題を解決します。創造性を維持しながら出力を制御するために、空間調整 (鋭いエッジ、深さ、ポーズ) を追加します。


1. イメージからイメージへのパイプライン

from diffusers import StableDiffusionImg2ImgPipeline
from PIL import Image

pipe = StableDiffusionImg2ImgPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16,
)
pipe.to("cuda")

init_image = Image.open("sketch.png").resize((1024, 1024))

image = pipe(
    prompt="a detailed digital painting of a castle",
    image=init_image,
    strength=0.75,      # 0.0 = no change, 1.0 = full generation
    guidance_scale=7.5,
    num_inference_steps=30,
).images[0]

強度パラメータ

strength=0.3  →  Giữ nguyên phần lớn input, chỉ enhance
strength=0.5  →  Balance giữa input và generation
strength=0.75 →  Generate nhiều, giữ composition
strength=1.0  →  Hoàn toàn mới (giống txt2img)

2. ControlNet — 空間制御

from diffusers import StableDiffusionXLControlNetPipeline, ControlNetModel
import cv2
import numpy as np

# Load ControlNet model (canny edge)
controlnet = ControlNetModel.from_pretrained(
    "diffusers/controlnet-canny-sdxl-1.0",
    torch_dtype=torch.float16,
)

pipe = StableDiffusionXLControlNetPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    controlnet=controlnet,
    torch_dtype=torch.float16,
)
pipe.to("cuda")

# Extract canny edges from reference image
image = cv2.imread("reference.jpg")
edges = cv2.Canny(image, 100, 200)
control_image = Image.fromarray(edges)

# Generate with ControlNet
result = pipe(
    prompt="a futuristic city, cyberpunk style",
    image=control_image,
    controlnet_conditioning_scale=0.7,  # strength of control
    num_inference_steps=30,
).images[0]

3. ControlNet の種類

キャニーエッジ

# Detect edges → preserve structure/outlines
edges = cv2.Canny(image, 100, 200)
# Use case: maintain composition, change style

深度マップ

from transformers import pipeline
depth_estimator = pipeline("depth-estimation", model="Intel/dpt-hybrid-midas")
depth = depth_estimator(image)["depth"]
# Use case: maintain 3D structure, change scene

OpenPose (人間のポーズ)

# Detect skeleton → control body pose
from controlnet_aux import OpenposeDetector
openpose = OpenposeDetector.from_pretrained("lllyasviel/ControlNet")
pose = openpose(image)
# Use case: maintain pose, change character appearance

セグメンテーション マップ

# Semantic regions → control spatial layout
from controlnet_aux import SamDetector
sam = SamDetector.from_pretrained("ybelkada/segment-anything")
seg_map = sam(image)
# Use case: control where each element goes

落書き/線画

# Hand-drawn sketches → detailed images
from controlnet_aux import PidiNetDetector
pidi = PidiNetDetector.from_pretrained("lllyasviel/Annotators")
scribble = pidi(image, safe=True)

4. マルチコントロールネット

from diffusers import StableDiffusionXLControlNetPipeline, ControlNetModel

# Combine multiple controlnets
controlnets = [
    ControlNetModel.from_pretrained("controlnet-canny-sdxl-1.0"),
    ControlNetModel.from_pretrained("controlnet-depth-sdxl-1.0"),
]

pipe = StableDiffusionXLControlNetPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    controlnet=controlnets,
    torch_dtype=torch.float16,
)

result = pipe(
    prompt="...",
    image=[canny_image, depth_image],
    controlnet_conditioning_scale=[0.7, 0.5],  # per-controlnet weight
).images[0]

5. IP アダプター — スタイル転送

from diffusers import StableDiffusionXLPipeline
from diffusers.utils import load_image

pipe = StableDiffusionXLPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16,
)
pipe.load_ip_adapter(
    "h94/IP-Adapter",
    subfolder="sdxl_models",
    weight_name="ip-adapter_sdxl.bin"
)

# Style image determines the aesthetic
style_image = load_image("style_reference.jpg")

pipe.set_ip_adapter_scale(0.6)  # strength of style transfer

result = pipe(
    prompt="a cat sitting on a windowsill",
    ip_adapter_image=style_image,
    num_inference_steps=30,
).images[0]

6. ControlNet コンディショニング_スケール ガイド

scale=0.0  →  Ignore control (pure text-to-image)
scale=0.3  →  Subtle guidance, high creativity
scale=0.5  →  Balanced control and creativity
scale=0.7  →  Strong control (recommended)
scale=1.0  →  Strict adherence to control image
scale=1.5+ →  Over-constrained (artifacts)

概要

エンジニアリング入力→出力
img2img参考画像→スタイルバリエーション
コントロールネットキャニーエッジライン → 構造を保持
コントロールネットの深さ深度マップ → 3D レイアウトを維持
コントロールネットのポーズスケルトン → コントロールボディポーズ
IPアダプタースタイルイメージ→トランスファー美学
マルチコントロールネット複数のコントロール → 結合

📌 次の記事: AI によるインペイント、アウトペイント、画像編集。