Chuyển đến nội dung chính

レッスン 6: 画像生成のための迅速なエンジニアリング

テキスト プロンプトの構造: 件名、スタイル、品質、否定的なプロンプト。構文の重み付けと強調を促します。安定した拡散のプロンプトのヒント。旅の途中でのプロンプトパターン。 DALL-E を促す戦略。

🧠 AI と ML — レッスン 5 レッスン 6: 画像の迅速なエンジニアリング 世代

生成 AI: AI を使用して画像とビデオを作成する

パート 2: 拡散モデル — 革新的なイメージの作成

xdev.asia

はじめに

画像生成のためのプロンプト エンジニアリングは、出力を制御するための効果的なテキスト プロンプトを作成する技術です。 LLM プロンプトとは異なり、画像プロンプトは、主題、スタイル、照明、構成、カメラ アングルといった 視覚的な説明 に重点を置いています。


1. 効果的なプロンプトの構造

[Subject] + [Medium/Style] + [Details] + [Quality] + [Lighting] + [Camera]

Ví dụ:
"A samurai warrior standing in a bamboo forest,
 digital painting, intricate armor details,
 4k, highly detailed, dramatic lighting,
 cinematic composition, wide angle shot"

テンプレート

コンポーネント例役割
件名「猫」「城」本題
アクション「走る」、「座る」アクション
設定「森の中」、「火星で」背景
スタイル「油絵」「アニメ」スタイル
品質「4K」、「高精細」品質
照明「ゴールデンアワー」「ネオン」ライト
カメラ「クローズアップ」、「広角」撮影角度
気分「平和的」、「ドラマチック」感情

2. スタイルキーワード

Photography styles:
  "portrait photography, 85mm lens, shallow depth of field, bokeh"
  "street photography, candid, natural light, grainy film"
  "macro photography, extreme close-up, water droplets"

Art styles:
  "oil painting, impressionist, thick brushstrokes"
  "watercolor, soft edges, pastel colors"
  "digital art, concept art, artstation"
  "anime style, studio ghibli, cel shading"
  "pixel art, 16-bit, retro gaming"

3D/Rendering:
  "3D render, octane render, unreal engine 5"
  "isometric, low poly, miniature"
  "photorealistic, ray tracing, cinema 4D"

3. 否定的なプロンプト

# Negative prompts: what to AVOID in generation

# Universal negative prompt
negative = """
blurry, low quality, low resolution, deformed, distorted,
bad anatomy, bad proportions, extra limbs, extra fingers,
ugly, poorly drawn, watermark, text, signature,
cropped, out of frame, worst quality, jpeg artifacts
"""

# Photo-specific negative
photo_negative = """
cartoon, anime, illustration, painting, drawing,
3d render, CGI, artificial, fake looking
"""

# Art-specific negative
art_negative = """
photograph, photo, realistic, 3d render,
blurry, low effort, amateur
"""

4. プロンプトの重み付け

# Hugging Face Diffusers: compel library
from compel import Compel

compel = Compel(tokenizer=pipe.tokenizer, text_encoder=pipe.text_encoder)

# Tăng weight: ++ hoặc (word)1.5
prompt = "a (beautiful)1.3 sunset over the (ocean)1.5, dramatic clouds"
conditioning = compel(prompt)

# Giảm weight
prompt = "a cat sitting on a (table)0.5"  # less emphasis on table

# Blend prompts
prompt = '"a cat".blend("a dog", 0.7, 0.3)'  # 70% cat, 30% dog

A1111 構文 (WebUI)

# Tăng emphasis: (word:1.3) hoặc ((word))
a ((beautiful)) sunset over the (ocean:1.5)

# Giảm emphasis: [word] hoặc (word:0.7)
a cat sitting on a [table]

# Alternate: [word1|word2] — switch mỗi step
a [cat|dog] sitting in a garden

5. 実際のプロンプトパターン

キャラクターデザイン

"Character concept art of a female cyberpunk hacker,
neon blue hair, glowing circuit tattoos on arms,
wearing a black leather jacket with LED strips,
standing in a rainy Tokyo alley at night,
neon signs reflecting in puddles,
digital art, artstation trending, 8k, highly detailed"

製品写真

"Professional product photography of a luxury watch,
silver metallic case, black leather strap,
on a dark marble surface, studio lighting,
soft shadows, shallow depth of field,
commercial photography, 4k, Canon EOS R5"

アーキテクチャ

"Modern minimalist house, floor to ceiling windows,
concrete and wood materials, infinity pool,
surrounded by tropical garden, golden hour lighting,
architectural photography, wide angle, drone shot"

ファンタジー/ゲームアート

"Epic fantasy landscape, floating islands in the sky,
waterfalls cascading into clouds below,
ancient ruins with glowing runes,
dragon flying in the distance,
volumetric lighting, god rays, matte painting,
concept art, 4k, trending on artstation"

6. 系統的な即時テスト

# Grid search cho prompt optimization
subjects = ["a cat", "a robot", "a dragon"]
styles = ["oil painting", "digital art", "photograph"]
lightings = ["golden hour", "studio", "neon"]

for subject in subjects:
    for style in styles:
        for lighting in lightings:
            prompt = f"{subject}, {style}, {lighting} lighting, 4k"
            image = pipe(prompt, num_inference_steps=25).images[0]
            image.save(f"grid_{subject}_{style}_{lighting}.png")

7. ヒントとベストプラクティス

✅ DO:
- Đặt subject ở đầu prompt
- Sử dụng comma để phân tách concepts
- Thêm quality modifiers: "4k, highly detailed, sharp"
- Dùng negative prompts để loại bỏ artifacts
- Test nhiều seeds cho cùng prompt
- Tham khảo community prompts (Civitai, PromptHero)

❌ DON'T:
- Viết câu quá dài phức tạp (CLIP max 77 tokens)
- Dùng từ mơ hồ: "nice", "cool", "good"
- Mâu thuẫn: "realistic anime style"
- Bỏ qua negative prompts
- Kỳ vọng chính xác: AI interpret, không follow literal

概要

エンジニアリング説明
プロンプト解剖主題 + スタイル + 詳細 + 品質
否定的なプロンプト不要な要素を削除する
即時重み付け各パートの強調を増減
スタイルキーワード写真、アート、3D レンダリング スタイル
体系的なテストグリッド検索パラメータ

📌 次の投稿: ControlNet と Image-to-Image — 参照画像を使用して出力を制御します。