簡介
圖像生成提示工程是編寫有效文字提示來控制輸出的藝術。與 LLM 提示不同,影像提示著重於視覺描述——主體、風格、燈光、構圖、拍攝角度。
1. 有效提示的剖析
[Subject] + [Medium/Style] + [Details] + [Quality] + [Lighting] + [Camera]
Ví dụ:
"A samurai warrior standing in a bamboo forest,
digital painting, intricate armor details,
4k, highly detailed, dramatic lighting,
cinematic composition, wide angle shot"
模板
| 組件 | 範例 | 角色 |
|---|---|---|
| 主題 | 「一隻貓」、「一座城堡」 | 主題 |
| 行動 | 「跑步」、「坐著」 | 行動 |
| 設定 | 「在森林裡」、「在火星上」 | 背景 |
| 風格 | 「油畫」、「動漫」 | 風格 |
| 品質 | “4k”、“非常詳細” | 品質 |
| 燈光 | 「黃金時刻」、「霓虹燈」 | 光 |
| 相機 | “特寫”、“廣角” | 拍攝角度 |
| 心情 | 「和平」、「戲劇性」 | 情感 |
2. 風格關鍵字
Photography styles:
"portrait photography, 85mm lens, shallow depth of field, bokeh"
"street photography, candid, natural light, grainy film"
"macro photography, extreme close-up, water droplets"
Art styles:
"oil painting, impressionist, thick brushstrokes"
"watercolor, soft edges, pastel colors"
"digital art, concept art, artstation"
"anime style, studio ghibli, cel shading"
"pixel art, 16-bit, retro gaming"
3D/Rendering:
"3D render, octane render, unreal engine 5"
"isometric, low poly, miniature"
"photorealistic, ray tracing, cinema 4D"
3. 負麵提示
# Negative prompts: what to AVOID in generation
# Universal negative prompt
negative = """
blurry, low quality, low resolution, deformed, distorted,
bad anatomy, bad proportions, extra limbs, extra fingers,
ugly, poorly drawn, watermark, text, signature,
cropped, out of frame, worst quality, jpeg artifacts
"""
# Photo-specific negative
photo_negative = """
cartoon, anime, illustration, painting, drawing,
3d render, CGI, artificial, fake looking
"""
# Art-specific negative
art_negative = """
photograph, photo, realistic, 3d render,
blurry, low effort, amateur
"""
4. 提示加權
# Hugging Face Diffusers: compel library
from compel import Compel
compel = Compel(tokenizer=pipe.tokenizer, text_encoder=pipe.text_encoder)
# Tăng weight: ++ hoặc (word)1.5
prompt = "a (beautiful)1.3 sunset over the (ocean)1.5, dramatic clouds"
conditioning = compel(prompt)
# Giảm weight
prompt = "a cat sitting on a (table)0.5" # less emphasis on table
# Blend prompts
prompt = '"a cat".blend("a dog", 0.7, 0.3)' # 70% cat, 30% dog
A1111 文法 (WebUI)
# Tăng emphasis: (word:1.3) hoặc ((word))
a ((beautiful)) sunset over the (ocean:1.5)
# Giảm emphasis: [word] hoặc (word:0.7)
a cat sitting on a [table]
# Alternate: [word1|word2] — switch mỗi step
a [cat|dog] sitting in a garden
5. 實際提示模式
角色設計
"Character concept art of a female cyberpunk hacker,
neon blue hair, glowing circuit tattoos on arms,
wearing a black leather jacket with LED strips,
standing in a rainy Tokyo alley at night,
neon signs reflecting in puddles,
digital art, artstation trending, 8k, highly detailed"
產品攝影
"Professional product photography of a luxury watch,
silver metallic case, black leather strap,
on a dark marble surface, studio lighting,
soft shadows, shallow depth of field,
commercial photography, 4k, Canon EOS R5"
架構
"Modern minimalist house, floor to ceiling windows,
concrete and wood materials, infinity pool,
surrounded by tropical garden, golden hour lighting,
architectural photography, wide angle, drone shot"
幻想/遊戲藝術
"Epic fantasy landscape, floating islands in the sky,
waterfalls cascading into clouds below,
ancient ruins with glowing runes,
dragon flying in the distance,
volumetric lighting, god rays, matte painting,
concept art, 4k, trending on artstation"
6. 系統的即時測試
# Grid search cho prompt optimization
subjects = ["a cat", "a robot", "a dragon"]
styles = ["oil painting", "digital art", "photograph"]
lightings = ["golden hour", "studio", "neon"]
for subject in subjects:
for style in styles:
for lighting in lightings:
prompt = f"{subject}, {style}, {lighting} lighting, 4k"
image = pipe(prompt, num_inference_steps=25).images[0]
image.save(f"grid_{subject}_{style}_{lighting}.png")
7. 提示與最佳實踐
✅ DO:
- Đặt subject ở đầu prompt
- Sử dụng comma để phân tách concepts
- Thêm quality modifiers: "4k, highly detailed, sharp"
- Dùng negative prompts để loại bỏ artifacts
- Test nhiều seeds cho cùng prompt
- Tham khảo community prompts (Civitai, PromptHero)
❌ DON'T:
- Viết câu quá dài phức tạp (CLIP max 77 tokens)
- Dùng từ mơ hồ: "nice", "cool", "good"
- Mâu thuẫn: "realistic anime style"
- Bỏ qua negative prompts
- Kỳ vọng chính xác: AI interpret, không follow literal
總結
| 工程 | 描述 |
|---|---|
| 提示解剖學 | 主題+風格+細節+品質 |
| 負面提示 | 刪除不需要的元素 |
| 及時稱重 | 增加/減少每個部分的重點 |
| 風格關鍵字 | 攝影、藝術、3D 渲染風格 |
| 系統測試 | 網格搜尋參數 |
📌 下一篇文章: ControlNet 與影像到影像 — 使用參考影像控制輸出。