Chuyển đến nội dung chính

第 6 課:影像生成的快速工程

文字提示剖析:主題、風格、品質、負面提示。提示權重和強調語法。穩定擴散提示提示。中途提示模式。 DALL-E 提示策略。

🧠 人工智慧與機器學習 — 第 5 課 第 6 課:影像快速工程 世代

生成式 AI:使用 AI 創建圖像和視頻

第 2 部分:擴散模型 — 革命性的影像創建

亞洲開發網

簡介

圖像生成提示工程是編寫有效文字提示來控制輸出的藝術。與 LLM 提示不同,影像提示著重於視覺描述——主體、風格、燈光、構圖、拍攝角度。


1. 有效提示的剖析

[Subject] + [Medium/Style] + [Details] + [Quality] + [Lighting] + [Camera]

Ví dụ:
"A samurai warrior standing in a bamboo forest,
 digital painting, intricate armor details,
 4k, highly detailed, dramatic lighting,
 cinematic composition, wide angle shot"

模板

組件範例角色
主題「一隻貓」、「一座城堡」主題
行動「跑步」、「坐著」行動
設定「在森林裡」、「在火星上」背景
風格「油畫」、「動漫」風格
品質“4k”、“非常詳細”品質
燈光「黃金時刻」、「霓虹燈」光
相機“特寫”、“廣角”拍攝角度
心情「和平」、「戲劇性」情感

2. 風格關鍵字

Photography styles:
  "portrait photography, 85mm lens, shallow depth of field, bokeh"
  "street photography, candid, natural light, grainy film"
  "macro photography, extreme close-up, water droplets"

Art styles:
  "oil painting, impressionist, thick brushstrokes"
  "watercolor, soft edges, pastel colors"
  "digital art, concept art, artstation"
  "anime style, studio ghibli, cel shading"
  "pixel art, 16-bit, retro gaming"

3D/Rendering:
  "3D render, octane render, unreal engine 5"
  "isometric, low poly, miniature"
  "photorealistic, ray tracing, cinema 4D"

3. 負麵提示

# Negative prompts: what to AVOID in generation

# Universal negative prompt
negative = """
blurry, low quality, low resolution, deformed, distorted,
bad anatomy, bad proportions, extra limbs, extra fingers,
ugly, poorly drawn, watermark, text, signature,
cropped, out of frame, worst quality, jpeg artifacts
"""

# Photo-specific negative
photo_negative = """
cartoon, anime, illustration, painting, drawing,
3d render, CGI, artificial, fake looking
"""

# Art-specific negative
art_negative = """
photograph, photo, realistic, 3d render,
blurry, low effort, amateur
"""

4. 提示加權

# Hugging Face Diffusers: compel library
from compel import Compel

compel = Compel(tokenizer=pipe.tokenizer, text_encoder=pipe.text_encoder)

# Tăng weight: ++ hoặc (word)1.5
prompt = "a (beautiful)1.3 sunset over the (ocean)1.5, dramatic clouds"
conditioning = compel(prompt)

# Giảm weight
prompt = "a cat sitting on a (table)0.5"  # less emphasis on table

# Blend prompts
prompt = '"a cat".blend("a dog", 0.7, 0.3)'  # 70% cat, 30% dog

A1111 文法 (WebUI)

# Tăng emphasis: (word:1.3) hoặc ((word))
a ((beautiful)) sunset over the (ocean:1.5)

# Giảm emphasis: [word] hoặc (word:0.7)
a cat sitting on a [table]

# Alternate: [word1|word2] — switch mỗi step
a [cat|dog] sitting in a garden

5. 實際提示模式

角色設計

"Character concept art of a female cyberpunk hacker,
neon blue hair, glowing circuit tattoos on arms,
wearing a black leather jacket with LED strips,
standing in a rainy Tokyo alley at night,
neon signs reflecting in puddles,
digital art, artstation trending, 8k, highly detailed"

產品攝影

"Professional product photography of a luxury watch,
silver metallic case, black leather strap,
on a dark marble surface, studio lighting,
soft shadows, shallow depth of field,
commercial photography, 4k, Canon EOS R5"

架構

"Modern minimalist house, floor to ceiling windows,
concrete and wood materials, infinity pool,
surrounded by tropical garden, golden hour lighting,
architectural photography, wide angle, drone shot"

幻想/遊戲藝術

"Epic fantasy landscape, floating islands in the sky,
waterfalls cascading into clouds below,
ancient ruins with glowing runes,
dragon flying in the distance,
volumetric lighting, god rays, matte painting,
concept art, 4k, trending on artstation"

6. 系統的即時測試

# Grid search cho prompt optimization
subjects = ["a cat", "a robot", "a dragon"]
styles = ["oil painting", "digital art", "photograph"]
lightings = ["golden hour", "studio", "neon"]

for subject in subjects:
    for style in styles:
        for lighting in lightings:
            prompt = f"{subject}, {style}, {lighting} lighting, 4k"
            image = pipe(prompt, num_inference_steps=25).images[0]
            image.save(f"grid_{subject}_{style}_{lighting}.png")

7. 提示與最佳實踐

✅ DO:
- Đặt subject ở đầu prompt
- Sử dụng comma để phân tách concepts
- Thêm quality modifiers: "4k, highly detailed, sharp"
- Dùng negative prompts để loại bỏ artifacts
- Test nhiều seeds cho cùng prompt
- Tham khảo community prompts (Civitai, PromptHero)

❌ DON'T:
- Viết câu quá dài phức tạp (CLIP max 77 tokens)
- Dùng từ mơ hồ: "nice", "cool", "good"
- Mâu thuẫn: "realistic anime style"
- Bỏ qua negative prompts
- Kỳ vọng chính xác: AI interpret, không follow literal

總結

工程描述
提示解剖學主題+風格+細節+品質
負面提示刪除不需要的元素
及時稱重增加/減少每個部分的重點
風格關鍵字攝影、藝術、3D 渲染風格
系統測試網格搜尋參數

📌 下一篇文章: ControlNet 與影像到影像 — 使用參考影像控制輸出。