Introduction
Prompt Engineering for image generation is the art of writing effective text prompts to control output. Unlike LLM prompts, image prompts focus on visual description — subject, style, lighting, composition, camera angle.
1. Anatomy of an effective Prompt
[Subject] + [Medium/Style] + [Details] + [Quality] + [Lighting] + [Camera]
Ví dụ:
"A samurai warrior standing in a bamboo forest,
digital painting, intricate armor details,
4k, highly detailed, dramatic lighting,
cinematic composition, wide angle shot"
Templates
| Components | Example | Role |
|---|---|---|
| Subject | "a cat", "a castle" | Main subject |
| Action | "running", "sitting" | Action |
| Setting | "in a forest", "on Mars" | Background |
| Style | "oil painting", "anime" | Style |
| Quality | "4k", "highly detailed" | Quality |
| Lighting | "golden hour", "neon" | Light |
| Camera | "close-up", "wide angle" | Shooting angle |
| Mood | "peaceful", "dramatic" | Emotions |
2. Style Keywords
Photography styles:
"portrait photography, 85mm lens, shallow depth of field, bokeh"
"street photography, candid, natural light, grainy film"
"macro photography, extreme close-up, water droplets"
Art styles:
"oil painting, impressionist, thick brushstrokes"
"watercolor, soft edges, pastel colors"
"digital art, concept art, artstation"
"anime style, studio ghibli, cel shading"
"pixel art, 16-bit, retro gaming"
3D/Rendering:
"3D render, octane render, unreal engine 5"
"isometric, low poly, miniature"
"photorealistic, ray tracing, cinema 4D"
3. Negative Prompts
# Negative prompts: what to AVOID in generation
# Universal negative prompt
negative = """
blurry, low quality, low resolution, deformed, distorted,
bad anatomy, bad proportions, extra limbs, extra fingers,
ugly, poorly drawn, watermark, text, signature,
cropped, out of frame, worst quality, jpeg artifacts
"""
# Photo-specific negative
photo_negative = """
cartoon, anime, illustration, painting, drawing,
3d render, CGI, artificial, fake looking
"""
# Art-specific negative
art_negative = """
photograph, photo, realistic, 3d render,
blurry, low effort, amateur
"""
4. Prompt Weighting
# Hugging Face Diffusers: compel library
from compel import Compel
compel = Compel(tokenizer=pipe.tokenizer, text_encoder=pipe.text_encoder)
# Tăng weight: ++ hoặc (word)1.5
prompt = "a (beautiful)1.3 sunset over the (ocean)1.5, dramatic clouds"
conditioning = compel(prompt)
# Giảm weight
prompt = "a cat sitting on a (table)0.5" # less emphasis on table
# Blend prompts
prompt = '"a cat".blend("a dog", 0.7, 0.3)' # 70% cat, 30% dog
A1111 Syntax (WebUI)
# Tăng emphasis: (word:1.3) hoặc ((word))
a ((beautiful)) sunset over the (ocean:1.5)
# Giảm emphasis: [word] hoặc (word:0.7)
a cat sitting on a [table]
# Alternate: [word1|word2] — switch mỗi step
a [cat|dog] sitting in a garden
5. Actual Prompt Patterns
Character Design
"Character concept art of a female cyberpunk hacker,
neon blue hair, glowing circuit tattoos on arms,
wearing a black leather jacket with LED strips,
standing in a rainy Tokyo alley at night,
neon signs reflecting in puddles,
digital art, artstation trending, 8k, highly detailed"
Product Photography
"Professional product photography of a luxury watch,
silver metallic case, black leather strap,
on a dark marble surface, studio lighting,
soft shadows, shallow depth of field,
commercial photography, 4k, Canon EOS R5"
Architecture
"Modern minimalist house, floor to ceiling windows,
concrete and wood materials, infinity pool,
surrounded by tropical garden, golden hour lighting,
architectural photography, wide angle, drone shot"
Fantasy/Game Art
"Epic fantasy landscape, floating islands in the sky,
waterfalls cascading into clouds below,
ancient ruins with glowing runes,
dragon flying in the distance,
volumetric lighting, god rays, matte painting,
concept art, 4k, trending on artstation"
6. Systematic Prompt Testing
# Grid search cho prompt optimization
subjects = ["a cat", "a robot", "a dragon"]
styles = ["oil painting", "digital art", "photograph"]
lightings = ["golden hour", "studio", "neon"]
for subject in subjects:
for style in styles:
for lighting in lightings:
prompt = f"{subject}, {style}, {lighting} lighting, 4k"
image = pipe(prompt, num_inference_steps=25).images[0]
image.save(f"grid_{subject}_{style}_{lighting}.png")
7. Tips & Best Practices
✅ DO:
- Đặt subject ở đầu prompt
- Sử dụng comma để phân tách concepts
- Thêm quality modifiers: "4k, highly detailed, sharp"
- Dùng negative prompts để loại bỏ artifacts
- Test nhiều seeds cho cùng prompt
- Tham khảo community prompts (Civitai, PromptHero)
❌ DON'T:
- Viết câu quá dài phức tạp (CLIP max 77 tokens)
- Dùng từ mơ hồ: "nice", "cool", "good"
- Mâu thuẫn: "realistic anime style"
- Bỏ qua negative prompts
- Kỳ vọng chính xác: AI interpret, không follow literal
Summary
| Engineering | Description |
|---|---|
| Prompt anatomy | Subject + Style + Details + Quality |
| Negative prompts | Remove unwanted elements |
| Prompt weighting | Increase/decrease emphasis for each part |
| Style keywords | Photography, art, 3D rendering styles |
| Systematic testing | Grid search parameters |
📌 Next post: ControlNet & Image-to-Image — control output with reference images.