Chuyển đến nội dung chính

第 11 課:中途、通量與新興模型

Midjourney API 和 Discord 整合。 Flux:架構和功能。 Google Imagen 3. Adob​​e Firefly API。比較平台之間的品質、速度、成本。多模型編排。

🧠 人工智慧與機器學習 — 第 10 課 第 11 課:中途、通量與新興模型

生成式 AI:使用 AI 創建圖像和視頻

第 4 部分:DALL-E、中途和商業 API

亞洲開發網

簡介

影像生成 2026 市場非常多樣化。除了Stable Diffusion和DALL-E之外,還有Midjourney(最佳美學)、Flux(強開源)、Google Imagen 3、Adobe Firefly(商業安全)。本文對整合進行了比較和指導。


1. 中途

Đặc điểm:
- Aesthetic quality tốt nhất (đặc biệt art, illustration)
- Chạy qua Discord bot hoặc Web UI
- Closed-source, subscription-based
- Hỗ trợ: text-to-image, image-to-image, vary, upscale

Pricing (2026):
- Basic: $10/month (~200 images)
- Standard: $30/month (~900 images)
- Pro: $60/month (unlimited relaxed)

中途參數

/imagine prompt: a dragon flying over mountains --ar 16:9 --v 6 --stylize 750

Parameters:
--ar 16:9      → Aspect ratio
--v 6          → Version
--stylize 750  → Creativity level (0-1000)
--chaos 50     → Variation (0-100)
--quality 2    → Detail level
--no text      → Exclude elements
--tile          → Seamless pattern
--seed 12345   → Reproducibility

2. 通量

# Flux: open-source từ Black Forest Labs (ex-Stability AI team)
# Kiến trúc: DiT (Diffusion Transformer) + T5 text encoder
# Quality ngang DALL-E 3, open-source

from diffusers import FluxPipeline
import torch

pipe = FluxPipeline.from_pretrained(
    "black-forest-labs/FLUX.1-dev",
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")

image = pipe(
    prompt="A cat holding a sign that says 'Hello World'",
    num_inference_steps=30,
    guidance_scale=3.5,
    height=1024,
    width=1024,
).images[0]

通量變體

型號許可證品質速度
Flux.1 專業版僅限 API最佳快速
Flux.1 開發非商業太棒了中等
Flux.1 施內爾阿帕契2.0好非常快(4步)

3. 谷歌圖片 3

# Google Imagen 3 via Vertex AI
from google.cloud import aiplatform
from vertexai.preview.vision_models import ImageGenerationModel

model = ImageGenerationModel.from_pretrained("imagen-3.0-generate-001")

response = model.generate_images(
    prompt="A peaceful Japanese garden with cherry blossoms",
    number_of_images=4,
    aspect_ratio="16:9",
    safety_filter_level="block_some",
    person_generation="dont_allow",
)

response[0].save("imagen_output.png")

4.Adobe Firefly API

# Adobe Firefly: trained on licensed content → commercially safe
import requests

headers = {
    "Authorization": f"Bearer {FIREFLY_TOKEN}",
    "Content-Type": "application/json",
}

response = requests.post(
    "https://firefly-api.adobe.io/v2/images/generate",
    headers=headers,
    json={
        "prompt": "A modern office space with plants",
        "contentClass": "photo",  # photo, art
        "size": {"width": 2048, "height": 2048},
        "n": 1,
        "styles": {"presets": ["photo"]},
    }
)

5. 比較平台

特色標清/標清XL達爾-E 3中途助焊劑圖 3螢火蟲
品質太棒了優秀最佳(藝術)優秀優秀好
圖片中的文字可憐太棒了好最佳好好
提示關注好優秀好優秀優秀好
速度快速(本地)〜10 秒〜30秒中〜10 秒〜15秒
成本/圖片免費(GPU)0.04-0.08 美元~$0.03免費/API~$0.04~$0.04
開源是的沒有沒有部分沒有沒有
商業用途是的是的是(付費)變化是的是的
微調是的沒有沒有是的沒有沒有
自架是的沒有沒有是的沒有沒有

6. 多模型編排

class ImageGeneratorOrchestrator:
    """Route to best model based on use case"""

    def __init__(self):
        self.dalle = OpenAI()
        self.sd_pipe = StableDiffusionXLPipeline.from_pretrained(...)

    async def generate(self, prompt, use_case="general"):
        if use_case == "text_in_image":
            # Flux/DALL-E best for text rendering
            return await self._dalle_generate(prompt)
        elif use_case == "artistic":
            # Local SD with LoRA for custom styles
            return self._sd_generate(prompt)
        elif use_case == "commercial":
            # Firefly for copyright-safe content
            return await self._firefly_generate(prompt)
        elif use_case == "batch":
            # Local SD for cost efficiency
            return self._sd_generate(prompt)
        else:
            return await self._dalle_generate(prompt)

    async def _dalle_generate(self, prompt):
        response = self.dalle.images.generate(
            model="dall-e-3", prompt=prompt, size="1024x1024"
        )
        return response.data[0].url

    def _sd_generate(self, prompt):
        return self.sd_pipe(prompt, num_inference_steps=25).images[0]

7. 為用例選擇模型

📸 Product Photography → DALL-E 3, Imagen 3
🎨 Art & Illustration → Midjourney, SD + LoRA
📝 Text in Image → Flux, DALL-E 3
🏢 Commercial (copyright safe) → Adobe Firefly
🔧 Custom Style → Stable Diffusion + LoRA
💰 Budget / Batch → Stable Diffusion (local)
⚡ Real-time → Flux Schnell, SD Turbo
🔬 Research → Stable Diffusion, Flux Dev

總結

平台最適合定價模式
達爾-E 3一般,文字渲染每個圖像 API
中途藝術、美學訂閱
助焊劑開源,文字免費/API
圖片 3企業、Google雲端每張圖片
螢火蟲版權安全的廣告每張圖片
標清/標清XL客製化、大量、自架免費(GPU 成本)

📌 下一篇文章: 影片產生 — Sora、Runway、Kling 和 Pika。