Chuyển đến nội dung chính

レッスン 11: ミッドジャーニー、フラックス、新興モデル

Midjourney API と Discord の統合。 Flux: アーキテクチャと機能。 Google Imagen 3。Adobe Firefly API。プラットフォーム間の品質、速度、コストを比較します。マルチモデルのオーケストレーション。

🧠 AI と ML — レッスン 10 レッスン 11: ミッドジャーニー、フラックス、新興モデル

生成 AI: AI を使用して画像とビデオを作成する

パート 4: DALL-E、Midjourney、商用 API

xdev.asia

はじめに

2026 年の画像生成市場は非常に多様です。 Stable Diffusion と DALL-E に加えて、Midjourney (最高の美しさ)、Flux (強力なオープンソース)、Google Imagen 3、Adobe Firefly (商用安全) もあります。この記事では、統合を比較し、ガイドします。


1. 旅の途中

Đặc điểm:
- Aesthetic quality tốt nhất (đặc biệt art, illustration)
- Chạy qua Discord bot hoặc Web UI
- Closed-source, subscription-based
- Hỗ trợ: text-to-image, image-to-image, vary, upscale

Pricing (2026):
- Basic: $10/month (~200 images)
- Standard: $30/month (~900 images)
- Pro: $60/month (unlimited relaxed)

旅の途中のパラメータ

/imagine prompt: a dragon flying over mountains --ar 16:9 --v 6 --stylize 750

Parameters:
--ar 16:9      → Aspect ratio
--v 6          → Version
--stylize 750  → Creativity level (0-1000)
--chaos 50     → Variation (0-100)
--quality 2    → Detail level
--no text      → Exclude elements
--tile          → Seamless pattern
--seed 12345   → Reproducibility

2. フラックス

# Flux: open-source từ Black Forest Labs (ex-Stability AI team)
# Kiến trúc: DiT (Diffusion Transformer) + T5 text encoder
# Quality ngang DALL-E 3, open-source

from diffusers import FluxPipeline
import torch

pipe = FluxPipeline.from_pretrained(
    "black-forest-labs/FLUX.1-dev",
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")

image = pipe(
    prompt="A cat holding a sign that says 'Hello World'",
    num_inference_steps=30,
    guidance_scale=3.5,
    height=1024,
    width=1024,
).images[0]

磁束のバリエーション

モデルライセンス品質スピード
Flux.1 プロAPI のみベスト速い
Flux.1 開発非営利すばらしい中
Flux.1 シュネルアパッチ2.0良い非常に速い (4 ステップ)

3. Google Imagen 3

# Google Imagen 3 via Vertex AI
from google.cloud import aiplatform
from vertexai.preview.vision_models import ImageGenerationModel

model = ImageGenerationModel.from_pretrained("imagen-3.0-generate-001")

response = model.generate_images(
    prompt="A peaceful Japanese garden with cherry blossoms",
    number_of_images=4,
    aspect_ratio="16:9",
    safety_filter_level="block_some",
    person_generation="dont_allow",
)

response[0].save("imagen_output.png")

4. Adobe Firefly API

# Adobe Firefly: trained on licensed content → commercially safe
import requests

headers = {
    "Authorization": f"Bearer {FIREFLY_TOKEN}",
    "Content-Type": "application/json",
}

response = requests.post(
    "https://firefly-api.adobe.io/v2/images/generate",
    headers=headers,
    json={
        "prompt": "A modern office space with plants",
        "contentClass": "photo",  # photo, art
        "size": {"width": 2048, "height": 2048},
        "n": 1,
        "styles": {"presets": ["photo"]},
    }
)

5. プラットフォームの比較

特長SD/SDXLダルイー3旅の途中フラックスイマージェン3ホタル
品質すばらしい素晴らしいベスト(アート)素晴らしい素晴らしい良い
画像内のテキスト悪いすばらしい良いベスト良い良い
次のプロンプト良い素晴らしい良い素晴らしい素晴らしい良い
スピード高速 (ローカル)~10代~30代中~10代~15秒
コスト/イメージ無料 (GPU)$0.04-0.08~$0.03無料/API~$0.04~$0.04
オープンソースはいいいえいいえ部分的いいえいいえ
商用利用はいはいはい (有料)さまざまはいはい
微調整可能はいいいえいいえはいいいえいいえ
自己ホスト型はいいいえいいえはいいいえいいえ

6. マルチモデルのオーケストレーション

class ImageGeneratorOrchestrator:
    """Route to best model based on use case"""

    def __init__(self):
        self.dalle = OpenAI()
        self.sd_pipe = StableDiffusionXLPipeline.from_pretrained(...)

    async def generate(self, prompt, use_case="general"):
        if use_case == "text_in_image":
            # Flux/DALL-E best for text rendering
            return await self._dalle_generate(prompt)
        elif use_case == "artistic":
            # Local SD with LoRA for custom styles
            return self._sd_generate(prompt)
        elif use_case == "commercial":
            # Firefly for copyright-safe content
            return await self._firefly_generate(prompt)
        elif use_case == "batch":
            # Local SD for cost efficiency
            return self._sd_generate(prompt)
        else:
            return await self._dalle_generate(prompt)

    async def _dalle_generate(self, prompt):
        response = self.dalle.images.generate(
            model="dall-e-3", prompt=prompt, size="1024x1024"
        )
        return response.data[0].url

    def _sd_generate(self, prompt):
        return self.sd_pipe(prompt, num_inference_steps=25).images[0]

7. ユースケースのモデルを選択する

📸 Product Photography → DALL-E 3, Imagen 3
🎨 Art & Illustration → Midjourney, SD + LoRA
📝 Text in Image → Flux, DALL-E 3
🏢 Commercial (copyright safe) → Adobe Firefly
🔧 Custom Style → Stable Diffusion + LoRA
💰 Budget / Batch → Stable Diffusion (local)
⚡ Real-time → Flux Schnell, SD Turbo
🔬 Research → Stable Diffusion, Flux Dev

概要

プラットフォームこんな方に最適価格モデル
ダルイー3一般、テキスト レンダリング画像ごとの API
旅の途中アート、美学定期購読
フラックスオープンソース、テキスト無料/API
イマージェン3エンタープライズ、Google Cloud画像ごと
ホタル著作権保護された商用画像ごと
SD/SDXLカスタム、バッチ、セルフホスト無料 (GPU コスト)

📌 次の記事: ビデオ生成 — ソラ、ランウェイ、クリング、ピカ。