Chuyển đến nội dung chính

レッスン 12: ビデオ生成 — ソラ、ランウェイ、クリング、ピカ

ビデオ生成の風景 2026。テキストからビデオへ: Sora、Runway Gen-3、Kling、Pika Labs。画像からビデオへ。 AIによる動画編集。時間的な一貫性の課題。 API 統合パターン。

🧠 AI と ML — レッスン 11 レッスン 12: ビデオ生成 — ソラ、ランウェイ、 クリング&ピカ

生成 AI: AI を使用して画像とビデオを作成する

パート 5: ビデオ生成とマルチモーダル

xdev.asia

はじめに

2024 年から 2026 年は、Sora (OpenAI)、Runway Gen-3、Kling (Kuaishou)、そして Pika Labs に至る AI ビデオ生成 にとって大きな飛躍です。 「おもちゃのデモ」の品質は、多くのユースケースで「本番環境に対応した」ものになっています。


1. 2026 年のビデオ生成の展望

プラットフォーム開発者強み最大持続時間
ソラオープンAIフォトリアリズム、物理学60代
滑走路 Gen-3滑走路クリエイティブコントロール、編集10~40代
クリングクアイショウモーション品質、手頃な価格10代
ピカピカラボスピード、使いやすさ10代
Luma ドリームマシンルマAI3D の一貫性5~10秒
安定したビデオ安定性AIオープンソース4秒

2. OpenAI Sora API

from openai import OpenAI

client = OpenAI()

# Text-to-Video
response = client.videos.generate(
    model="sora",
    prompt="""
    A golden retriever playing fetch on a sunny beach.
    The waves gently crash in the background.
    The dog runs towards the camera with a ball in its mouth.
    Shot on a cinematic camera, shallow depth of field.
    """,
    duration=10,           # seconds
    resolution="1080p",
    aspect_ratio="16:9",
)

video_url = response.data[0].url

3. 滑走路 Gen-3 API

import requests

RUNWAY_API_KEY = "your_api_key"

# Text-to-video
response = requests.post(
    "https://api.runwayml.com/v1/generate/video",
    headers={"Authorization": f"Bearer {RUNWAY_API_KEY}"},
    json={
        "prompt": "A timelapse of a flower blooming",
        "model": "gen3",
        "duration": 10,
        "resolution": "1280x768",
    }
)

task_id = response.json()["task_id"]

# Poll for completion
import time
while True:
    status = requests.get(
        f"https://api.runwayml.com/v1/tasks/{task_id}",
        headers={"Authorization": f"Bearer {RUNWAY_API_KEY}"},
    ).json()
    if status["status"] == "completed":
        video_url = status["output"]["video_url"]
        break
    time.sleep(5)

4. 画像からビデオへの変換

# Animate a still image into a video
import base64

with open("landscape.png", "rb") as f:
    image_b64 = base64.b64encode(f.read()).decode()

response = requests.post(
    "https://api.runwayml.com/v1/generate/video",
    headers={"Authorization": f"Bearer {RUNWAY_API_KEY}"},
    json={
        "prompt": "camera slowly pans right, clouds moving",
        "image": image_b64,
        "model": "gen3",
        "duration": 5,
    }
)

5. 安定したビデオの拡散 (オープンソース)

from diffusers import StableVideoDiffusionPipeline
from diffusers.utils import load_image, export_to_video
import torch

pipe = StableVideoDiffusionPipeline.from_pretrained(
    "stabilityai/stable-video-diffusion-img2vid-xt",
    torch_dtype=torch.float16,
)
pipe.to("cuda")

# Image-to-video
image = load_image("input_frame.png").resize((1024, 576))

frames = pipe(
    image,
    num_frames=25,          # number of frames
    decode_chunk_size=4,
    motion_bucket_id=127,    # amount of motion (0-255)
    fps=7,
    num_inference_steps=25,
).frames[0]

export_to_video(frames, "output.mp4", fps=7)

6. 時間的一貫性

Challenge: giữ cho video nhất quán qua các frame

Vấn đề thường gặp:
- Flickering: brightness/color thay đổi giữa frames
- Morphing: objects thay đổi shape
- Disappearing: objects xuất hiện/biến mất
- Physics: vật thể di chuyển không tự nhiên

Giải pháp:
- Temporal attention layers (Sora)
- Frame interpolation (FILM, RIFE)
- Optical flow guidance
- Longer context windows

7. ビデオ編集パイプライン

# Complete video creation workflow

class VideoCreationPipeline:
    def __init__(self, sora_client, runway_key):
        self.sora = sora_client
        self.runway_key = runway_key

    async def create_video(self, script):
        """Full pipeline: script → storyboard → video → edit"""
        # 1. Generate storyboard (images for each scene)
        scenes = self.parse_script(script)
        storyboard = []
        for scene in scenes:
            img = self.sora.images.generate(
                model="dall-e-3",
                prompt=scene["visual_description"],
            )
            storyboard.append(img)

        # 2. Animate each scene
        clips = []
        for img, scene in zip(storyboard, scenes):
            clip = await self.animate_scene(img, scene["motion"])
            clips.append(clip)

        # 3. Concatenate clips
        final = self.concatenate_clips(clips)
        return final

    def parse_script(self, script):
        """Parse script into scenes"""
        # Use LLM to break down script
        response = self.sora.chat.completions.create(
            model="gpt-4",
            messages=[{
                "role": "user",
                "content": f"Break this into video scenes: {script}"
            }]
        )
        return eval(response.choices[0].message.content)

8. コストの比較

プラットフォーム10 秒あたりのコスト品質スピード
ソラ~$0.50-2.00ベスト~2分
滑走路 Gen-3~$0.25-0.50すばらしい~1分
クリング~$0.10-0.20良い~30代
ピカ~$0.05-0.15良い~20代
SVD (ローカル)GPU コストのみまともな~5分

概要

コンセプト説明
テキストからビデオへテキストの説明からビデオを生成
画像からビデオへ静止画をアニメーション化する
時間的一貫性フレーム間で一貫性を保つ
ビデオパイプライン台本 → 絵コンテ → アニメーション → 編集

📌 次の記事: オーディオと音楽の生成 — AI を使用してサウンドと音楽を作成します。