ビデオ生成の風景 2026。テキストからビデオへ: Sora、Runway Gen-3、Kling、Pika Labs。画像からビデオへ。 AIによる動画編集。時間的な一貫性の課題。 API 統合パターン。
生成 AI: AI を使用して画像とビデオを作成する
パート 5: ビデオ生成とマルチモーダル
xdev.asia
はじめに
2024 年から 2026 年は、Sora (OpenAI)、Runway Gen-3、Kling (Kuaishou)、そして Pika Labs に至る AI ビデオ生成 にとって大きな飛躍です。 「おもちゃのデモ」の品質は、多くのユースケースで「本番環境に対応した」ものになっています。
1. 2026 年のビデオ生成の展望
プラットフォーム
開発者
強み
最大持続時間
ソラ
オープンAI
フォトリアリズム、物理学
60代
滑走路 Gen-3
滑走路
クリエイティブコントロール、編集
10~40代
クリング
クアイショウ
モーション品質、手頃な価格
10代
ピカ
ピカラボ
スピード、使いやすさ
10代
Luma ドリームマシン
ルマAI
3D の一貫性
5~10秒
安定したビデオ
安定性AI
オープンソース
4秒
2. OpenAI Sora API
from openai import OpenAI
client = OpenAI()
# Text-to-Video
response = client.videos.generate(
model="sora",
prompt="""
A golden retriever playing fetch on a sunny beach.
The waves gently crash in the background.
The dog runs towards the camera with a ball in its mouth.
Shot on a cinematic camera, shallow depth of field.
""",
duration=10, # seconds
resolution="1080p",
aspect_ratio="16:9",
)
video_url = response.data[0].url
3. 滑走路 Gen-3 API
import requests
RUNWAY_API_KEY = "your_api_key"
# Text-to-video
response = requests.post(
"https://api.runwayml.com/v1/generate/video",
headers={"Authorization": f"Bearer {RUNWAY_API_KEY}"},
json={
"prompt": "A timelapse of a flower blooming",
"model": "gen3",
"duration": 10,
"resolution": "1280x768",
}
)
task_id = response.json()["task_id"]
# Poll for completion
import time
while True:
status = requests.get(
f"https://api.runwayml.com/v1/tasks/{task_id}",
headers={"Authorization": f"Bearer {RUNWAY_API_KEY}"},
).json()
if status["status"] == "completed":
video_url = status["output"]["video_url"]
break
time.sleep(5)
4. 画像からビデオへの変換
# Animate a still image into a video
import base64
with open("landscape.png", "rb") as f:
image_b64 = base64.b64encode(f.read()).decode()
response = requests.post(
"https://api.runwayml.com/v1/generate/video",
headers={"Authorization": f"Bearer {RUNWAY_API_KEY}"},
json={
"prompt": "camera slowly pans right, clouds moving",
"image": image_b64,
"model": "gen3",
"duration": 5,
}
)
5. 安定したビデオの拡散 (オープンソース)
from diffusers import StableVideoDiffusionPipeline
from diffusers.utils import load_image, export_to_video
import torch
pipe = StableVideoDiffusionPipeline.from_pretrained(
"stabilityai/stable-video-diffusion-img2vid-xt",
torch_dtype=torch.float16,
)
pipe.to("cuda")
# Image-to-video
image = load_image("input_frame.png").resize((1024, 576))
frames = pipe(
image,
num_frames=25, # number of frames
decode_chunk_size=4,
motion_bucket_id=127, # amount of motion (0-255)
fps=7,
num_inference_steps=25,
).frames[0]
export_to_video(frames, "output.mp4", fps=7)
6. 時間的一貫性
Challenge: giữ cho video nhất quán qua các frame
Vấn đề thường gặp:
- Flickering: brightness/color thay đổi giữa frames
- Morphing: objects thay đổi shape
- Disappearing: objects xuất hiện/biến mất
- Physics: vật thể di chuyển không tự nhiên
Giải pháp:
- Temporal attention layers (Sora)
- Frame interpolation (FILM, RIFE)
- Optical flow guidance
- Longer context windows
7. ビデオ編集パイプライン
# Complete video creation workflow
class VideoCreationPipeline:
def __init__(self, sora_client, runway_key):
self.sora = sora_client
self.runway_key = runway_key
async def create_video(self, script):
"""Full pipeline: script → storyboard → video → edit"""
# 1. Generate storyboard (images for each scene)
scenes = self.parse_script(script)
storyboard = []
for scene in scenes:
img = self.sora.images.generate(
model="dall-e-3",
prompt=scene["visual_description"],
)
storyboard.append(img)
# 2. Animate each scene
clips = []
for img, scene in zip(storyboard, scenes):
clip = await self.animate_scene(img, scene["motion"])
clips.append(clip)
# 3. Concatenate clips
final = self.concatenate_clips(clips)
return final
def parse_script(self, script):
"""Parse script into scenes"""
# Use LLM to break down script
response = self.sora.chat.completions.create(
model="gpt-4",
messages=[{
"role": "user",
"content": f"Break this into video scenes: {script}"
}]
)
return eval(response.choices[0].message.content)