Chuyển đến nội dung chính

第 12 課:影片產生 — Sora、Runway、Kling 和 Pika

2026 年影片生成前景。文字轉影片:Sora、Runway Gen-3、Kling、Pika Labs。圖像到影片。使用 AI 進行影片編輯。時間一致性挑戰。 API 整合模式。

🧠 人工智慧與機器學習 — 第 11 課 第 12 課:影片產生 — Sora、Runway、 克林與皮卡

生成式 AI:使用 AI 創建圖像和視頻

第 5 部分:視訊生成和多模式

亞洲開發網

簡介

2024-2026 年是 AI 視訊世代 的巨大飛躍——從 Sora (OpenAI)、Runway Gen-3、Kling (快手) 到 Pika Labs。對於許多用例來說,「玩具演示」的品質已成為「生產就緒」。


1. 2026 年影片生成格局

平台開發商優勢最長持續時間
索拉開放人工智慧照相寫實主義,物理學60 年代
跑道 Gen-3跑道創意控制、編輯10-40 秒
克林快手運動品質,價格實惠10 秒
皮卡皮卡實驗室速度快,易用性10 秒
盧瑪夢想機魯瑪人工智慧3D 一致性5-10 秒
視訊穩定穩定性人工智慧開源4s

2.OpenAI Sora API

from openai import OpenAI

client = OpenAI()

# Text-to-Video
response = client.videos.generate(
    model="sora",
    prompt="""
    A golden retriever playing fetch on a sunny beach.
    The waves gently crash in the background.
    The dog runs towards the camera with a ball in its mouth.
    Shot on a cinematic camera, shallow depth of field.
    """,
    duration=10,           # seconds
    resolution="1080p",
    aspect_ratio="16:9",
)

video_url = response.data[0].url

3. Runway Gen-3 API

import requests

RUNWAY_API_KEY = "your_api_key"

# Text-to-video
response = requests.post(
    "https://api.runwayml.com/v1/generate/video",
    headers={"Authorization": f"Bearer {RUNWAY_API_KEY}"},
    json={
        "prompt": "A timelapse of a flower blooming",
        "model": "gen3",
        "duration": 10,
        "resolution": "1280x768",
    }
)

task_id = response.json()["task_id"]

# Poll for completion
import time
while True:
    status = requests.get(
        f"https://api.runwayml.com/v1/tasks/{task_id}",
        headers={"Authorization": f"Bearer {RUNWAY_API_KEY}"},
    ).json()
    if status["status"] == "completed":
        video_url = status["output"]["video_url"]
        break
    time.sleep(5)

4. 影像到視頻

# Animate a still image into a video
import base64

with open("landscape.png", "rb") as f:
    image_b64 = base64.b64encode(f.read()).decode()

response = requests.post(
    "https://api.runwayml.com/v1/generate/video",
    headers={"Authorization": f"Bearer {RUNWAY_API_KEY}"},
    json={
        "prompt": "camera slowly pans right, clouds moving",
        "image": image_b64,
        "model": "gen3",
        "duration": 5,
    }
)

5.穩定的視訊擴散(開源)

from diffusers import StableVideoDiffusionPipeline
from diffusers.utils import load_image, export_to_video
import torch

pipe = StableVideoDiffusionPipeline.from_pretrained(
    "stabilityai/stable-video-diffusion-img2vid-xt",
    torch_dtype=torch.float16,
)
pipe.to("cuda")

# Image-to-video
image = load_image("input_frame.png").resize((1024, 576))

frames = pipe(
    image,
    num_frames=25,          # number of frames
    decode_chunk_size=4,
    motion_bucket_id=127,    # amount of motion (0-255)
    fps=7,
    num_inference_steps=25,
).frames[0]

export_to_video(frames, "output.mp4", fps=7)

6. 時間一致性

Challenge: giữ cho video nhất quán qua các frame

Vấn đề thường gặp:
- Flickering: brightness/color thay đổi giữa frames
- Morphing: objects thay đổi shape
- Disappearing: objects xuất hiện/biến mất
- Physics: vật thể di chuyển không tự nhiên

Giải pháp:
- Temporal attention layers (Sora)
- Frame interpolation (FILM, RIFE)
- Optical flow guidance
- Longer context windows

7. 影片編輯管道

# Complete video creation workflow

class VideoCreationPipeline:
    def __init__(self, sora_client, runway_key):
        self.sora = sora_client
        self.runway_key = runway_key

    async def create_video(self, script):
        """Full pipeline: script → storyboard → video → edit"""
        # 1. Generate storyboard (images for each scene)
        scenes = self.parse_script(script)
        storyboard = []
        for scene in scenes:
            img = self.sora.images.generate(
                model="dall-e-3",
                prompt=scene["visual_description"],
            )
            storyboard.append(img)

        # 2. Animate each scene
        clips = []
        for img, scene in zip(storyboard, scenes):
            clip = await self.animate_scene(img, scene["motion"])
            clips.append(clip)

        # 3. Concatenate clips
        final = self.concatenate_clips(clips)
        return final

    def parse_script(self, script):
        """Parse script into scenes"""
        # Use LLM to break down script
        response = self.sora.chat.completions.create(
            model="gpt-4",
            messages=[{
                "role": "user",
                "content": f"Break this into video scenes: {script}"
            }]
        )
        return eval(response.choices[0].message.content)

8. 成本比較

平台每 10 秒的成本品質速度
索拉~$0.50-2.00最佳約 2 分鐘
跑道 Gen-3~$0.25-0.50太棒了約 1 分鐘
克林~$0.10-0.20好〜30秒
皮卡~$0.05-0.15好~20 秒
SVD(本地)僅 GPU 成本體面約 5 分鐘

總結

概念描述
文字轉影片從文字描述產生影片
圖像到影片靜態圖像動畫
時間一致性跨框架一致
影片管道腳本→故事板→動畫→編輯

📌 下一篇文章: 音訊和音樂產生 — 使用 AI 創作聲音和音樂。