Chuyển đến nội dung chính

第 17 課:產生人工智慧中的人工智慧安全、道德和版權

內容安全:NSFW偵測,提示過濾。 Deepfake 檢測和預防。人工智慧生成內容的版權問題。 AI 影像浮水印 (C2PA)。負責任的人工智慧部署指南。

🧠 人工智慧與機器學習 — 第 16 課 第 17 課:人工智慧安全、道德與版權 在產生人工智慧領域

生成式 AI:使用 AI 創建圖像和視頻

第六部分:生產與實際應用

亞洲開發網

簡介

將生成式人工智慧部署到生產中需要承擔巨大的責任——從內容安全、深度造假到版權合規性。本文介紹了負責任的人工智慧部署的最佳實踐、工具和指南。


1. 內容安全-NSFW 偵測

from transformers import pipeline

# NSFW classifier
nsfw_detector = pipeline(
    "image-classification",
    model="Falconsai/nsfw_image_detection"
)

def check_safety(image):
    """Check if generated image is safe"""
    results = nsfw_detector(image)
    for result in results:
        if result["label"] == "nsfw" and result["score"] > 0.8:
            return False, "NSFW content detected"
    return True, "Safe"

# Integrate into generation pipeline
def safe_generate(pipe, prompt, **kwargs):
    # 1. Check prompt safety (text filter)
    if not is_prompt_safe(prompt):
        raise ValueError("Prompt contains prohibited content")

    # 2. Generate image
    image = pipe(prompt, **kwargs).images[0]

    # 3. Check output safety
    is_safe, reason = check_safety(image)
    if not is_safe:
        raise ValueError(f"Generated image flagged: {reason}")

    return image

2. 提示安全過濾器

import re

BLOCKED_PATTERNS = [
    r"real\s+person",
    r"celebrity",
    r"child|minor|underage",
    r"violence|gore|blood",
    r"weapon|gun|knife",
]

BLOCKED_NAMES = {
    # Don't generate images of real people without consent
    "public_figures_list"  # load from maintained list
}

def is_prompt_safe(prompt):
    """Check prompt against safety rules"""
    prompt_lower = prompt.lower()

    # Pattern matching
    for pattern in BLOCKED_PATTERNS:
        if re.search(pattern, prompt_lower):
            return False

    # Name matching
    for name in BLOCKED_NAMES:
        if name.lower() in prompt_lower:
            return False

    return True

3. Deepfake 偵測

# Detecting AI-generated images
from transformers import pipeline

deepfake_detector = pipeline(
    "image-classification",
    model="umm-maybe/AI-image-detector"
)

def detect_ai_generated(image):
    """Detect if image is AI-generated"""
    results = deepfake_detector(image)
    for r in results:
        if r["label"] == "artificial" and r["score"] > 0.7:
            return True, r["score"]
    return False, 0.0

# Usage in content moderation
is_ai, confidence = detect_ai_generated(uploaded_image)
if is_ai:
    print(f"AI-generated image detected (confidence: {confidence:.1%})")

4. C2PA 浮水印 — 內容來源

# C2PA: Coalition for Content Provenance and Authenticity
# Embeds metadata proving origin of AI-generated content

# Using c2pa-python library
# pip install c2pa-python

from c2pa import Builder, SigningAlg

def add_provenance(image_path, output_path):
    """Add C2PA provenance metadata to AI-generated image"""
    builder = Builder()

    # Add assertion about AI generation
    builder.add_assertion(
        "c2pa.ai_generated",
        {
            "model": "stable-diffusion-xl",
            "prompt": "a cat in space",
            "generated_at": "2026-03-31T10:00:00Z",
            "platform": "xdev.asia GenAI Platform",
        }
    )

    # Sign and save
    builder.sign_file(
        image_path,
        output_path,
        signing_alg=SigningAlg.ES256,
    )

# Verify provenance
from c2pa import Reader

def verify_provenance(image_path):
    """Check C2PA metadata"""
    reader = Reader(image_path)
    if reader.manifest_store:
        for manifest in reader.manifest_store.manifests:
            print(f"Creator: {manifest.claim_generator}")
            for assertion in manifest.assertions:
                print(f"  Assertion: {assertion.label}")
    else:
        print("No provenance metadata found")

5. 版權注意事項

⚖️ Legal Landscape 2026:

Training data copyright:
- Nhiều vụ kiện đang diễn ra (Getty vs Stability AI, etc.)
- EU AI Act yêu cầu disclosure training data
- Opt-out mechanisms cho artists

Generated content:
- US: AI-generated content không có copyright protection (nếu không có human authorship)
- Significant human creative input → có thể được bảo hộ
- Company-specific policies khác nhau

Best practices:
✅ Sử dụng models trained on licensed data (Firefly, Getty)
✅ Add C2PA metadata cho AI-generated content
✅ Maintain provenance records
✅ Respect opt-out requests
✅ Disclose AI usage
❌ Don't replicate copyrighted characters/brands
❌ Don't claim AI art as human-made
❌ Don't use AI to create counterfeit content

6. 偏見與公平

# Audit generative models for bias
def audit_representation(pipe, category, prompts, num_per_prompt=10):
    """Generate images and audit demographic representation"""
    results = []
    for prompt in prompts:
        for _ in range(num_per_prompt):
            image = pipe(prompt).images[0]
            # Analyze demographic attributes
            analysis = analyze_demographics(image)
            results.append({
                "prompt": prompt,
                "demographics": analysis,
            })

    # Report
    print(f"\n=== Bias Audit: {category} ===")
    # Aggregate and report demographic distribution
    # Flag under-representation or stereotyping
    return results

# Example audit
prompts = [
    "a doctor in a hospital",
    "a CEO in an office",
    "a teacher in a classroom",
    "a engineer at work",
]
audit_representation(pipe, "Occupations", prompts)

7. 負責任的部署清單

Pre-deployment:
□ Content safety filters (input + output)
□ NSFW detection enabled
□ Prompt blocklist maintained
□ Rate limiting configured
□ Usage logging enabled
□ C2PA watermarking integrated

Ongoing:
□ Monitor for misuse patterns
□ Update safety filters regularly
□ Respond to abuse reports
□ Audit for bias quarterly
□ Update legal compliance

User-facing:
□ Clear terms of use
□ Disclosure of AI-generation
□ Report mechanism for problematic content
□ Age verification (if applicable)
□ Transparency about capabilities and limitations

總結

面積工具/實踐
NSFW 偵測隼賽分級機
提示過濾模式匹配、黑名單
Deepfake 偵測AI 影像偵測器模型
來源C2PA 浮水印
版權所有授權型號、揭露
偏見定期審核,多樣化訓練

📌 下一篇文章: Capstone — 建立完整的人工智慧創意平台。