はじめに
AIはフリーテキストで回答→コードで解析するのが難しい。実稼働環境では、下流システムが処理するための 構造化 出力 (JSON、テーブル、固定リストなど) が必要です。
例: 「レビューの感情を分析する」 → AI は「これは肯定的なレビューです」と回答します (フリーテキスト、解析は困難)。構造化された出力:
{"sentiment": "positive", "score": 0.85, "keywords": ["tốt", "nhanh"]}→ コード解析が簡単!
この記事の内容は次のとおりです。
- JSON モード — AI に JSON を強制的に返す
- JSON スキーマ — 正確な構造を定義する
- 関数呼び出し — ツールの使用による構造化された出力
- Pydantic + Instructor — タイプセーフな検証
1. JSON モード — AI が JSON を返すようにする
1.1 基本的なプロンプト
❌ Prompt kém:
"Phân tích review này và cho biết sentiment"
→ AI: "Review này có sentiment tích cực vì người dùng khen sản phẩm tốt..."
✅ Prompt tốt:
"Phân tích review này. Trả lời CHÍNH XÁC bằng JSON:
{
"sentiment": "positive" | "negative" | "neutral",
"score": 0.0-1.0,
"keywords": ["từ khóa 1", "từ khóa 2"],
"summary": "tóm tắt 1 câu"
}"
→ AI: {"sentiment": "positive", "score": 0.85, "keywords": ["tốt", "nhanh"], "summary": "..."}
1.2 OpenAI JSON モード
"""OpenAI JSON Mode — đảm bảo output là valid JSON"""
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o-mini",
response_format={"type": "json_object"}, # ← JSON Mode ON
messages=[
{"role": "system", "content": "Trả lời bằng JSON."},
{"role": "user", "content": """Phân tích sentiment:
Review: "Sản phẩm tốt, giao hàng nhanh, sẽ mua lại!"
JSON format:
{"sentiment": "positive|negative|neutral", "score": 0-1, "keywords": [...]}"""},
],
)
import json
result = json.loads(response.choices[0].message.content)
print(result)
# {"sentiment": "positive", "score": 0.92, "keywords": ["tốt", "nhanh", "mua lại"]}
1.3 JSON モードの問題
JSON Mode chỉ đảm bảo output là VALID JSON,
KHÔNG đảm bảo schema đúng!
Bạn yêu cầu: {"sentiment": "...", "score": ...}
AI có thể trả: {"feeling": "good", "rating": 5} ← Sai field names!
→ Cần JSON Schema để enforce cấu trúc chính xác.
2. JSON スキーマ — 構造化出力
2.1 OpenAI 構造化出力 (2024+)
"""Structured Outputs: define schema, AI PHẢI tuân thủ"""
from openai import OpenAI
from pydantic import BaseModel
class SentimentAnalysis(BaseModel):
sentiment: str # "positive", "negative", "neutral"
score: float # 0.0 - 1.0
keywords: list[str]
summary: str
client = OpenAI()
response = client.beta.chat.completions.parse(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "Phân tích sentiment review."},
{"role": "user", "content": "Sản phẩm tốt, giao hàng nhanh!"},
],
response_format=SentimentAnalysis, # ← Schema enforcement
)
result = response.choices[0].message.parsed
print(result.sentiment) # "positive"
print(result.score) # 0.92
print(result.keywords) # ["tốt", "nhanh"]
2.2 複雑なスキーマ
"""Schema phức tạp: nested objects, enums, optional fields"""
from pydantic import BaseModel, Field
from typing import Optional
from enum import Enum
class Sentiment(str, Enum):
positive = "positive"
negative = "negative"
neutral = "neutral"
class Aspect(BaseModel):
category: str = Field(description="Khía cạnh: quality, price, delivery, service")
sentiment: Sentiment
text: str = Field(description="Đoạn text liên quan")
class ReviewAnalysis(BaseModel):
overall_sentiment: Sentiment
overall_score: float = Field(ge=0, le=1, description="0=rất tiêu cực, 1=rất tích cực")
aspects: list[Aspect] = Field(description="Phân tích từng khía cạnh")
recommendation: bool = Field(description="Có nên mua không?")
summary: str
# Output:
# {
# "overall_sentiment": "positive",
# "overall_score": 0.85,
# "aspects": [
# {"category": "quality", "sentiment": "positive", "text": "Sản phẩm tốt"},
# {"category": "delivery", "sentiment": "positive", "text": "giao hàng nhanh"},
# {"category": "price", "sentiment": "neutral", "text": "giá hợp lý"}
# ],
# "recommendation": true,
# "summary": "..."
# }
💡 演習 1: CV/履歴書から情報を抽出するためのスキーマを作成します: 名前、電子メール、スキル、経験 (リスト)、教育レベル。 3 つの異なる CV を使用してテストします。
3. 関数呼び出し
3.1 構造化された出力に使用するツール
"""Function calling: AI "gọi function" với arguments có cấu trúc"""
from openai import OpenAI
client = OpenAI()
tools = [{
"type": "function",
"function": {
"name": "save_contact",
"description": "Lưu thông tin liên hệ",
"parameters": {
"type": "object",
"properties": {
"name": {"type": "string", "description": "Họ tên"},
"phone": {"type": "string", "description": "Số điện thoại"},
"email": {"type": "string", "description": "Email"},
"company": {"type": "string", "description": "Công ty"},
},
"required": ["name"],
},
},
}]
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content":
"Anh Minh, SĐT 0901234567, email [email protected], công ty XDev"}],
tools=tools,
tool_choice={"type": "function", "function": {"name": "save_contact"}},
)
# AI trả về structured arguments
args = json.loads(response.choices[0].message.tool_calls[0].function.arguments)
# {"name": "Minh", "phone": "0901234567", "email": "[email protected]", "company": "XDev"}
3.2 関数呼び出しと JSON スキーマをいつ使用するか?
| 特長 | JSON スキーマ | 関数呼び出し |
|---|---|---|
| 出力形式 | JSON オブジェクト | 関数の引数 |
| スキーマの適用 | ✅ 厳格 | ✅ 厳格 |
| 複数の出力 | ❌ 1 オブジェクト | ✅ 複数のツール呼び出し |
| ストリーミング | ✅ | ✅ |
| 使用例 | データ抽出 | アクション + 抽出 |
4. LangChain 構造化出力
4.1 with_structed_output()
"""LangChain: structured output dễ dàng"""
from langchain_openai import ChatOpenAI
from pydantic import BaseModel, Field
class ExtractedInfo(BaseModel):
"""Thông tin trích xuất từ email"""
sender: str = Field(description="Người gửi")
subject: str = Field(description="Chủ đề")
action_items: list[str] = Field(description="Việc cần làm")
priority: str = Field(description="high, medium, low")
deadline: str | None = Field(description="Deadline nếu có")
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
structured_llm = llm.with_structured_output(ExtractedInfo)
email = """Chào team,
Cần hoàn thành report Q3 trước thứ 6 tuần này.
Minh review data, Hùng viết slides.
Ưu tiên cao vì CEO cần trình bày thứ 2.
Thanks."""
result = structured_llm.invoke(f"Extract thông tin từ email:\n{email}")
print(result.action_items) # ["Hoàn thành report Q3", "Review data", "Viết slides"]
print(result.priority) # "high"
print(result.deadline) # "Thứ 6 tuần này"
4.2 インストラクター — タイプセーフ + 再試行
"""Instructor: Pydantic validation + auto retry"""
# pip install instructor
import instructor
from openai import OpenAI
from pydantic import BaseModel, field_validator
client = instructor.from_openai(OpenAI())
class UserInfo(BaseModel):
name: str
age: int
email: str
@field_validator("age")
@classmethod
def validate_age(cls, v):
if not 0 < v < 150:
raise ValueError("Age must be between 1 and 149")
return v
@field_validator("email")
@classmethod
def validate_email(cls, v):
if "@" not in v:
raise ValueError("Invalid email format")
return v
# Instructor tự retry nếu validation fail!
result = client.chat.completions.create(
model="gpt-4o-mini",
response_model=UserInfo,
max_retries=3, # Retry tối đa 3 lần nếu validation fail
messages=[{"role": "user", "content": "Minh, 30 tuổi, [email protected]"}],
)
print(result) # UserInfo(name="Minh", age=30, email="[email protected]")
💡 演習 2: Instructor を使用して抽出パイプラインを作成します: 入力 = 製品説明段落 → 出力 = スキーマ (名前、価格、カテゴリ、機能、評価)。バリデータを追加: 価格 > 0、評価 1 ~ 5。
5. エラー処理とエッジケース
5.1 再試行パターン
"""Retry khi JSON parse fail"""
import json
from tenacity import retry, stop_after_attempt, retry_if_exception_type
@retry(
stop=stop_after_attempt(3),
retry=retry_if_exception_type(json.JSONDecodeError),
)
def extract_json(text: str, prompt: str) -> dict:
response = client.chat.completions.create(
model="gpt-4o-mini",
response_format={"type": "json_object"},
messages=[
{"role": "system", "content": prompt},
{"role": "user", "content": text},
],
)
return json.loads(response.choices[0].message.content)
5.2 フォールバック戦略
Strategy khi structured output fail:
1. JSON Schema (strict) ← Thử đầu tiên
↓ fail
2. JSON Mode + prompt ← Fallback
↓ fail
3. Text output + regex parse ← Last resort
概要
| コンセプト | 覚えておいてください |
|---|---|
| JSON モード | スキーマを強制するのではなく、有効な JSON を確認する |
| 構造化された出力 | スキーマの強制、Pydantic モデル |
| 関数呼び出し | ツールの使用形式、複数の呼び出し |
| with_structed_output() | LangChain ラッパー、使いやすい |
| インストラクター | Pydantic 検証 + 自動再試行 |
| 再試行 | 解析が失敗した場合のテナシティーの再試行 |
一般的な演習
- ✅ 2 つの小さな演習 (1、2) を完了します。
- 電子メール分類子: 入力 = 電子メール → 出力 =
{category, priority, action_items, sentiment, response_draft}。 Pydantic スキーマ + 検証を使用します。 10 通のメールでテストします。 - データ パイプライン: 20 件のレビューをクロール → 構造化データを抽出 (講師) → CSV に保存 → 分析。構造化出力と正規表現解析の精度を比較します。
- マルチモデル: 構造化出力品質の比較: GPT-4o-mini 対 Claude 対 Gemini。どのモデルがスキーマに最もよく準拠していますか?
次の記事: コード生成のためのプロンプト エンジニアリング — AI が高品質のコードを生成し、コードをレビューし、デバッグし、テストを生成するためのプロンプトを作成します。