はじめに
テキスト分類は、スパム検出、センチメント分析、サポート チケット分類に至るまで、最も一般的な NLP 問題です。この記事では、データの準備→モデルの選択→トレーニング→評価→デプロイの順に説明します。
1. テキスト分類の種類
| タイプ | 入力 | 出力 | 例 |
|---|---|---|---|
| バイナリ | テキスト | 0/1 | スパム vs スパムではない |
| マルチクラス | テキスト | N 個のラベルのうち 1 個 | トピック分類 |
| マルチレーベル | テキスト | N 個のラベル (同時に多数) | 記事のタグ |
| アスペクトベースの感情 | テキスト + アスペクト | 側面ごとのセンチメント | 製品レビュー |
2. PhoBERT を使用したエンドツーエンドのパイプライン
2.1 データの準備
import pandas as pd
from datasets import Dataset, DatasetDict
from sklearn.model_selection import train_test_split
# Ví dụ: phân loại sentiment tiếng Việt
data = pd.DataFrame({
"text": [
"Sản phẩm rất tốt, giao hàng nhanh",
"Chất lượng tệ, không đáng tiền",
"Bình thường, không có gì đặc biệt",
# ... thêm data
],
"label": [2, 0, 1], # 0=negative, 1=neutral, 2=positive
})
# Split
train_df, test_df = train_test_split(data, test_size=0.2, random_state=42)
dataset = DatasetDict({
"train": Dataset.from_pandas(train_df),
"test": Dataset.from_pandas(test_df),
})
2.2 トークン化
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")
def tokenize_fn(examples):
return tokenizer(
examples["text"],
truncation=True,
padding="max_length",
max_length=128,
)
tokenized = dataset.map(tokenize_fn, batched=True)
2.3 微調整
from transformers import (
AutoModelForSequenceClassification,
Trainer,
TrainingArguments,
)
import numpy as np
from sklearn.metrics import accuracy_score, f1_score
model = AutoModelForSequenceClassification.from_pretrained(
"vinai/phobert-base-v2",
num_labels=3,
)
def compute_metrics(eval_pred):
logits, labels = eval_pred
predictions = np.argmax(logits, axis=-1)
return {
"accuracy": accuracy_score(labels, predictions),
"f1_macro": f1_score(labels, predictions, average="macro"),
"f1_weighted": f1_score(labels, predictions, average="weighted"),
}
training_args = TrainingArguments(
output_dir="./sentiment-phobert",
num_train_epochs=5,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
learning_rate=2e-5,
weight_decay=0.01,
evaluation_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="f1_macro",
fp16=True,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["test"],
compute_metrics=compute_metrics,
)
trainer.train()
2.4 評価
from sklearn.metrics import classification_report, confusion_matrix
import seaborn as sns
import matplotlib.pyplot as plt
# Predictions
preds = trainer.predict(tokenized["test"])
y_pred = np.argmax(preds.predictions, axis=-1)
y_true = preds.label_ids
# Classification Report
labels = ["Negative", "Neutral", "Positive"]
print(classification_report(y_true, y_pred, target_names=labels))
# Confusion Matrix
cm = confusion_matrix(y_true, y_pred)
sns.heatmap(cm, annot=True, fmt='d', xticklabels=labels, yticklabels=labels)
plt.xlabel("Predicted")
plt.ylabel("True")
plt.title("Confusion Matrix")
plt.show()
3. FastAPI を使用してデプロイする
from fastapi import FastAPI
from transformers import pipeline
from pydantic import BaseModel
app = FastAPI()
# Load model
classifier = pipeline(
"sentiment-analysis",
model="./sentiment-phobert",
tokenizer="vinai/phobert-base-v2",
)
class TextInput(BaseModel):
text: str
@app.post("/predict")
def predict(input_data: TextInput):
result = classifier(input_data.text)
return {"sentiment": result[0]["label"], "score": result[0]["score"]}
概要
| ステップ | ツール |
|---|---|
| データ準備 | パンダ、データセット |
| トークン化 | AutoTokenizer (PhoBERT) |
| トレーニング | トレーナー API |
| 評価 | sklearn メトリクス、混同行列 |
| 導入 | FastAPI + パイプライン |
次の記事
レッスン 14: 固有表現認識 (NER) — テキストからエンティティ (人、組織、場所) を抽出します。