Chuyển đến nội dung chính

レッスン 13: テキストの分類と感情分析

エンドツーエンドのテキスト分類パイプライン。センチメント分析: バイナリ、マルチクラス、アスペクトベース。ベトナム語分類用に BERT/PhoBERT を微調整します。評価: 精度、F1、混同行列。 FastAPI を使用してモデルをデプロイします。

🧠 AI と ML — レッスン 12 レッスン 13: テキストの分類と感情 分析

NLP の基礎から上級まで: 自然言語処理をマスターする

パート 5: NLP の応用問題 — 実践プロジェクト

xdev.asia

はじめに

テキスト分類は、スパム検出、センチメント分析、サポート チケット分類に至るまで、最も一般的な NLP 問題です。この記事では、データの準備→モデルの選択→トレーニング→評価→デプロイの順に説明します。


1. テキスト分類の種類

タイプ入力出力例
バイナリテキスト0/1スパム vs スパムではない
マルチクラステキストN 個のラベルのうち 1 個トピック分類
マルチレーベルテキストN 個のラベル (同時に多数)記事のタグ
アスペクトベースの感情テキスト + アスペクト側面ごとのセンチメント製品レビュー

2. PhoBERT を使用したエンドツーエンドのパイプライン

2.1 データの準備

import pandas as pd
from datasets import Dataset, DatasetDict
from sklearn.model_selection import train_test_split

# Ví dụ: phân loại sentiment tiếng Việt
data = pd.DataFrame({
    "text": [
        "Sản phẩm rất tốt, giao hàng nhanh",
        "Chất lượng tệ, không đáng tiền",
        "Bình thường, không có gì đặc biệt",
        # ... thêm data
    ],
    "label": [2, 0, 1],  # 0=negative, 1=neutral, 2=positive
})

# Split
train_df, test_df = train_test_split(data, test_size=0.2, random_state=42)

dataset = DatasetDict({
    "train": Dataset.from_pandas(train_df),
    "test": Dataset.from_pandas(test_df),
})

2.2 トークン化

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")

def tokenize_fn(examples):
    return tokenizer(
        examples["text"],
        truncation=True,
        padding="max_length",
        max_length=128,
    )

tokenized = dataset.map(tokenize_fn, batched=True)

2.3 微調整

from transformers import (
    AutoModelForSequenceClassification,
    Trainer,
    TrainingArguments,
)
import numpy as np
from sklearn.metrics import accuracy_score, f1_score

model = AutoModelForSequenceClassification.from_pretrained(
    "vinai/phobert-base-v2",
    num_labels=3,
)

def compute_metrics(eval_pred):
    logits, labels = eval_pred
    predictions = np.argmax(logits, axis=-1)
    return {
        "accuracy": accuracy_score(labels, predictions),
        "f1_macro": f1_score(labels, predictions, average="macro"),
        "f1_weighted": f1_score(labels, predictions, average="weighted"),
    }

training_args = TrainingArguments(
    output_dir="./sentiment-phobert",
    num_train_epochs=5,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=32,
    learning_rate=2e-5,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    metric_for_best_model="f1_macro",
    fp16=True,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
    eval_dataset=tokenized["test"],
    compute_metrics=compute_metrics,
)

trainer.train()

2.4 評価

from sklearn.metrics import classification_report, confusion_matrix
import seaborn as sns
import matplotlib.pyplot as plt

# Predictions
preds = trainer.predict(tokenized["test"])
y_pred = np.argmax(preds.predictions, axis=-1)
y_true = preds.label_ids

# Classification Report
labels = ["Negative", "Neutral", "Positive"]
print(classification_report(y_true, y_pred, target_names=labels))

# Confusion Matrix
cm = confusion_matrix(y_true, y_pred)
sns.heatmap(cm, annot=True, fmt='d', xticklabels=labels, yticklabels=labels)
plt.xlabel("Predicted")
plt.ylabel("True")
plt.title("Confusion Matrix")
plt.show()

3. FastAPI を使用してデプロイする

from fastapi import FastAPI
from transformers import pipeline
from pydantic import BaseModel

app = FastAPI()

# Load model
classifier = pipeline(
    "sentiment-analysis",
    model="./sentiment-phobert",
    tokenizer="vinai/phobert-base-v2",
)

class TextInput(BaseModel):
    text: str

@app.post("/predict")
def predict(input_data: TextInput):
    result = classifier(input_data.text)
    return {"sentiment": result[0]["label"], "score": result[0]["score"]}

概要

ステップツール
データ準備パンダ、データセット
トークン化AutoTokenizer (PhoBERT)
トレーニングトレーナー API
評価sklearn メトリクス、混同行列
導入FastAPI + パイプライン

次の記事

レッスン 14: 固有表現認識 (NER) — テキストからエンティティ (人、組織、場所) を抽出します。