簡介
文字分類是最受歡迎的 NLP 問題——從垃圾郵件檢測、情緒分析到支持票證分類。本文從 A-Z 進行指導:準備資料 → 選擇模型 → 訓練 → 評估 → 部署。
1. 文字分類的類型
| 類型 | 輸入 | 輸出 | 範例 |
|---|---|---|---|
| 二進位 | 文字 | 0/1 | 垃圾郵件與非垃圾郵件 |
| 多類別 | 文字 | N 個標籤中的 1 個 | 主題分類 |
| 多標籤 | 文字 | N 個標籤(同時多個) | 文章標籤 |
| 基於方面的情感 | 文字+外觀 | 每個方面的情緒 | 產品評論 |
2. 使用 PhoBERT 的端對端管道
2.1 準備數據
import pandas as pd
from datasets import Dataset, DatasetDict
from sklearn.model_selection import train_test_split
# Ví dụ: phân loại sentiment tiếng Việt
data = pd.DataFrame({
"text": [
"Sản phẩm rất tốt, giao hàng nhanh",
"Chất lượng tệ, không đáng tiền",
"Bình thường, không có gì đặc biệt",
# ... thêm data
],
"label": [2, 0, 1], # 0=negative, 1=neutral, 2=positive
})
# Split
train_df, test_df = train_test_split(data, test_size=0.2, random_state=42)
dataset = DatasetDict({
"train": Dataset.from_pandas(train_df),
"test": Dataset.from_pandas(test_df),
})
2.2 標記化
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")
def tokenize_fn(examples):
return tokenizer(
examples["text"],
truncation=True,
padding="max_length",
max_length=128,
)
tokenized = dataset.map(tokenize_fn, batched=True)
2.3 微調
from transformers import (
AutoModelForSequenceClassification,
Trainer,
TrainingArguments,
)
import numpy as np
from sklearn.metrics import accuracy_score, f1_score
model = AutoModelForSequenceClassification.from_pretrained(
"vinai/phobert-base-v2",
num_labels=3,
)
def compute_metrics(eval_pred):
logits, labels = eval_pred
predictions = np.argmax(logits, axis=-1)
return {
"accuracy": accuracy_score(labels, predictions),
"f1_macro": f1_score(labels, predictions, average="macro"),
"f1_weighted": f1_score(labels, predictions, average="weighted"),
}
training_args = TrainingArguments(
output_dir="./sentiment-phobert",
num_train_epochs=5,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
learning_rate=2e-5,
weight_decay=0.01,
evaluation_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="f1_macro",
fp16=True,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["test"],
compute_metrics=compute_metrics,
)
trainer.train()
2.4 評估
from sklearn.metrics import classification_report, confusion_matrix
import seaborn as sns
import matplotlib.pyplot as plt
# Predictions
preds = trainer.predict(tokenized["test"])
y_pred = np.argmax(preds.predictions, axis=-1)
y_true = preds.label_ids
# Classification Report
labels = ["Negative", "Neutral", "Positive"]
print(classification_report(y_true, y_pred, target_names=labels))
# Confusion Matrix
cm = confusion_matrix(y_true, y_pred)
sns.heatmap(cm, annot=True, fmt='d', xticklabels=labels, yticklabels=labels)
plt.xlabel("Predicted")
plt.ylabel("True")
plt.title("Confusion Matrix")
plt.show()
3. 使用 FastAPI 部署
from fastapi import FastAPI
from transformers import pipeline
from pydantic import BaseModel
app = FastAPI()
# Load model
classifier = pipeline(
"sentiment-analysis",
model="./sentiment-phobert",
tokenizer="vinai/phobert-base-v2",
)
class TextInput(BaseModel):
text: str
@app.post("/predict")
def predict(input_data: TextInput):
result = classifier(input_data.text)
return {"sentiment": result[0]["label"], "score": result[0]["score"]}
總結
| 步驟 | 工具 |
|---|---|
| 資料準備 | 熊貓,資料集 |
| 代幣化 | 自動分詞器(PhoBERT) |
| 培訓 | 培訓師 API |
| 評價 | sklearn 指標、混淆矩陣 |
| 部署 | FastAPI + 管道 |
下一篇文章
第 14 課:命名實體識別 (NER) — 從文本中提取實體:人員、組織、位置。