Chuyển đến nội dung chính

第 13 課:文本分類與情緒分析

端到端的文字分類管道。情緒分析:二元、多類、基於面向。針對越南語分類微調 BERT/PhoBERT。評價:準確率、F1、混淆矩陣。使用 FastAPI 部署模型。

🧠 人工智慧與機器學習 — 第 12 課 第 13 課:文本分類與情感 分析

NLP 從基礎到進階:掌握自然語言處理

第 5 部分:應用 NLP 問題 — 實作項目

亞洲開發網

簡介

文字分類是最受歡迎的 NLP 問題——從垃圾郵件檢測、情緒分析到支持票證分類。本文從 A-Z 進行指導:準備資料 → 選擇模型 → 訓練 → 評估 → 部署。


1. 文字分類的類型

類型輸入輸出範例
二進位文字0/1垃圾郵件與非垃圾郵件
多類別文字N 個標籤中的 1 個主題分類
多標籤文字N 個標籤(同時多個)文章標籤
基於方面的情感文字+外觀每個方面的情緒產品評論

2. 使用 PhoBERT 的端對端管道

2.1 準備數據

import pandas as pd
from datasets import Dataset, DatasetDict
from sklearn.model_selection import train_test_split

# Ví dụ: phân loại sentiment tiếng Việt
data = pd.DataFrame({
    "text": [
        "Sản phẩm rất tốt, giao hàng nhanh",
        "Chất lượng tệ, không đáng tiền",
        "Bình thường, không có gì đặc biệt",
        # ... thêm data
    ],
    "label": [2, 0, 1],  # 0=negative, 1=neutral, 2=positive
})

# Split
train_df, test_df = train_test_split(data, test_size=0.2, random_state=42)

dataset = DatasetDict({
    "train": Dataset.from_pandas(train_df),
    "test": Dataset.from_pandas(test_df),
})

2.2 標記化

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")

def tokenize_fn(examples):
    return tokenizer(
        examples["text"],
        truncation=True,
        padding="max_length",
        max_length=128,
    )

tokenized = dataset.map(tokenize_fn, batched=True)

2.3 微調

from transformers import (
    AutoModelForSequenceClassification,
    Trainer,
    TrainingArguments,
)
import numpy as np
from sklearn.metrics import accuracy_score, f1_score

model = AutoModelForSequenceClassification.from_pretrained(
    "vinai/phobert-base-v2",
    num_labels=3,
)

def compute_metrics(eval_pred):
    logits, labels = eval_pred
    predictions = np.argmax(logits, axis=-1)
    return {
        "accuracy": accuracy_score(labels, predictions),
        "f1_macro": f1_score(labels, predictions, average="macro"),
        "f1_weighted": f1_score(labels, predictions, average="weighted"),
    }

training_args = TrainingArguments(
    output_dir="./sentiment-phobert",
    num_train_epochs=5,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=32,
    learning_rate=2e-5,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    metric_for_best_model="f1_macro",
    fp16=True,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
    eval_dataset=tokenized["test"],
    compute_metrics=compute_metrics,
)

trainer.train()

2.4 評估

from sklearn.metrics import classification_report, confusion_matrix
import seaborn as sns
import matplotlib.pyplot as plt

# Predictions
preds = trainer.predict(tokenized["test"])
y_pred = np.argmax(preds.predictions, axis=-1)
y_true = preds.label_ids

# Classification Report
labels = ["Negative", "Neutral", "Positive"]
print(classification_report(y_true, y_pred, target_names=labels))

# Confusion Matrix
cm = confusion_matrix(y_true, y_pred)
sns.heatmap(cm, annot=True, fmt='d', xticklabels=labels, yticklabels=labels)
plt.xlabel("Predicted")
plt.ylabel("True")
plt.title("Confusion Matrix")
plt.show()

3. 使用 FastAPI 部署

from fastapi import FastAPI
from transformers import pipeline
from pydantic import BaseModel

app = FastAPI()

# Load model
classifier = pipeline(
    "sentiment-analysis",
    model="./sentiment-phobert",
    tokenizer="vinai/phobert-base-v2",
)

class TextInput(BaseModel):
    text: str

@app.post("/predict")
def predict(input_data: TextInput):
    result = classifier(input_data.text)
    return {"sentiment": result[0]["label"], "score": result[0]["score"]}

總結

步驟工具
資料準備熊貓,資料集
代幣化自動分詞器(PhoBERT)
培訓培訓師 API
評價sklearn 指標、混淆矩陣
部署FastAPI + 管道

下一篇文章

第 14 課:命名實體識別 (NER) — 從文本中提取實體:人員、組織、位置。