Chuyển đến nội dung chính

Lesson 13: Text Classification & Sentiment Analysis

Text classification pipeline end-to-end. Sentiment analysis: binary, multi-class, aspect-based. Fine-tune BERT/PhoBERT for Vietnamese classification. Evaluation: accuracy, F1, confusion matrix. Deploy model with FastAPI.

🧠 AI & ML — Lesson 12 Lesson 13: Text Classification & Sentiment Analysis

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 5: Applied NLP problems — Hands-on Projects

xdev.asia

Introduction

Text Classification is the most popular NLP problem — from spam detection, sentiment analysis, to support ticket classification. This article guides from A-Z: prepare data → choose model → train → evaluate → deploy.


1. Types of Text Classification

TypeInputOutputExample
BinaryText0/1Spam vs Not spam
Multi-classText1 of N labelsTopic classification
Multi-labelTextN labels (many simultaneously)Tags for articles
Aspect-based SentimentText + AspectSentiment per aspectProduct reviews

2. End-to-End Pipeline with PhoBERT

2.1 Prepare Data

import pandas as pd
from datasets import Dataset, DatasetDict
from sklearn.model_selection import train_test_split

# Ví dụ: phân loại sentiment tiếng Việt
data = pd.DataFrame({
    "text": [
        "Sản phẩm rất tốt, giao hàng nhanh",
        "Chất lượng tệ, không đáng tiền",
        "Bình thường, không có gì đặc biệt",
        # ... thêm data
    ],
    "label": [2, 0, 1],  # 0=negative, 1=neutral, 2=positive
})

# Split
train_df, test_df = train_test_split(data, test_size=0.2, random_state=42)

dataset = DatasetDict({
    "train": Dataset.from_pandas(train_df),
    "test": Dataset.from_pandas(test_df),
})

2.2 Tokenize

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")

def tokenize_fn(examples):
    return tokenizer(
        examples["text"],
        truncation=True,
        padding="max_length",
        max_length=128,
    )

tokenized = dataset.map(tokenize_fn, batched=True)

2.3 Fine-tune

from transformers import (
    AutoModelForSequenceClassification,
    Trainer,
    TrainingArguments,
)
import numpy as np
from sklearn.metrics import accuracy_score, f1_score

model = AutoModelForSequenceClassification.from_pretrained(
    "vinai/phobert-base-v2",
    num_labels=3,
)

def compute_metrics(eval_pred):
    logits, labels = eval_pred
    predictions = np.argmax(logits, axis=-1)
    return {
        "accuracy": accuracy_score(labels, predictions),
        "f1_macro": f1_score(labels, predictions, average="macro"),
        "f1_weighted": f1_score(labels, predictions, average="weighted"),
    }

training_args = TrainingArguments(
    output_dir="./sentiment-phobert",
    num_train_epochs=5,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=32,
    learning_rate=2e-5,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    metric_for_best_model="f1_macro",
    fp16=True,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
    eval_dataset=tokenized["test"],
    compute_metrics=compute_metrics,
)

trainer.train()

2.4 Evaluation

from sklearn.metrics import classification_report, confusion_matrix
import seaborn as sns
import matplotlib.pyplot as plt

# Predictions
preds = trainer.predict(tokenized["test"])
y_pred = np.argmax(preds.predictions, axis=-1)
y_true = preds.label_ids

# Classification Report
labels = ["Negative", "Neutral", "Positive"]
print(classification_report(y_true, y_pred, target_names=labels))

# Confusion Matrix
cm = confusion_matrix(y_true, y_pred)
sns.heatmap(cm, annot=True, fmt='d', xticklabels=labels, yticklabels=labels)
plt.xlabel("Predicted")
plt.ylabel("True")
plt.title("Confusion Matrix")
plt.show()

3. Deploy with FastAPI

from fastapi import FastAPI
from transformers import pipeline
from pydantic import BaseModel

app = FastAPI()

# Load model
classifier = pipeline(
    "sentiment-analysis",
    model="./sentiment-phobert",
    tokenizer="vinai/phobert-base-v2",
)

class TextInput(BaseModel):
    text: str

@app.post("/predict")
def predict(input_data: TextInput):
    result = classifier(input_data.text)
    return {"sentiment": result[0]["label"], "score": result[0]["score"]}

Summary

StepTools
Data preparationpandas, datasets
TokenizationAutoTokenizer (PhoBERT)
TrainingTrainer API
Evaluationsklearn metrics, confusion matrix
DeploymentFastAPI + pipeline

Next article

Lesson 14: Named Entity Recognition (NER) — Extract entities: people, organizations, locations from text.