Chuyển đến nội dung chính

Bài 18: NLP Pipeline Production — MLOps cho NLP

Production NLP pipeline: data ingestion → preprocessing → inference → post-processing. Model serving: FastAPI, Triton, vLLM. Monitoring: data drift, model drift. CI/CD cho NLP models. Logging và error handling. Scaling considerations.

🧠 AI & ML — Bài 17 Bài 18: NLP Pipeline Production — MLOps cho NLP

NLP từ Cơ bản đến Nâng cao: Làm chủ Xử lý Ngôn ngữ Tự nhiên

Phần 6: NLP Production & Xu hướng Hiện đại

xdev.asia

Giới thiệu

Xây model NLP đạt accuracy cao mới chỉ là 30% công việc — 70% còn lại là đưa lên production, monitoring, và maintain. Bài này hướng dẫn xây dựng NLP pipeline production-ready.


1. Production NLP Pipeline

┌─────────────────────────────────────────────────────────────┐
│                    NLP PRODUCTION PIPELINE                   │
│                                                             │
│  ┌──────────┐   ┌──────────┐   ┌──────────┐   ┌─────────┐ │
│  │  Input    │──▶│  Preproc │──▶│  Model   │──▶│  Post   │ │
│  │  API      │   │  Engine  │   │  Server  │   │  Proc   │ │
│  └──────────┘   └──────────┘   └──────────┘   └────┬────┘ │
│       │              │              │               │       │
│       ▼              ▼              ▼               ▼       │
│  ┌──────────────────────────────────────────────────────┐   │
│  │              Monitoring & Logging                     │   │
│  │  • Latency  • Throughput  • Data Drift  • Errors     │   │
│  └──────────────────────────────────────────────────────┘   │
└─────────────────────────────────────────────────────────────┘

2. Model Serving với FastAPI

from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from transformers import pipeline
import logging
import time

app = FastAPI(title="NLP API")
logger = logging.getLogger(__name__)

# Load models on startup
classifier = None

@app.on_event("startup")
async def load_models():
    global classifier
    classifier = pipeline(
        "sentiment-analysis",
        model="./models/sentiment-phobert",
        device=0,  # GPU
    )
    logger.info("Models loaded successfully")

class TextRequest(BaseModel):
    text: str
    max_length: int = 512

class PredictionResponse(BaseModel):
    label: str
    score: float
    latency_ms: float

@app.post("/predict", response_model=PredictionResponse)
async def predict(request: TextRequest):
    start = time.time()

    if not request.text.strip():
        raise HTTPException(status_code=400, detail="Text cannot be empty")

    # Truncate input
    text = request.text[:request.max_length]

    result = classifier(text)[0]

    latency = (time.time() - start) * 1000

    logger.info(f"Prediction: {result['label']} ({result['score']:.4f}) in {latency:.1f}ms")

    return PredictionResponse(
        label=result["label"],
        score=result["score"],
        latency_ms=round(latency, 2),
    )

@app.get("/health")
async def health():
    return {"status": "healthy", "model_loaded": classifier is not None}

3. Batch Processing

from transformers import pipeline

classifier = pipeline("sentiment-analysis", device=0, batch_size=32)

# Batch inference — nhanh hơn nhiều so với từng câu
texts = ["Text 1...", "Text 2...", ...]  # Hàng nghìn texts
results = classifier(texts)  # Tự động batch

4. Model Optimization

4.1 ONNX Runtime

from optimum.onnxruntime import ORTModelForSequenceClassification
from transformers import AutoTokenizer

# Export và load ONNX model
model = ORTModelForSequenceClassification.from_pretrained(
    "./models/sentiment-phobert",
    export=True,
)
tokenizer = AutoTokenizer.from_pretrained("./models/sentiment-phobert")

# Inference nhanh hơn 2-5x
inputs = tokenizer("NLP rất thú vị", return_tensors="pt")
outputs = model(**inputs)

4.2 Quantization

from optimum.onnxruntime import ORTQuantizer
from optimum.onnxruntime.configuration import AutoQuantizationConfig

quantizer = ORTQuantizer.from_pretrained(model)
qconfig = AutoQuantizationConfig.avx512_vnni(is_static=False)
quantizer.quantize(save_dir="./quantized-model", quantization_config=qconfig)
# Model size giảm ~4x, speed tăng ~2x

5. Monitoring

Data Drift Detection

from scipy.stats import ks_2samp
import numpy as np

def detect_text_drift(reference_lengths, current_lengths, threshold=0.05):
    """Detect data drift bằng KS test trên text length distribution."""
    stat, p_value = ks_2samp(reference_lengths, current_lengths)
    is_drift = p_value < threshold
    return {
        "is_drift": is_drift,
        "ks_statistic": stat,
        "p_value": p_value,
    }

Metrics Dashboard

import prometheus_client as prom

# Prometheus metrics
PREDICTION_COUNTER = prom.Counter(
    'nlp_predictions_total',
    'Total predictions',
    ['model', 'label']
)
PREDICTION_LATENCY = prom.Histogram(
    'nlp_prediction_latency_seconds',
    'Prediction latency',
    ['model']
)
CONFIDENCE_HISTOGRAM = prom.Histogram(
    'nlp_confidence_score',
    'Prediction confidence distribution',
    ['model'],
    buckets=[0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.99, 1.0]
)

6. CI/CD cho NLP Models

# .github/workflows/nlp-pipeline.yml
name: NLP Model CI/CD

on:
  push:
    paths: ['models/**', 'src/**']

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Run unit tests
        run: pytest tests/ -v
      - name: Run model quality tests
        run: python scripts/evaluate_model.py --threshold 0.85
      - name: Check for data drift
        run: python scripts/check_drift.py

Tổng kết

Khía cạnhTools/Practices
ServingFastAPI, Triton, vLLM
OptimizationONNX, quantization, batching
MonitoringPrometheus, drift detection
CI/CDGitHub Actions, model quality gates
ScalingLoad balancer, horizontal scaling

Bài tiếp theo

Bài 19: LLM & NLP Hiện đại — RAG, Agents, và xu hướng NLP năm 2026.