Chuyển đến nội dung chính

Lesson 18: NLP Pipeline Production — MLOps for NLP

Production NLP pipeline: data ingestion → preprocessing → inference → post-processing. Model serving: FastAPI, Triton, vLLM. Monitoring: data drift, model drift. CI/CD for NLP models. Logging and error handling. Scaling considerations.

🧠 AI & ML — Lesson 17 Lesson 18: NLP Pipeline Production — MLOps for NLP

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 6: NLP Production & Modern Trends

xdev.asia

Introduction

Building a highly accurate NLP model is only 30% of the work — the remaining 70% is putting it into production, monitoring, and maintaining. This article guides on building NLP pipeline production-ready.


1. Production NLP Pipeline

┌─────────────────────────────────────────────────────────────┐
│                    NLP PRODUCTION PIPELINE                   │
│                                                             │
│  ┌──────────┐   ┌──────────┐   ┌──────────┐   ┌─────────┐ │
│  │  Input    │──▶│  Preproc │──▶│  Model   │──▶│  Post   │ │
│  │  API      │   │  Engine  │   │  Server  │   │  Proc   │ │
│  └──────────┘   └──────────┘   └──────────┘   └────┬────┘ │
│       │              │              │               │       │
│       ▼              ▼              ▼               ▼       │
│  ┌──────────────────────────────────────────────────────┐   │
│  │              Monitoring & Logging                     │   │
│  │  • Latency  • Throughput  • Data Drift  • Errors     │   │
│  └──────────────────────────────────────────────────────┘   │
└─────────────────────────────────────────────────────────────┘

2. Model Serving with FastAPI

from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from transformers import pipeline
import logging
import time

app = FastAPI(title="NLP API")
logger = logging.getLogger(__name__)

# Load models on startup
classifier = None

@app.on_event("startup")
async def load_models():
    global classifier
    classifier = pipeline(
        "sentiment-analysis",
        model="./models/sentiment-phobert",
        device=0,  # GPU
    )
    logger.info("Models loaded successfully")

class TextRequest(BaseModel):
    text: str
    max_length: int = 512

class PredictionResponse(BaseModel):
    label: str
    score: float
    latency_ms: float

@app.post("/predict", response_model=PredictionResponse)
async def predict(request: TextRequest):
    start = time.time()

    if not request.text.strip():
        raise HTTPException(status_code=400, detail="Text cannot be empty")

    # Truncate input
    text = request.text[:request.max_length]

    result = classifier(text)[0]

    latency = (time.time() - start) * 1000

    logger.info(f"Prediction: {result['label']} ({result['score']:.4f}) in {latency:.1f}ms")

    return PredictionResponse(
        label=result["label"],
        score=result["score"],
        latency_ms=round(latency, 2),
    )

@app.get("/health")
async def health():
    return {"status": "healthy", "model_loaded": classifier is not None}

3. Batch Processing

from transformers import pipeline

classifier = pipeline("sentiment-analysis", device=0, batch_size=32)

# Batch inference — nhanh hơn nhiều so với từng câu
texts = ["Text 1...", "Text 2...", ...]  # Hàng nghìn texts
results = classifier(texts)  # Tự động batch

4. Model Optimization

4.1 ONNX Runtime

from optimum.onnxruntime import ORTModelForSequenceClassification
from transformers import AutoTokenizer

# Export và load ONNX model
model = ORTModelForSequenceClassification.from_pretrained(
    "./models/sentiment-phobert",
    export=True,
)
tokenizer = AutoTokenizer.from_pretrained("./models/sentiment-phobert")

# Inference nhanh hơn 2-5x
inputs = tokenizer("NLP rất thú vị", return_tensors="pt")
outputs = model(**inputs)

4.2 Quantization

from optimum.onnxruntime import ORTQuantizer
from optimum.onnxruntime.configuration import AutoQuantizationConfig

quantizer = ORTQuantizer.from_pretrained(model)
qconfig = AutoQuantizationConfig.avx512_vnni(is_static=False)
quantizer.quantize(save_dir="./quantized-model", quantization_config=qconfig)
# Model size giảm ~4x, speed tăng ~2x

5. Monitoring

Data Drift Detection

from scipy.stats import ks_2samp
import numpy as np

def detect_text_drift(reference_lengths, current_lengths, threshold=0.05):
    """Detect data drift bằng KS test trên text length distribution."""
    stat, p_value = ks_2samp(reference_lengths, current_lengths)
    is_drift = p_value < threshold
    return {
        "is_drift": is_drift,
        "ks_statistic": stat,
        "p_value": p_value,
    }

Metrics Dashboard

import prometheus_client as prom

# Prometheus metrics
PREDICTION_COUNTER = prom.Counter(
    'nlp_predictions_total',
    'Total predictions',
    ['model', 'label']
)
PREDICTION_LATENCY = prom.Histogram(
    'nlp_prediction_latency_seconds',
    'Prediction latency',
    ['model']
)
CONFIDENCE_HISTOGRAM = prom.Histogram(
    'nlp_confidence_score',
    'Prediction confidence distribution',
    ['model'],
    buckets=[0.5, 0.6, 0.7, 0.8, 0.9, 0.95, 0.99, 1.0]
)

6. CI/CD for NLP Models

# .github/workflows/nlp-pipeline.yml
name: NLP Model CI/CD

on:
  push:
    paths: ['models/**', 'src/**']

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Run unit tests
        run: pytest tests/ -v
      - name: Run model quality tests
        run: python scripts/evaluate_model.py --threshold 0.85
      - name: Check for data drift
        run: python scripts/check_drift.py

Summary

AspectTools/Practice
ServingFastAPI, Triton, vLLM
OptimizationONNX, quantization, batching
MonitoringPrometheus, drift detection
CI/CDGitHub Actions, model quality gates
ScalingLoad balancer, horizontal scaling

Next article

Lesson 19: LLM & Modern NLP — RAG, Agents, and NLP trends in 2026.