Chuyển đến nội dung chính

Lesson 19: Modern LLM & NLP — RAGs, Agents, and 2026 Trends

From traditional NLP to LLM era. Retrieval-Augmented Generation. In-context learning vs fine-tuning. Prompt engineering for NLP tasks. AI Agents for NLP workflows. Multimodal NLP. Trends: small language models, synthetic data, constitutional AI.

🧠 AI & ML — Lesson 18 Lesson 19: Modern LLM & NLP — RAG, Agents, and Trends 2026

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 6: NLP Production & Modern Trends

xdev.asia

Introduction

NLP in 2026 is completely different than it was 5 years ago. LLMs have changed the way almost every NLP problem is approached. This article summarizes the most modern trends and techniques.


1. Traditional NLP vs LLM Era

TraditionalLLM Era
Each task needs its own modelAn LLM solves many tasks
Need labeled dataZero/few-shot, prompt engineering
Train → Evaluate → DeployPrompt → Test → RAG/Fine-tune → Deploy
BERT + task-specific headGPT-4 / Gemini + prompt
Weeks to buildHours to prototype

When to still use traditional NLP?

  • Latency critical: BERT inference ~5ms vs LLM ~500ms
  • Cost sensitive: Fine-tuned small model << LLM API
  • Offline: On-device, no internet
  • Specific domain: When very high precision is needed (medical, legal)

2. Retrieval-Augmented Generation (RAG)

┌────────────────────────────────────────────────────────┐
│                    RAG PIPELINE                         │
│                                                        │
│  User Query                                            │
│      │                                                 │
│      ▼                                                 │
│  ┌──────────┐    ┌───────────────┐                    │
│  │  Embed   │───▶│ Vector Search │── Top-K docs       │
│  │  Query   │    │ (FAISS/PGVector)│                   │
│  └──────────┘    └───────────────┘                    │
│                         │                              │
│                         ▼                              │
│  ┌──────────────────────────────────────────┐         │
│  │  LLM (GPT-4 / Gemini)                    │         │
│  │  System: "Answer based on context below"  │         │
│  │  Context: [retrieved documents]            │         │
│  │  Question: [user query]                    │         │
│  └──────────────────────────────────────────┘         │
│                         │                              │
│                         ▼                              │
│                     Answer                             │
└────────────────────────────────────────────────────────┘

RAG cho NLP Tasks

Before (train model)After (RAG)
Fine-tune BERT cho QARAG + LLM: retrieve docs → generate answer
Train classifier on labeled dataFew-shot examples + LLM
Build NER pipelineLLM extract entities with prompt

3. Prompt Engineering cho NLP Tasks

from openai import OpenAI

client = OpenAI()

# NER bằng prompt (không cần train!)
def extract_entities(text):
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{
            "role": "system",
            "content": """Extract named entities from Vietnamese text.
Return JSON: {"persons": [], "organizations": [], "locations": [], "dates": []}"""
        }, {
            "role": "user",
            "content": text
        }],
        response_format={"type": "json_object"},
    )
    return response.choices[0].message.content

# Classification bằng prompt
def classify_text(text, categories):
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{
            "role": "system",
            "content": f"Classify text into one of: {categories}. Return only the category name."
        }, {
            "role": "user",
            "content": text
        }],
    )
    return response.choices[0].message.content

4. AI Agents cho NLP Workflows

# Agent tự động phân tích document
# 1. Extract entities → 2. Classify sentiment → 3. Summarize → 4. Store results

from langchain.agents import AgentExecutor, create_openai_tools_agent
from langchain.tools import tool

@tool
def extract_entities_tool(text: str) -> dict:
    """Extract named entities from text."""
    ner = pipeline("ner", grouped_entities=True)
    return ner(text)

@tool
def classify_sentiment_tool(text: str) -> str:
    """Classify text sentiment."""
    classifier = pipeline("sentiment-analysis")
    return classifier(text)[0]

# Agent combines many NLP tools
# → Decide which tools to use and in what order

5. NLP Trends 2026

5.1 Small Language Models (SLMs)

  • Phi-3, Gemma 2, LLaMA 3.2 (1B-7B params)
  • Can run on laptop, mobile
  • Fine-tune easily on consumer GPUs
  • Good enough for many NLP tasks

5.2 Multimodal NLP

  • GPT-4o, Gemini: text + image + audio + video
  • NLP is not just about text anymore — multimodal understanding
  • Document AI: OCR + NLP cho invoice, form, report

5.3 Synthetic Data

  • Use large LLM to generate training data for small models
  • Reduce labeling costs 10-100x
  • Quality control: LLM-as-judge

5.4 Structured Generation

# Make sure LLM output is always in the correct format
from pydantic import BaseModel

class NEROutput(BaseModel):
    persons: list[str]
    organizations: list[str]
    locations: list[str]

# With Instructor, Outlines, or JSON mode

6. Decision Framework: Which approach to choose?

Do you need to solve an NLP task?
    │
    ├── Fast prototype? ──→ LLM API + Prompt Engineering
    │
    ├── Cost sensitive? ──→ Fine-tune small model (BERT/PhoBERT)
    │
    ├── Need knowledge base? ──→ RAG Pipeline
    │
    ├── Latency < 50ms? ──→ Distilled/Quantized model
    │
    └── Complex workflow? ──→ AI Agent + NLP Tools

Summary

TrendsMeaning
LLM-firstPrototype with LLM, optimize later
RAGCombining retrieval + generation
SLMsSmall but powerful, runs edge
MultimodalText + Image + Audio
AgentsAutomate NLP workflows

Next article

Lesson 20: Capstone Project — Building an end-to-end NLP Platform: classification + NER + QA for a real domain.