Lesson 19: Modern LLM & NLP — RAGs, Agents, and 2026 Trends
From traditional NLP to LLM era. Retrieval-Augmented Generation. In-context learning vs fine-tuning. Prompt engineering for NLP tasks. AI Agents for NLP workflows. Multimodal NLP. Trends: small language models, synthetic data, constitutional AI.
NLP from Basics to Advanced: Mastering Natural Language Processing
Part 6: NLP Production & Modern Trends
xdev.asia
Introduction
NLP in 2026 is completely different than it was 5 years ago. LLMs have changed the way almost every NLP problem is approached. This article summarizes the most modern trends and techniques.
1. Traditional NLP vs LLM Era
Traditional
LLM Era
Each task needs its own model
An LLM solves many tasks
Need labeled data
Zero/few-shot, prompt engineering
Train → Evaluate → Deploy
Prompt → Test → RAG/Fine-tune → Deploy
BERT + task-specific head
GPT-4 / Gemini + prompt
Weeks to build
Hours to prototype
When to still use traditional NLP?
Latency critical: BERT inference ~5ms vs LLM ~500ms
Cost sensitive: Fine-tuned small model << LLM API
Offline: On-device, no internet
Specific domain: When very high precision is needed (medical, legal)
from openai import OpenAI
client = OpenAI()
# NER bằng prompt (không cần train!)
def extract_entities(text):
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{
"role": "system",
"content": """Extract named entities from Vietnamese text.
Return JSON: {"persons": [], "organizations": [], "locations": [], "dates": []}"""
}, {
"role": "user",
"content": text
}],
response_format={"type": "json_object"},
)
return response.choices[0].message.content
# Classification bằng prompt
def classify_text(text, categories):
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{
"role": "system",
"content": f"Classify text into one of: {categories}. Return only the category name."
}, {
"role": "user",
"content": text
}],
)
return response.choices[0].message.content
4. AI Agents cho NLP Workflows
# Agent tự động phân tích document
# 1. Extract entities → 2. Classify sentiment → 3. Summarize → 4. Store results
from langchain.agents import AgentExecutor, create_openai_tools_agent
from langchain.tools import tool
@tool
def extract_entities_tool(text: str) -> dict:
"""Extract named entities from text."""
ner = pipeline("ner", grouped_entities=True)
return ner(text)
@tool
def classify_sentiment_tool(text: str) -> str:
"""Classify text sentiment."""
classifier = pipeline("sentiment-analysis")
return classifier(text)[0]
# Agent combines many NLP tools
# → Decide which tools to use and in what order
5. NLP Trends 2026
5.1 Small Language Models (SLMs)
Phi-3, Gemma 2, LLaMA 3.2 (1B-7B params)
Can run on laptop, mobile
Fine-tune easily on consumer GPUs
Good enough for many NLP tasks
5.2 Multimodal NLP
GPT-4o, Gemini: text + image + audio + video
NLP is not just about text anymore — multimodal understanding
Document AI: OCR + NLP cho invoice, form, report
5.3 Synthetic Data
Use large LLM to generate training data for small models
Reduce labeling costs 10-100x
Quality control: LLM-as-judge
5.4 Structured Generation
# Make sure LLM output is always in the correct format
from pydantic import BaseModel
class NEROutput(BaseModel):
persons: list[str]
organizations: list[str]
locations: list[str]
# With Instructor, Outlines, or JSON mode
6. Decision Framework: Which approach to choose?
Do you need to solve an NLP task?
│
├── Fast prototype? ──→ LLM API + Prompt Engineering
│
├── Cost sensitive? ──→ Fine-tune small model (BERT/PhoBERT)
│
├── Need knowledge base? ──→ RAG Pipeline
│
├── Latency < 50ms? ──→ Distilled/Quantized model
│
└── Complex workflow? ──→ AI Agent + NLP Tools
Summary
Trends
Meaning
LLM-first
Prototype with LLM, optimize later
RAG
Combining retrieval + generation
SLMs
Small but powerful, runs edge
Multimodal
Text + Image + Audio
Agents
Automate NLP workflows
Next article
Lesson 20: Capstone Project — Building an end-to-end NLP Platform: classification + NER + QA for a real domain.