Introduction
Natural Language Processing (NLP) — Natural Language Processing — is a field at the intersection of Computer Science, Artificial Intelligence and Linguistics, studying how to help computers understand, analyze and generate human language.
💡 One sentence: NLP teaches computers to "read", "understand" and "write" natural language.
From Google Search, Gmail Smart Compose, ChatGPT to virtual assistant Siri — all are based on NLP.
1. Where is NLP in AI?
┌─────────────────────────────────────────────────────────┐
│ ARTIFICIAL INTELLIGENCE │
│ │
│ ┌─────────────────────────────────────────────────┐ │
│ │ MACHINE LEARNING │ │
│ │ │ │
│ │ ┌────────────────────────────────────────┐ │ │
│ │ │ DEEP LEARNING │ │ │
│ │ │ │ │ │
│ │ │ ┌──────────┐ ┌──────────────────┐ │ │ │
│ │ │ │ NLP │ │ Computer Vision │ │ │ │
│ │ │ │ │ │ │ │ │ │
│ │ │ │ • Text │ │ • Image │ │ │ │
│ │ │ │ • Speech │ │ • Video │ │ │ │
│ │ │ └──────────┘ └──────────────────┘ │ │ │
│ │ └────────────────────────────────────────┘ │ │
│ └─────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────┘
2. History of NLP — From Rule-based to Transformer
| Period | Method | Features | Example |
|---|---|---|---|
| 1950s–1980s | Rule-based | Write regex, grammar rules manually | ELIZA chatbot |
| 1990s–2000s | Statistical | Probability, n-gram, HMM, CRF | Spam filters, POS tagging |
| 2010s | ML/Deep Learning | Word2Vec, RNN, LSTM, CNN | Sentiment analysis |
| 2017 | Transformer | Self-attention, parallelization | Improved Google Translate |
| 2018–now | Pre-trained LMs | BERT, GPT, T5, LLaMA | ChatGPT, Gemini, Claude |
Important turning point
- 2013 — Word2Vec: First time representing words using dense meaningful vectors
- 2017 — Transformer: "Attention Is All You Need" changes everything
- 2018 — BERT & GPT: Transfer learning for NLP, pre-train once, use everywhere
- 2022 to present — LLM era: ChatGPT, emergent ability, reasoning
3. Core problems in NLP
3.1 Classification by level
┌────────────────────────────────────────────────────────┐
│ CÁC BÀI TOÁN NLP │
│ │
│ 📝 CẤP ĐỘ TỪ (Token-level) │
│ ├── POS Tagging: gán nhãn từ loại (danh từ, động từ) │
│ ├── NER: trích xuất thực thể (người, địa điểm, tổ chức)│
│ └── Word Segmentation: tách từ (quan trọng cho tiếng Việt)│
│ │
│ 📄 CẤP ĐỘ CÂU/TÀI LIỆU (Sequence-level) │
│ ├── Text Classification: phân loại văn bản │
│ ├── Sentiment Analysis: phân tích cảm xúc │
│ └── Topic Modeling: phát hiện chủ đề │
│ │
│ 🔄 CẤP ĐỘ SINH (Generation) │
│ ├── Machine Translation: dịch máy │
│ ├── Text Summarization: tóm tắt │
│ ├── Question Answering: hỏi đáp │
│ └── Text Generation: sinh văn bản (ChatGPT, Gemini) │
│ │
│ 🔗 CẤP ĐỘ QUAN HỆ (Relation) │
│ ├── Semantic Similarity: đo độ tương đồng nghĩa │
│ ├── Textual Entailment: suy luận logic │
│ └── Coreference Resolution: xác định đại từ │
└────────────────────────────────────────────────────────┘
3.2 Practical example
| Math problem | Input | Output | Application |
|---|---|---|---|
| Sentiment Analysis | "This product is great!" | Positive (0.95) | Review monitoring |
| NER | "Nguyen Van A works at FPT" | PER: Nguyen Van A, ORG: FPT | Extract information |
| Translation | "Hello, how are you?" | "Hello, how are you?" | Google Translate |
| Summarization | 1000 word article | 50 word summary | Auto news |
| QA | Context + "Who is Apple's CEO?" | "Tim Cook" | Chatbots, search |
4. NLP Pipeline overview
No matter what problem you solve, the basic NLP pipeline includes the following steps:
Input Text
│
▼
┌──────────────────┐
│ 1. Preprocessing │ ← Tokenization, cleaning, normalization
└────────┬─────────┘
│
▼
┌──────────────────┐
│ 2. Representation│ ← BoW, TF-IDF, Word Embeddings, BERT
└────────┬─────────┘
│
▼
┌──────────────────┐
│ 3. Modeling │ ← ML/DL model: classification, NER, generation
└────────┬─────────┘
│
▼
┌──────────────────┐
│ 4. Post-processing│ ← Decode, format, threshold, filter
└────────┬─────────┘
│
▼
Output
5. Demo: NLP in 5 minutes with Python
# Cài đặt: pip install transformers torch
from transformers import pipeline
# 1. Sentiment Analysis
sentiment = pipeline("sentiment-analysis")
result = sentiment("Khóa học NLP này thực sự hay quá!")
print(result)
# [{'label': 'POSITIVE', 'score': 0.9998}]
# 2. Named Entity Recognition
ner = pipeline("ner", grouped_entities=True)
entities = ner("Elon Musk là CEO của Tesla và SpaceX tại California")
for e in entities:
print(f" {e['word']}: {e['entity_group']} ({e['score']:.2f})")
# Elon Musk: PER (0.99)
# Tesla: ORG (0.98)
# SpaceX: ORG (0.97)
# California: LOC (0.99)
# 3. Question Answering
qa = pipeline("question-answering")
answer = qa(
question="NLP là gì?",
context="NLP (Natural Language Processing) là lĩnh vực AI giúp máy tính hiểu ngôn ngữ tự nhiên."
)
print(f"Answer: {answer['answer']} (score: {answer['score']:.2f})")
# 4. Summarization
summarizer = pipeline("summarization")
summary = summarizer("Your long text here...", max_length=50)
print(summary)
# 5. Translation
translator = pipeline("translation_en_to_vi", model="Helsinki-NLP/opus-mt-en-vi")
result = translator("Natural Language Processing is amazing!")
print(result)
🎯 Only with Hugging Face
pipeline, you have run 5 different NLP problems — no need to understand the theory!
6. NLP for Vietnamese — Quick Overview
Vietnamese has specific challenges:
| Challenge | Example | Solution |
|---|---|---|
| Word segmentation | "student" vs "study" + "student" | VnCoreNLP, underthesea |
| Bar mark | "study" ≠ "learn" ≠ "paint" | Accent normalization |
| Few resources | Fewer datasets compared to English | PhoBERT, ViT5, VLSP datasets |
| Compound words | "computer", "keyboard" | Dictionary-based segmentation |
# Demo NLP tiếng Việt với underthesea
from underthesea import word_tokenize, pos_tag, ner
text = "Nguyễn Phú Trọng làm việc tại Hà Nội"
# Word segmentation
print(word_tokenize(text))
# ['Nguyễn_Phú_Trọng', 'làm_việc', 'tại', 'Hà_Nội']
# POS Tagging
print(pos_tag(text))
# [('Nguyễn_Phú_Trọng', 'Np'), ('làm_việc', 'V'), ('tại', 'E'), ('Hà_Nội', 'Np')]
📌 Lesson 17 will delve into NLP for Vietnamese.
Summary
| Concept | Meaning |
|---|---|
| NLP | The field of AI helps computers understand natural language |
| History | Rule-based → Statistical → Deep Learning → Transformer → LLM |
| Math problem | Classification, NER, QA, Summarization, Translation, Generation |
| Pipelines | Preprocessing → Representation → Modeling → Post-processing |
| Vietnamese | Word segmentation challenge, low resources but growing |
Next article
Lesson 2: Text Preprocessing — Delve into the first and most important step: Tokenization, cleaning, normalization. "Garbage in, garbage out" — data preprocessing determines 80% of success in NLP.