Chuyển đến nội dung chính

Lesson 1: What is NLP? — Overview of the field of Natural Language Processing

Definition of NLP, history of development from rule-based to deep learning. Core problems: classification, NER, POS tagging, parsing, generation, QA, summarization. NLP pipeline overview. Simple end-to-end demo with Python.

🧠 AI & ML — Lesson 0 Lesson 1: What is NLP? — Overview of the Processing field Natural Language Physics

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 1: NLP Foundations — Understanding Language Through a Computer Lens

xdev.asia

Introduction

Natural Language Processing (NLP) — Natural Language Processing — is a field at the intersection of Computer Science, Artificial Intelligence and Linguistics, studying how to help computers understand, analyze and generate human language.

💡 One sentence: NLP teaches computers to "read", "understand" and "write" natural language.

From Google Search, Gmail Smart Compose, ChatGPT to virtual assistant Siri — all are based on NLP.


1. Where is NLP in AI?

┌─────────────────────────────────────────────────────────┐
│                  ARTIFICIAL INTELLIGENCE                 │
│                                                         │
│   ┌─────────────────────────────────────────────────┐   │
│   │              MACHINE LEARNING                    │   │
│   │                                                  │   │
│   │   ┌────────────────────────────────────────┐    │   │
│   │   │          DEEP LEARNING                  │    │   │
│   │   │                                         │    │   │
│   │   │   ┌──────────┐  ┌──────────────────┐   │    │   │
│   │   │   │    NLP    │  │ Computer Vision  │   │    │   │
│   │   │   │          │  │                  │   │    │   │
│   │   │   │ • Text   │  │ • Image          │   │    │   │
│   │   │   │ • Speech │  │ • Video          │   │    │   │
│   │   │   └──────────┘  └──────────────────┘   │    │   │
│   │   └────────────────────────────────────────┘    │   │
│   └─────────────────────────────────────────────────┘   │
└─────────────────────────────────────────────────────────┘

2. History of NLP — From Rule-based to Transformer

PeriodMethodFeaturesExample
1950s–1980sRule-basedWrite regex, grammar rules manuallyELIZA chatbot
1990s–2000sStatisticalProbability, n-gram, HMM, CRFSpam filters, POS tagging
2010sML/Deep LearningWord2Vec, RNN, LSTM, CNNSentiment analysis
2017TransformerSelf-attention, parallelizationImproved Google Translate
2018–nowPre-trained LMsBERT, GPT, T5, LLaMAChatGPT, Gemini, Claude

Important turning point

  1. 2013 — Word2Vec: First time representing words using dense meaningful vectors
  2. 2017 — Transformer: "Attention Is All You Need" changes everything
  3. 2018 — BERT & GPT: Transfer learning for NLP, pre-train once, use everywhere
  4. 2022 to present — LLM era: ChatGPT, emergent ability, reasoning

3. Core problems in NLP

3.1 Classification by level

┌────────────────────────────────────────────────────────┐
│                    CÁC BÀI TOÁN NLP                    │
│                                                        │
│  📝 CẤP ĐỘ TỪ (Token-level)                          │
│  ├── POS Tagging: gán nhãn từ loại (danh từ, động từ) │
│  ├── NER: trích xuất thực thể (người, địa điểm, tổ chức)│
│  └── Word Segmentation: tách từ (quan trọng cho tiếng Việt)│
│                                                        │
│  📄 CẤP ĐỘ CÂU/TÀI LIỆU (Sequence-level)            │
│  ├── Text Classification: phân loại văn bản            │
│  ├── Sentiment Analysis: phân tích cảm xúc             │
│  └── Topic Modeling: phát hiện chủ đề                  │
│                                                        │
│  🔄 CẤP ĐỘ SINH (Generation)                          │
│  ├── Machine Translation: dịch máy                     │
│  ├── Text Summarization: tóm tắt                       │
│  ├── Question Answering: hỏi đáp                       │
│  └── Text Generation: sinh văn bản (ChatGPT, Gemini)   │
│                                                        │
│  🔗 CẤP ĐỘ QUAN HỆ (Relation)                        │
│  ├── Semantic Similarity: đo độ tương đồng nghĩa       │
│  ├── Textual Entailment: suy luận logic                 │
│  └── Coreference Resolution: xác định đại từ           │
└────────────────────────────────────────────────────────┘

3.2 Practical example

Math problemInputOutputApplication
Sentiment Analysis"This product is great!"Positive (0.95)Review monitoring
NER"Nguyen Van A works at FPT"PER: Nguyen Van A, ORG: FPTExtract information
Translation"Hello, how are you?""Hello, how are you?"Google Translate
Summarization1000 word article50 word summaryAuto news
QAContext + "Who is Apple's CEO?""Tim Cook"Chatbots, search

4. NLP Pipeline overview

No matter what problem you solve, the basic NLP pipeline includes the following steps:

Input Text
    │
    ▼
┌──────────────────┐
│  1. Preprocessing │ ← Tokenization, cleaning, normalization
└────────┬─────────┘
         │
         ▼
┌──────────────────┐
│  2. Representation│ ← BoW, TF-IDF, Word Embeddings, BERT
└────────┬─────────┘
         │
         ▼
┌──────────────────┐
│  3. Modeling      │ ← ML/DL model: classification, NER, generation
└────────┬─────────┘
         │
         ▼
┌──────────────────┐
│  4. Post-processing│ ← Decode, format, threshold, filter
└────────┬─────────┘
         │
         ▼
    Output

5. Demo: NLP in 5 minutes with Python

# Cài đặt: pip install transformers torch

from transformers import pipeline

# 1. Sentiment Analysis
sentiment = pipeline("sentiment-analysis")
result = sentiment("Khóa học NLP này thực sự hay quá!")
print(result)
# [{'label': 'POSITIVE', 'score': 0.9998}]

# 2. Named Entity Recognition
ner = pipeline("ner", grouped_entities=True)
entities = ner("Elon Musk là CEO của Tesla và SpaceX tại California")
for e in entities:
    print(f"  {e['word']}: {e['entity_group']} ({e['score']:.2f})")
# Elon Musk: PER (0.99)
# Tesla: ORG (0.98)
# SpaceX: ORG (0.97)
# California: LOC (0.99)

# 3. Question Answering
qa = pipeline("question-answering")
answer = qa(
    question="NLP là gì?",
    context="NLP (Natural Language Processing) là lĩnh vực AI giúp máy tính hiểu ngôn ngữ tự nhiên."
)
print(f"Answer: {answer['answer']} (score: {answer['score']:.2f})")

# 4. Summarization
summarizer = pipeline("summarization")
summary = summarizer("Your long text here...", max_length=50)
print(summary)

# 5. Translation
translator = pipeline("translation_en_to_vi", model="Helsinki-NLP/opus-mt-en-vi")
result = translator("Natural Language Processing is amazing!")
print(result)

🎯 Only with Hugging Face pipeline, you have run 5 different NLP problems — no need to understand the theory!


6. NLP for Vietnamese — Quick Overview

Vietnamese has specific challenges:

ChallengeExampleSolution
Word segmentation"student" vs "study" + "student"VnCoreNLP, underthesea
Bar mark"study" ≠ "learn" ≠ "paint"Accent normalization
Few resourcesFewer datasets compared to EnglishPhoBERT, ViT5, VLSP datasets
Compound words"computer", "keyboard"Dictionary-based segmentation
# Demo NLP tiếng Việt với underthesea
from underthesea import word_tokenize, pos_tag, ner

text = "Nguyễn Phú Trọng làm việc tại Hà Nội"

# Word segmentation
print(word_tokenize(text))
# ['Nguyễn_Phú_Trọng', 'làm_việc', 'tại', 'Hà_Nội']

# POS Tagging
print(pos_tag(text))
# [('Nguyễn_Phú_Trọng', 'Np'), ('làm_việc', 'V'), ('tại', 'E'), ('Hà_Nội', 'Np')]

📌 Lesson 17 will delve into NLP for Vietnamese.


Summary

ConceptMeaning
NLPThe field of AI helps computers understand natural language
HistoryRule-based → Statistical → Deep Learning → Transformer → LLM
Math problemClassification, NER, QA, Summarization, Translation, Generation
PipelinesPreprocessing → Representation → Modeling → Post-processing
VietnameseWord segmentation challenge, low resources but growing

Next article

Lesson 2: Text Preprocessing — Delve into the first and most important step: Tokenization, cleaning, normalization. "Garbage in, garbage out" — data preprocessing determines 80% of success in NLP.