Introduction
"Garbage in, garbage out" — In NLP, 80% of success comes from preparing and cleaning text data.
Text data in "real life" is always dirty: contains HTML tags, emojis, special characters, abbreviations, spelling errors, incorrect encoding... Before adding it to any model, you need to have a standard preprocessing pipeline.
1. Overview of Text Preprocessing Pipeline
Raw Text
│
▼
┌────────────────────────┐
│ 1. Text Cleaning │ ← Loại HTML, URLs, special chars
├────────────────────────┤
│ 2. Unicode Normalization│ ← NFC/NFD, encoding fix
├────────────────────────┤
│ 3. Tokenization │ ← Tách thành tokens (từ/subword)
├────────────────────────┤
│ 4. Lowercasing │ ← Chuẩn hóa chữ hoa/thường
├────────────────────────┤
│ 5. Stopword Removal │ ← Loại từ không mang nghĩa
├────────────────────────┤
│ 6. Stemming/Lemma │ ← Đưa về dạng gốc
├────────────────────────┤
│ 7. Final Filtering │ ← Min length, frequency threshold
└────────────────────────┘
│
▼
Clean Tokens
⚠️ Note: All steps are not always needed. With BERT/GPT, you usually only need steps 1–3, because the model is already trained to handle the remaining steps.
2. Text Cleaning
2.1 Regex Power Tools
import re
def clean_text(text: str) -> str:
"""Pipeline làm sạch text cơ bản."""
# 1. Loại HTML tags
text = re.sub(r'<[^>]+>', '', text)
# 2. Loại URLs
text = re.sub(r'https?://\S+|www\.\S+', '[URL]', text)
# 3. Loại email
text = re.sub(r'\S+@\S+\.\S+', '[EMAIL]', text)
# 4. Loại số điện thoại
text = re.sub(r'\b\d{10,11}\b', '[PHONE]', text)
# 5. Chuẩn hóa whitespace
text = re.sub(r'\s+', ' ', text).strip()
return text
# Test
raw = """
<p>Liên hệ qua email: [email protected] hoặc
truy cập https://example.com để biết thêm chi tiết.
Hotline: 0901234567</p>
"""
print(clean_text(raw))
# "Liên hệ qua email: [EMAIL] hoặc truy cập [URL] để biết thêm chi tiết. Hotline: [PHONE]"
2.2 Handling Emoji and Special Characters
import emoji
def handle_emoji(text: str, mode: str = "remove") -> str:
"""Xử lý emoji: remove hoặc convert to text."""
if mode == "remove":
return emoji.replace_emoji(text, replace='')
elif mode == "text":
return emoji.demojize(text) # 😀 → :grinning_face:
return text
text = "Sản phẩm tuyệt vời! 😍🔥 5 sao ⭐⭐⭐⭐⭐"
print(handle_emoji(text, "remove"))
# "Sản phẩm tuyệt vời! 5 sao "
print(handle_emoji(text, "text"))
# "Sản phẩm tuyệt vời! :heart_eyes::fire: 5 sao :star::star::star::star::star:"
3. Unicode Normalization
import unicodedata
def normalize_unicode(text: str) -> str:
"""Chuẩn hóa Unicode: đặc biệt quan trọng cho tiếng Việt."""
# NFC: Composed form (khuyến nghị cho tiếng Việt)
text = unicodedata.normalize('NFC', text)
# Loại các ký tự control
text = ''.join(c for c in text if not unicodedata.category(c).startswith('C')
or c in '\n\t')
return text
# Ví dụ: 2 cách viết "ệ" trong Unicode
s1 = "Vi\u1ec7t" # ệ = single codepoint (NFC)
s2 = "Vie\u0302\u0323t" # ệ = e + circumflex + dot below (NFD)
print(s1 == s2) # False!
print(normalize_unicode(s1) == normalize_unicode(s2)) # True!
🇻🇳 Vietnamese: ALWAYS normalize to NFC before processing.
4. Tokenization
4.1 Word-level Tokenization
# Phương pháp đơn giản nhất: split by whitespace
text = "NLP là lĩnh vực thú vị"
tokens = text.split()
# ['NLP', 'là', 'lĩnh', 'vực', 'thú', 'vị']
# Với NLTK
import nltk
tokens = nltk.word_tokenize("I can't believe it's raining!")
# ["I", "ca", "n't", "believe", "it", "'s", "raining", "!"]
# Với spaCy
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("I can't believe it's raining!")
tokens = [token.text for token in doc]
# ["I", "ca", "n't", "believe", "it", "'s", "raining", "!"]
4.2 Sentence Tokenization
import nltk
nltk.download('punkt')
text = """NLP rất thú vị. Bạn nên học nó!
Đặc biệt là phần Transformer. Hãy bắt đầu ngay."""
sentences = nltk.sent_tokenize(text)
# ['NLP rất thú vị.', 'Bạn nên học nó!',
# 'Đặc biệt là phần Transformer.', 'Hãy bắt đầu ngay.']
5. Lowercasing, Stemming & Lemmatization
5.1 Lowercasing
text = "Natural Language PROCESSING is AMAZING"
text_lower = text.lower()
# "natural language processing is amazing"
⚠️ Be careful: lowercasing loses NER information (Apple company vs apple fruit).
5.2 Stemming vs Lemmatization
| Features | Stemming | Lemmatization |
|---|---|---|
| Method | Cut suffix (rule-based) | Look up dictionary + POS |
| Speed | Fast | Slower |
| Quality | Rough, possibly wrong | More accurate |
| Example | running → run, better → better | running → run, better → good |
# Stemming
from nltk.stem import PorterStemmer
stemmer = PorterStemmer()
words = ["running", "runs", "ran", "runner"]
print([stemmer.stem(w) for w in words])
# ['run', 'run', 'ran', 'runner']
# Lemmatization
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("He was running faster than the other runners")
print([(token.text, token.lemma_) for token in doc])
# [('He', 'he'), ('was', 'be'), ('running', 'run'),
# ('faster', 'fast'), ('than', 'than'), ('the', 'the'),
# ('other', 'other'), ('runners', 'runner')]
6. Stopword Removal
import spacy
nlp = spacy.load("en_core_web_sm")
text = "This is a very good example of text preprocessing in NLP"
doc = nlp(text)
# Loại stopwords
tokens = [token.text for token in doc if not token.is_stop and not token.is_punct]
print(tokens)
# ['good', 'example', 'text', 'preprocessing', 'NLP']
📌 With deep learning models (BERT, GPT), there is usually NO need to eliminate stopwords because the model uses context from all words.
7. Complete pipeline with spaCy
import spacy
import re
nlp = spacy.load("en_core_web_sm")
def preprocess_pipeline(text: str, remove_stopwords: bool = True) -> list[str]:
"""Pipeline preprocessing hoàn chỉnh."""
# 1. Clean
text = re.sub(r'<[^>]+>', '', text)
text = re.sub(r'https?://\S+', '', text)
text = re.sub(r'\s+', ' ', text).strip()
# 2. spaCy processing
doc = nlp(text.lower())
# 3. Filter tokens
tokens = []
for token in doc:
if token.is_punct or token.is_space:
continue
if remove_stopwords and token.is_stop:
continue
if len(token.lemma_) < 2:
continue
tokens.append(token.lemma_)
return tokens
text = """<p>NLP is an AMAZING field! Check https://example.com
for more info about Natural Language Processing.</p>"""
print(preprocess_pipeline(text))
# ['nlp', 'amazing', 'field', 'check', 'info', 'natural', 'language', 'processing']
Summary
| Step | When to use | When NOT |
|---|---|---|
| Text Cleaning | Always | — |
| Unicode Normalization | Vietnamese, multilingual | — |
| Tokenization | Always | — |
| Lowercasing | Classification, search | NER, when the case is important |
| Stopword Removal | Traditional ML, BoW/TF-IDF | Deep learning (BERT, GPT) |
| Stemming/Lemma | Search, IR, traditional ML | Deep learning |
Next article
Lesson 3: Tokenization Deep Dive — Dive into modern tokenization methods: BPE, WordPiece, SentencePiece — the foundation of all current LLMs.