簡介
「垃圾進,垃圾出」-在 NLP 中,80% 的成功來自於準備和清理文字資料。
「現實生活」中的文字資料總是髒:包含HTML標籤、表情符號、特殊字元、縮寫、拼字錯誤、不正確的編碼...在將其添加到任何模型之前,您需要有一個標準的預處理管道。
1. 文字預處理流程概述
Raw Text
│
▼
┌────────────────────────┐
│ 1. Text Cleaning │ ← Loại HTML, URLs, special chars
├────────────────────────┤
│ 2. Unicode Normalization│ ← NFC/NFD, encoding fix
├────────────────────────┤
│ 3. Tokenization │ ← Tách thành tokens (từ/subword)
├────────────────────────┤
│ 4. Lowercasing │ ← Chuẩn hóa chữ hoa/thường
├────────────────────────┤
│ 5. Stopword Removal │ ← Loại từ không mang nghĩa
├────────────────────────┤
│ 6. Stemming/Lemma │ ← Đưa về dạng gốc
├────────────────────────┤
│ 7. Final Filtering │ ← Min length, frequency threshold
└────────────────────────┘
│
▼
Clean Tokens
⚠️ 注意: 並不總是需要所有步驟。使用 BERT/GPT,您通常只需要步驟 1-3,因為模型已經經過訓練可以處理其餘步驟。
2. 文字清理
2.1 正規表示式電動工具
import re
def clean_text(text: str) -> str:
"""Pipeline làm sạch text cơ bản."""
# 1. Loại HTML tags
text = re.sub(r'<[^>]+>', '', text)
# 2. Loại URLs
text = re.sub(r'https?://\S+|www\.\S+', '[URL]', text)
# 3. Loại email
text = re.sub(r'\S+@\S+\.\S+', '[EMAIL]', text)
# 4. Loại số điện thoại
text = re.sub(r'\b\d{10,11}\b', '[PHONE]', text)
# 5. Chuẩn hóa whitespace
text = re.sub(r'\s+', ' ', text).strip()
return text
# Test
raw = """
<p>Liên hệ qua email: [email protected] hoặc
truy cập https://example.com để biết thêm chi tiết.
Hotline: 0901234567</p>
"""
print(clean_text(raw))
# "Liên hệ qua email: [EMAIL] hoặc truy cập [URL] để biết thêm chi tiết. Hotline: [PHONE]"
2.2 處理表情符號和特殊字符
import emoji
def handle_emoji(text: str, mode: str = "remove") -> str:
"""Xử lý emoji: remove hoặc convert to text."""
if mode == "remove":
return emoji.replace_emoji(text, replace='')
elif mode == "text":
return emoji.demojize(text) # 😀 → :grinning_face:
return text
text = "Sản phẩm tuyệt vời! 😍🔥 5 sao ⭐⭐⭐⭐⭐"
print(handle_emoji(text, "remove"))
# "Sản phẩm tuyệt vời! 5 sao "
print(handle_emoji(text, "text"))
# "Sản phẩm tuyệt vời! :heart_eyes::fire: 5 sao :star::star::star::star::star:"
3. Unicode 規範化
import unicodedata
def normalize_unicode(text: str) -> str:
"""Chuẩn hóa Unicode: đặc biệt quan trọng cho tiếng Việt."""
# NFC: Composed form (khuyến nghị cho tiếng Việt)
text = unicodedata.normalize('NFC', text)
# Loại các ký tự control
text = ''.join(c for c in text if not unicodedata.category(c).startswith('C')
or c in '\n\t')
return text
# Ví dụ: 2 cách viết "ệ" trong Unicode
s1 = "Vi\u1ec7t" # ệ = single codepoint (NFC)
s2 = "Vie\u0302\u0323t" # ệ = e + circumflex + dot below (NFD)
print(s1 == s2) # False!
print(normalize_unicode(s1) == normalize_unicode(s2)) # True!
🇻🇳 越南語:在處理之前始終標準化為 NFC。
4. 代幣化
4.1 字級標記化
# Phương pháp đơn giản nhất: split by whitespace
text = "NLP là lĩnh vực thú vị"
tokens = text.split()
# ['NLP', 'là', 'lĩnh', 'vực', 'thú', 'vị']
# Với NLTK
import nltk
tokens = nltk.word_tokenize("I can't believe it's raining!")
# ["I", "ca", "n't", "believe", "it", "'s", "raining", "!"]
# Với spaCy
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("I can't believe it's raining!")
tokens = [token.text for token in doc]
# ["I", "ca", "n't", "believe", "it", "'s", "raining", "!"]
4.2 句子標記化
import nltk
nltk.download('punkt')
text = """NLP rất thú vị. Bạn nên học nó!
Đặc biệt là phần Transformer. Hãy bắt đầu ngay."""
sentences = nltk.sent_tokenize(text)
# ['NLP rất thú vị.', 'Bạn nên học nó!',
# 'Đặc biệt là phần Transformer.', 'Hãy bắt đầu ngay.']
5. 小寫、字幹擷取與詞形還原
5.1 小寫
text = "Natural Language PROCESSING is AMAZING"
text_lower = text.lower()
# "natural language processing is amazing"
⚠️ 注意:小寫會失去 NER 訊息(Apple 公司 vs 蘋果水果)。
5.2 詞幹擷取與詞形還原
| 特點 | 詞幹 | 詞形還原 |
|---|---|---|
| 方法 | 剪下後綴(基於規則) | 查字典+POS |
| 速度 | 快 | 慢一點 |
| 品質 | 粗糙,可能是錯誤的 | 更準確 |
| 範例 | 跑步→跑步,更好→更好 | 跑→跑步,更好→好 |
# Stemming
from nltk.stem import PorterStemmer
stemmer = PorterStemmer()
words = ["running", "runs", "ran", "runner"]
print([stemmer.stem(w) for w in words])
# ['run', 'run', 'ran', 'runner']
# Lemmatization
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("He was running faster than the other runners")
print([(token.text, token.lemma_) for token in doc])
# [('He', 'he'), ('was', 'be'), ('running', 'run'),
# ('faster', 'fast'), ('than', 'than'), ('the', 'the'),
# ('other', 'other'), ('runners', 'runner')]
6. 停用詞刪除
import spacy
nlp = spacy.load("en_core_web_sm")
text = "This is a very good example of text preprocessing in NLP"
doc = nlp(text)
# Loại stopwords
tokens = [token.text for token in doc if not token.is_stop and not token.is_punct]
print(tokens)
# ['good', 'example', 'text', 'preprocessing', 'NLP']
📌 使用深度學習模型(BERT、GPT),通常不需要需要消除停用詞,因為該模型使用所有單字的上下文。
7. 使用 spaCy 完成管道
import spacy
import re
nlp = spacy.load("en_core_web_sm")
def preprocess_pipeline(text: str, remove_stopwords: bool = True) -> list[str]:
"""Pipeline preprocessing hoàn chỉnh."""
# 1. Clean
text = re.sub(r'<[^>]+>', '', text)
text = re.sub(r'https?://\S+', '', text)
text = re.sub(r'\s+', ' ', text).strip()
# 2. spaCy processing
doc = nlp(text.lower())
# 3. Filter tokens
tokens = []
for token in doc:
if token.is_punct or token.is_space:
continue
if remove_stopwords and token.is_stop:
continue
if len(token.lemma_) < 2:
continue
tokens.append(token.lemma_)
return tokens
text = """<p>NLP is an AMAZING field! Check https://example.com
for more info about Natural Language Processing.</p>"""
print(preprocess_pipeline(text))
# ['nlp', 'amazing', 'field', 'check', 'info', 'natural', 'language', 'processing']
總結
| 步驟 | 何時使用 | 當不 |
|---|---|---|
| 文字清理 | 永遠 | — |
| Unicode 規範化 | 越南語,多元語言 | — |
| 代幣化 | 永遠 | — |
| 小寫 | 分類、搜尋 | NER,當案件很重要時 |
| 停用詞刪除 | 傳統機器學習、BoW/TF-IDF | 深度學習(BERT、GPT) |
| 詞幹擷取/引理 | 搜尋、IR、傳統機器學習 | 深度學習 |
下一篇文章
第 3 課:深入剖析標記化 — 深入研究現代標記化方法:BPE、WordPiece、SentencePiece — 所有當前法學碩士的基礎。