Chuyển đến nội dung chính

レッスン 2: テキストの前処理 — テキストのクリーニングと標準化

トークン化 (単語、サブワード、文字レベル)。小文字化、ステミング、見出し語化。ストップワードの削除。テキストクリーニング用の正規表現。 Unicode とエンコーディングの問題。 Python と spaCy を使用した実践的なパイプライン前処理。

🧠 AI と ML — レッスン 1 レッスン 2: テキストの前処理 — クリーニングと ドキュメントの標準化

NLP の基礎から上級まで: 自然言語処理をマスターする

パート 1: NLP の基礎 — コンピューターのレンズを通して言語を理解する

xdev.asia

はじめに

「ガベージイン、ガベージアウト」 — NLP では、成功の 80% はテキスト データの準備とクリーニングから得られます。

「現実世界」のテキスト データは常に ダーティです。HTML タグ、絵文字、特殊文字、略語、スペル ミス、不正なエンコードが含まれています...モデルに追加する前に、標準の 前処理パイプラインが必要です。


1. テキスト前処理パイプラインの概要

Raw Text
    │
    ▼
┌────────────────────────┐
│ 1. Text Cleaning       │ ← Loại HTML, URLs, special chars
├────────────────────────┤
│ 2. Unicode Normalization│ ← NFC/NFD, encoding fix
├────────────────────────┤
│ 3. Tokenization        │ ← Tách thành tokens (từ/subword)
├────────────────────────┤
│ 4. Lowercasing         │ ← Chuẩn hóa chữ hoa/thường
├────────────────────────┤
│ 5. Stopword Removal    │ ← Loại từ không mang nghĩa
├────────────────────────┤
│ 6. Stemming/Lemma      │ ← Đưa về dạng gốc
├────────────────────────┤
│ 7. Final Filtering     │ ← Min length, frequency threshold
└────────────────────────┘
    │
    ▼
Clean Tokens

⚠️ 注: すべての手順が常に必要なわけではありません。 BERT/GPT を使用する場合、モデルは残りのステップを処理するようにすでにトレーニングされているため、通常はステップ 1 ~ 3 のみが必要です。


2. テキストのクリーニング

2.1 正規表現パワーツール

import re

def clean_text(text: str) -> str:
    """Pipeline làm sạch text cơ bản."""
    # 1. Loại HTML tags
    text = re.sub(r'<[^>]+>', '', text)

    # 2. Loại URLs
    text = re.sub(r'https?://\S+|www\.\S+', '[URL]', text)

    # 3. Loại email
    text = re.sub(r'\S+@\S+\.\S+', '[EMAIL]', text)

    # 4. Loại số điện thoại
    text = re.sub(r'\b\d{10,11}\b', '[PHONE]', text)

    # 5. Chuẩn hóa whitespace
    text = re.sub(r'\s+', ' ', text).strip()

    return text

# Test
raw = """
<p>Liên hệ qua email: [email protected] hoặc
truy cập https://example.com để biết thêm chi tiết.
Hotline: 0901234567</p>
"""
print(clean_text(raw))
# "Liên hệ qua email: [EMAIL] hoặc truy cập [URL] để biết thêm chi tiết. Hotline: [PHONE]"

2.2 絵文字と特殊文字の処理

import emoji

def handle_emoji(text: str, mode: str = "remove") -> str:
    """Xử lý emoji: remove hoặc convert to text."""
    if mode == "remove":
        return emoji.replace_emoji(text, replace='')
    elif mode == "text":
        return emoji.demojize(text)  # 😀 → :grinning_face:
    return text

text = "Sản phẩm tuyệt vời! 😍🔥 5 sao ⭐⭐⭐⭐⭐"
print(handle_emoji(text, "remove"))
# "Sản phẩm tuyệt vời!  5 sao "
print(handle_emoji(text, "text"))
# "Sản phẩm tuyệt vời! :heart_eyes::fire: 5 sao :star::star::star::star::star:"

3. Unicode の正規化

import unicodedata

def normalize_unicode(text: str) -> str:
    """Chuẩn hóa Unicode: đặc biệt quan trọng cho tiếng Việt."""
    # NFC: Composed form (khuyến nghị cho tiếng Việt)
    text = unicodedata.normalize('NFC', text)

    # Loại các ký tự control
    text = ''.join(c for c in text if not unicodedata.category(c).startswith('C')
                   or c in '\n\t')

    return text

# Ví dụ: 2 cách viết "ệ" trong Unicode
s1 = "Vi\u1ec7t"          # ệ = single codepoint (NFC)
s2 = "Vie\u0302\u0323t"   # ệ = e + circumflex + dot below (NFD)
print(s1 == s2)                                    # False!
print(normalize_unicode(s1) == normalize_unicode(s2))  # True!

🇻🇳 ベトナム語: 処理する前に必ず NFC に正規化してください。


4. トークン化

4.1 単語レベルのトークン化

# Phương pháp đơn giản nhất: split by whitespace
text = "NLP là lĩnh vực thú vị"
tokens = text.split()
# ['NLP', 'là', 'lĩnh', 'vực', 'thú', 'vị']

# Với NLTK
import nltk
tokens = nltk.word_tokenize("I can't believe it's raining!")
# ["I", "ca", "n't", "believe", "it", "'s", "raining", "!"]

# Với spaCy
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("I can't believe it's raining!")
tokens = [token.text for token in doc]
# ["I", "ca", "n't", "believe", "it", "'s", "raining", "!"]

4.2 文のトークン化

import nltk
nltk.download('punkt')

text = """NLP rất thú vị. Bạn nên học nó!
Đặc biệt là phần Transformer. Hãy bắt đầu ngay."""

sentences = nltk.sent_tokenize(text)
# ['NLP rất thú vị.', 'Bạn nên học nó!',
#  'Đặc biệt là phần Transformer.', 'Hãy bắt đầu ngay.']

5. 小文字化、語幹処理、および見出し語化

5.1 小文字化

text = "Natural Language PROCESSING is AMAZING"
text_lower = text.lower()
# "natural language processing is amazing"

⚠️ 注意してください: 小文字にすると NER 情報 (Apple 社と Apple 果実) が失われます。

5.2 ステミングと見出し語化

特長ステミング見出し語化
方法カットサフィックス (ルールベース)辞書を引く + POS
スピード速い遅い
品質大まかで、間違っている可能性がありますより正確に
例走る → 走る、良くなる → 良くなる走る → 走る、良くなる → 良い
# Stemming
from nltk.stem import PorterStemmer
stemmer = PorterStemmer()
words = ["running", "runs", "ran", "runner"]
print([stemmer.stem(w) for w in words])
# ['run', 'run', 'ran', 'runner']

# Lemmatization
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("He was running faster than the other runners")
print([(token.text, token.lemma_) for token in doc])
# [('He', 'he'), ('was', 'be'), ('running', 'run'),
#  ('faster', 'fast'), ('than', 'than'), ('the', 'the'),
#  ('other', 'other'), ('runners', 'runner')]

6. ストップワードの削除

import spacy

nlp = spacy.load("en_core_web_sm")

text = "This is a very good example of text preprocessing in NLP"
doc = nlp(text)

# Loại stopwords
tokens = [token.text for token in doc if not token.is_stop and not token.is_punct]
print(tokens)
# ['good', 'example', 'text', 'preprocessing', 'NLP']

📌 深層学習モデル (BERT、GPT) では、モデルがすべての単語のコンテキストを使用するため、通常、ストップワードを削除する必要はありません**。


7. spaCy によるパイプラインの完成

import spacy
import re

nlp = spacy.load("en_core_web_sm")

def preprocess_pipeline(text: str, remove_stopwords: bool = True) -> list[str]:
    """Pipeline preprocessing hoàn chỉnh."""
    # 1. Clean
    text = re.sub(r'<[^>]+>', '', text)
    text = re.sub(r'https?://\S+', '', text)
    text = re.sub(r'\s+', ' ', text).strip()

    # 2. spaCy processing
    doc = nlp(text.lower())

    # 3. Filter tokens
    tokens = []
    for token in doc:
        if token.is_punct or token.is_space:
            continue
        if remove_stopwords and token.is_stop:
            continue
        if len(token.lemma_) < 2:
            continue
        tokens.append(token.lemma_)

    return tokens

text = """<p>NLP is an AMAZING field! Check https://example.com
for more info about Natural Language Processing.</p>"""

print(preprocess_pipeline(text))
# ['nlp', 'amazing', 'field', 'check', 'info', 'natural', 'language', 'processing']

概要

ステップいつ使用するかそうでない場合
テキストのクリーニング常に—
Unicode の正規化ベトナム語、多言語—
トークン化常に—
小文字分類・検索NER、事件が重要な場合
ストップワードの削除従来の ML、BoW/TF-IDFディープラーニング (BERT、GPT)
ステミング/補題検索、IR、従来の MLディープラーニング

次の記事

レッスン 3: トークン化の詳細 — 最新のトークン化手法について詳しく説明します: BPE、WordPiece、SentencePiece — 現在のすべての LLM の基礎です。