Chuyển đến nội dung chính

Bài 5: Word Embeddings — Word2Vec, GloVe, FastText

Từ one-hot đến dense vectors. Word2Vec: CBOW vs Skip-gram, negative sampling. GloVe: co-occurrence matrix factorization. FastText: subword embeddings. Visualize với t-SNE/UMAP. Pre-trained embeddings cho tiếng Việt. Hands-on với Gensim.

🧠 AI & ML — Bài 4 Bài 5: Word Embeddings — Word2Vec, GloVe, FastText

NLP từ Cơ bản đến Nâng cao: Làm chủ Xử lý Ngôn ngữ Tự nhiên

Phần 2: Biểu diễn Ngôn ngữ — Từ BoW đến Word Embeddings

xdev.asia

Giới thiệu

"You shall know a word by the company it keeps" — J.R. Firth, 1957

Word embeddings biểu diễn mỗi từ bằng một dense vector (thường 100-300 chiều) sao cho các từ có nghĩa tương tự nằm gần nhau trong không gian vector. Đây là một trong những breakthrough lớn nhất trong lịch sử NLP.


1. Từ One-hot đến Dense Vectors

One-hot: Vấn đề

# Vocab: ["king", "queen", "man", "woman", "apple"]
king  = [1, 0, 0, 0, 0]
queen = [0, 1, 0, 0, 0]
man   = [0, 0, 1, 0, 0]

# cosine_similarity(king, queen) = 0  ← Không biểu diễn quan hệ nghĩa!
# Mỗi từ cách đều nhau trong không gian

Dense Embeddings: Giải pháp

# Word2Vec embeddings (ví dụ 4 chiều, thực tế 100-300)
king  = [0.8, 0.6, -0.2, 0.9]
queen = [0.7, 0.7, 0.8, 0.8]
man   = [0.9, 0.5, -0.3, 0.1]
woman = [0.8, 0.6, 0.7, 0.1]

# king - man + woman ≈ queen  ← "Phép toán" trên nghĩa từ!

2. Word2Vec

2.1 CBOW vs Skip-gram

CBOW (Continuous Bag of Words):
  Context → Target word
  ["the", "cat", "on", "the"] → "sat"

Skip-gram:
  Target word → Context
  "sat" → ["the", "cat", "on", "the"]
Đặc điểmCBOWSkip-gram
InputContext wordsTarget word
OutputTarget wordContext words
Tốc độNhanh hơnChậm hơn
Từ hiếmKém hơnTốt hơn
Dataset nhỏTốt hơn—

2.2 Hands-on với Gensim

from gensim.models import Word2Vec

# Corpus (list of tokenized sentences)
sentences = [
    ["xử", "lý", "ngôn", "ngữ", "tự", "nhiên"],
    ["machine", "learning", "và", "deep", "learning"],
    ["trí", "tuệ", "nhân", "tạo", "phát", "triển"],
    # ... thêm data
]

# Train Word2Vec
model = Word2Vec(
    sentences,
    vector_size=100,   # Số chiều embedding
    window=5,          # Context window
    min_count=1,       # Bỏ qua từ xuất hiện < min_count
    sg=1,              # 0=CBOW, 1=Skip-gram
    epochs=10,
)

# Tìm từ tương tự
print(model.wv.most_similar("ngôn", topn=5))

# Phép toán vector
result = model.wv.most_similar(
    positive=["queen", "man"],
    negative=["woman"],
    topn=1
)
print(result)  # [('king', 0.85)]

# Lưu và load
model.save("word2vec_vi.model")
loaded = Word2Vec.load("word2vec_vi.model")

3. GloVe (Global Vectors)

Ý tưởng

GloVe xây dựng co-occurrence matrix từ toàn bộ corpus, rồi factorize:

$$J = \sum_{i,j=1}^{V} f(X_{ij})(w_i^T \tilde{w}_j + b_i + \tilde{b}j - \log X{ij})^2$$

# Sử dụng pre-trained GloVe
import gensim.downloader as api

glove = api.load("glove-wiki-gigaword-100")  # 100d vectors

# Tìm từ tương tự
print(glove.most_similar("computer", topn=5))
# [('computers', 0.87), ('software', 0.81), ('technology', 0.78), ...]

# Analogy: king - man + woman = ?
print(glove.most_similar(positive=["king", "woman"], negative=["man"], topn=1))
# [('queen', 0.77)]

4. FastText

Ưu điểm: Subword Embeddings

FastText biểu diễn từ bằng tổng các n-gram ký tự — xử lý được từ OOV!

from gensim.models import FastText

model = FastText(
    sentences,
    vector_size=100,
    window=5,
    min_count=1,
    sg=1,  # Skip-gram
)

# Có thể lấy vector cho từ CHƯA THẤY BÃO GIỜ
vector = model.wv["từmớichưabaogiờthấy"]  # Vẫn work!
# Word2Vec sẽ báo KeyError

5. So sánh Word2Vec vs GloVe vs FastText

Đặc điểmWord2VecGloVeFastText
Phương phápPrediction (local context)Count (global statistics)Prediction + subword
OOV handlingKhôngKhôngCó (subword n-grams)
MorphologyKhôngKhôngCó
Training speedNhanhNhanhChậm hơn
Chất lượngTốtTốtTốt nhất cho morphologically-rich languages

6. Visualize Embeddings

from sklearn.manifold import TSNE
import matplotlib.pyplot as plt
import numpy as np

words = ["king", "queen", "man", "woman", "prince", "princess",
         "dog", "cat", "fish", "bird",
         "python", "java", "code", "programming"]

vectors = np.array([glove[w] for w in words])

# t-SNE giảm chiều xuống 2D
tsne = TSNE(n_components=2, random_state=42, perplexity=5)
vectors_2d = tsne.fit_transform(vectors)

plt.figure(figsize=(12, 8))
for i, word in enumerate(words):
    plt.scatter(vectors_2d[i, 0], vectors_2d[i, 1])
    plt.annotate(word, xy=(vectors_2d[i, 0], vectors_2d[i, 1]),
                 fontsize=12, ha='center', va='bottom')
plt.title("Word Embeddings Visualization (t-SNE)")
plt.show()

7. Pre-trained Embeddings cho Tiếng Việt

ResourceDimensionsVocabLink
PhoBERT embeddings76864Kvinai/phobert-base-v2
fastText Vietnamese3002Mcc.vi.300.bin
PhoW2V100/300500Kgithub.com/datquocnguyen/PhoW2V
import fasttext
import fasttext.util

# Download pre-trained Vietnamese FastText
fasttext.util.download_model('vi', if_exists='ignore')
ft = fasttext.load_model('cc.vi.300.bin')

# Sử dụng
vector = ft.get_word_vector("trí_tuệ_nhân_tạo")
similar = ft.get_nearest_neighbors("lập_trình", k=5)
print(similar)

Tổng kết

Khái niệmĐiểm chính
One-hotSparse, không biểu diễn nghĩa
Word2VecDense vectors, CBOW/Skip-gram, phép toán nghĩa
GloVeGlobal co-occurrence + factorization
FastTextSubword n-grams, xử lý OOV
Tiếng ViệtPhoW2V, fastText Vietnamese, PhoBERT

Bài tiếp theo

Bài 6: Sentence & Document Embeddings — Mở rộng từ word-level lên sentence-level: Sentence-BERT, E5, và ứng dụng semantic search.