Chuyển đến nội dung chính

Lesson 5: Word Embeddings — Word2Vec, GloVe, FastText

From one-hot to dense vectors. Word2Vec: CBOW vs Skip-gram, negative sampling. GloVe: co-occurrence matrix factorization. FastText: subword embeddings. Visualize with t-SNE/UMAP. Pre-trained embeddings for Vietnamese. Hands-on with Gensim.

🧠 AI & ML — Lesson 4 Lesson 5: Word Embeddings — Word2Vec, GloVe, FastText

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 2: Language Representation — From BoW to Word Embeddings

xdev.asia

Introduction

"You shall know a word by the company it keeps" — J.R. Firth, 1957

Word embeddings represent each word with a dense vector (usually 100-300 dimensions) so that words with similar meanings are close to each other in the vector space. This is one of the biggest breakthroughs in NLP history.


1. From One-hot to Dense Vectors

One-hot: Problem

# Vocab: ["king", "queen", "man", "woman", "apple"]
king  = [1, 0, 0, 0, 0]
queen = [0, 1, 0, 0, 0]
man   = [0, 0, 1, 0, 0]

# cosine_similarity(king, queen) = 0  ← Không biểu diễn quan hệ nghĩa!
# Mỗi từ cách đều nhau trong không gian

Dense Embeddings: Solution

# Word2Vec embeddings (ví dụ 4 chiều, thực tế 100-300)
king  = [0.8, 0.6, -0.2, 0.9]
queen = [0.7, 0.7, 0.8, 0.8]
man   = [0.9, 0.5, -0.3, 0.1]
woman = [0.8, 0.6, 0.7, 0.1]

# king - man + woman ≈ queen  ← "Phép toán" trên nghĩa từ!

2. Word2Vec

2.1 CBOW vs Skip-gram

CBOW (Continuous Bag of Words):
  Context → Target word
  ["the", "cat", "on", "the"] → "sat"

Skip-gram:
  Target word → Context
  "sat" → ["the", "cat", "on", "the"]
FeaturesCBOWSkip-gram
InputContext wordsTarget word
OutputTarget wordContext words
Speed ​​FasterSlower
Rare wordsInferiorBetter
Small datasetBetter—

2.2 Hands-on with Gensim

from gensim.models import Word2Vec

# Corpus (list of tokenized sentences)
sentences = [
    ["xử", "lý", "ngôn", "ngữ", "tự", "nhiên"],
    ["machine", "learning", "và", "deep", "learning"],
    ["trí", "tuệ", "nhân", "tạo", "phát", "triển"],
    # ... thêm data
]

# Train Word2Vec
model = Word2Vec(
    sentences,
    vector_size=100,   # Số chiều embedding
    window=5,          # Context window
    min_count=1,       # Bỏ qua từ xuất hiện < min_count
    sg=1,              # 0=CBOW, 1=Skip-gram
    epochs=10,
)

# Tìm từ tương tự
print(model.wv.most_similar("ngôn", topn=5))

# Phép toán vector
result = model.wv.most_similar(
    positive=["queen", "man"],
    negative=["woman"],
    topn=1
)
print(result)  # [('king', 0.85)]

# Lưu và load
model.save("word2vec_vi.model")
loaded = Word2Vec.load("word2vec_vi.model")

3. GloVe (Global Vectors)

Ideas

GloVe builds a co-occurrence matrix from the entire corpus, then factorizes:

$$J = \sum_{i,j=1}^{V} f(X_{ij})(w_i^T \tilde{w}_j + b_i + \tilde{b}j - \log X{ij})^2$$

# Sử dụng pre-trained GloVe
import gensim.downloader as api

glove = api.load("glove-wiki-gigaword-100")  # 100d vectors

# Tìm từ tương tự
print(glove.most_similar("computer", topn=5))
# [('computers', 0.87), ('software', 0.81), ('technology', 0.78), ...]

# Analogy: king - man + woman = ?
print(glove.most_similar(positive=["king", "woman"], negative=["man"], topn=1))
# [('queen', 0.77)]

4. FastText

Advantage: Subword Embeddings

FastText represents words as sum of n-grams of characters — handles OOV words!

from gensim.models import FastText

model = FastText(
    sentences,
    vector_size=100,
    window=5,
    min_count=1,
    sg=1,  # Skip-gram
)

# Có thể lấy vector cho từ CHƯA THẤY BÃO GIỜ
vector = model.wv["từmớichưabaogiờthấy"]  # Vẫn work!
# Word2Vec sẽ báo KeyError

5. Compare Word2Vec vs GloVe vs FastText

FeaturesWord2VecGloVeFastText
MethodPrediction (local context)Count (global statistics)Prediction + subword
OOV handlingNoNoYes (subword n-grams)
MorphologyNoNoYes
Training speedFastFastSlower
QualityGoodGoodBest for morphologically-rich languages ​​

6. Visualize Embeddings

from sklearn.manifold import TSNE
import matplotlib.pyplot as plt
import numpy as np

words = ["king", "queen", "man", "woman", "prince", "princess",
         "dog", "cat", "fish", "bird",
         "python", "java", "code", "programming"]

vectors = np.array([glove[w] for w in words])

# t-SNE giảm chiều xuống 2D
tsne = TSNE(n_components=2, random_state=42, perplexity=5)
vectors_2d = tsne.fit_transform(vectors)

plt.figure(figsize=(12, 8))
for i, word in enumerate(words):
    plt.scatter(vectors_2d[i, 0], vectors_2d[i, 1])
    plt.annotate(word, xy=(vectors_2d[i, 0], vectors_2d[i, 1]),
                 fontsize=12, ha='center', va='bottom')
plt.title("Word Embeddings Visualization (t-SNE)")
plt.show()

7. Pre-trained Embeddings for Vietnamese

ResourcesDimensionsVocabLink
PhoBERT embeddings76864Kvinai/phobert-base-v2
fastText Vietnamese3002Mcc.vi.300.bin
PhoW2V100/300500Kgithub.com/datquocnguyen/PhoW2V
import fasttext
import fasttext.util

# Download pre-trained Vietnamese FastText
fasttext.util.download_model('vi', if_exists='ignore')
ft = fasttext.load_model('cc.vi.300.bin')

# Sử dụng
vector = ft.get_word_vector("trí_tuệ_nhân_tạo")
similar = ft.get_nearest_neighbors("lập_trình", k=5)
print(similar)

Summary

ConceptKey points
One-hotSparse, does not express meaning
Word2VecDense vectors, CBOW/Skip-gram, semantic operations
GloVeGlobal co-occurrence + factorization
FastTextSubword n-grams, OOV processing
VietnamesePhoW2V, fastText Vietnamese, PhoBERT

Next article

Lesson 6: Sentence & Document Embeddings — Expanding from word-level to sentence-level: Sentence-BERT, E5, and semantic search applications.