Introduction
"You shall know a word by the company it keeps" — J.R. Firth, 1957
Word embeddings represent each word with a dense vector (usually 100-300 dimensions) so that words with similar meanings are close to each other in the vector space. This is one of the biggest breakthroughs in NLP history.
1. From One-hot to Dense Vectors
One-hot: Problem
# Vocab: ["king", "queen", "man", "woman", "apple"]
king = [1, 0, 0, 0, 0]
queen = [0, 1, 0, 0, 0]
man = [0, 0, 1, 0, 0]
# cosine_similarity(king, queen) = 0 ← Không biểu diễn quan hệ nghĩa!
# Mỗi từ cách đều nhau trong không gian
Dense Embeddings: Solution
# Word2Vec embeddings (ví dụ 4 chiều, thực tế 100-300)
king = [0.8, 0.6, -0.2, 0.9]
queen = [0.7, 0.7, 0.8, 0.8]
man = [0.9, 0.5, -0.3, 0.1]
woman = [0.8, 0.6, 0.7, 0.1]
# king - man + woman ≈ queen ← "Phép toán" trên nghĩa từ!
2. Word2Vec
2.1 CBOW vs Skip-gram
CBOW (Continuous Bag of Words):
Context → Target word
["the", "cat", "on", "the"] → "sat"
Skip-gram:
Target word → Context
"sat" → ["the", "cat", "on", "the"]
| Features | CBOW | Skip-gram |
|---|---|---|
| Input | Context words | Target word |
| Output | Target word | Context words |
| Speed | Faster | Slower |
| Rare words | Inferior | Better |
| Small dataset | Better | — |
2.2 Hands-on with Gensim
from gensim.models import Word2Vec
# Corpus (list of tokenized sentences)
sentences = [
["xử", "lý", "ngôn", "ngữ", "tự", "nhiên"],
["machine", "learning", "và", "deep", "learning"],
["trí", "tuệ", "nhân", "tạo", "phát", "triển"],
# ... thêm data
]
# Train Word2Vec
model = Word2Vec(
sentences,
vector_size=100, # Số chiều embedding
window=5, # Context window
min_count=1, # Bỏ qua từ xuất hiện < min_count
sg=1, # 0=CBOW, 1=Skip-gram
epochs=10,
)
# Tìm từ tương tự
print(model.wv.most_similar("ngôn", topn=5))
# Phép toán vector
result = model.wv.most_similar(
positive=["queen", "man"],
negative=["woman"],
topn=1
)
print(result) # [('king', 0.85)]
# Lưu và load
model.save("word2vec_vi.model")
loaded = Word2Vec.load("word2vec_vi.model")
3. GloVe (Global Vectors)
Ideas
GloVe builds a co-occurrence matrix from the entire corpus, then factorizes:
$$J = \sum_{i,j=1}^{V} f(X_{ij})(w_i^T \tilde{w}_j + b_i + \tilde{b}j - \log X{ij})^2$$
# Sử dụng pre-trained GloVe
import gensim.downloader as api
glove = api.load("glove-wiki-gigaword-100") # 100d vectors
# Tìm từ tương tự
print(glove.most_similar("computer", topn=5))
# [('computers', 0.87), ('software', 0.81), ('technology', 0.78), ...]
# Analogy: king - man + woman = ?
print(glove.most_similar(positive=["king", "woman"], negative=["man"], topn=1))
# [('queen', 0.77)]
4. FastText
Advantage: Subword Embeddings
FastText represents words as sum of n-grams of characters — handles OOV words!
from gensim.models import FastText
model = FastText(
sentences,
vector_size=100,
window=5,
min_count=1,
sg=1, # Skip-gram
)
# Có thể lấy vector cho từ CHƯA THẤY BÃO GIỜ
vector = model.wv["từmớichưabaogiờthấy"] # Vẫn work!
# Word2Vec sẽ báo KeyError
5. Compare Word2Vec vs GloVe vs FastText
| Features | Word2Vec | GloVe | FastText |
|---|---|---|---|
| Method | Prediction (local context) | Count (global statistics) | Prediction + subword |
| OOV handling | No | No | Yes (subword n-grams) |
| Morphology | No | No | Yes |
| Training speed | Fast | Fast | Slower |
| Quality | Good | Good | Best for morphologically-rich languages |
6. Visualize Embeddings
from sklearn.manifold import TSNE
import matplotlib.pyplot as plt
import numpy as np
words = ["king", "queen", "man", "woman", "prince", "princess",
"dog", "cat", "fish", "bird",
"python", "java", "code", "programming"]
vectors = np.array([glove[w] for w in words])
# t-SNE giảm chiều xuống 2D
tsne = TSNE(n_components=2, random_state=42, perplexity=5)
vectors_2d = tsne.fit_transform(vectors)
plt.figure(figsize=(12, 8))
for i, word in enumerate(words):
plt.scatter(vectors_2d[i, 0], vectors_2d[i, 1])
plt.annotate(word, xy=(vectors_2d[i, 0], vectors_2d[i, 1]),
fontsize=12, ha='center', va='bottom')
plt.title("Word Embeddings Visualization (t-SNE)")
plt.show()
7. Pre-trained Embeddings for Vietnamese
| Resources | Dimensions | Vocab | Link |
|---|---|---|---|
| PhoBERT embeddings | 768 | 64K | vinai/phobert-base-v2 |
| fastText Vietnamese | 300 | 2M | cc.vi.300.bin |
| PhoW2V | 100/300 | 500K | github.com/datquocnguyen/PhoW2V |
import fasttext
import fasttext.util
# Download pre-trained Vietnamese FastText
fasttext.util.download_model('vi', if_exists='ignore')
ft = fasttext.load_model('cc.vi.300.bin')
# Sử dụng
vector = ft.get_word_vector("trí_tuệ_nhân_tạo")
similar = ft.get_nearest_neighbors("lập_trình", k=5)
print(similar)
Summary
| Concept | Key points |
|---|---|
| One-hot | Sparse, does not express meaning |
| Word2Vec | Dense vectors, CBOW/Skip-gram, semantic operations |
| GloVe | Global co-occurrence + factorization |
| FastText | Subword n-grams, OOV processing |
| Vietnamese | PhoW2V, fastText Vietnamese, PhoBERT |
Next article
Lesson 6: Sentence & Document Embeddings — Expanding from word-level to sentence-level: Sentence-BERT, E5, and semantic search applications.