Introduction
Word embeddings represent each word — but in reality, we need to compare sentences or documents with each other. Sentence embeddings turn each sentence into a unique vector, allowing for the calculation of semantic similarity — the foundation of semantic search, RAG, and many modern NLP applications.
1. From Word to Sentence Embeddings
Simple method: Average Pooling
import numpy as np
def average_embedding(words, model):
"""Trung bình vector của các từ trong câu."""
vectors = [model[w] for w in words if w in model]
if not vectors:
return np.zeros(model.vector_size)
return np.mean(vectors, axis=0)
sentence = ["xử", "lý", "ngôn", "ngữ", "tự", "nhiên"]
sent_vec = average_embedding(sentence, word2vec_model)
# Shape: (100,) — một vector duy nhất cho cả câu
⚠️ Average pooling loses order information and context — "not good" will be closer to "good"!
2. Sentence-BERT (SBERT)
Ideas
Fine-tune BERT with Siamese network to create high quality sentence embeddings.
from sentence_transformers import SentenceTransformer, util
# Load pre-trained model
model = SentenceTransformer('all-MiniLM-L6-v2')
# Encode sentences
sentences = [
"NLP là lĩnh vực xử lý ngôn ngữ tự nhiên",
"Xử lý ngôn ngữ tự nhiên thuộc về AI",
"Hôm nay thời tiết đẹp quá",
]
embeddings = model.encode(sentences)
print(embeddings.shape) # (3, 384)
# Tính cosine similarity
similarities = util.cos_sim(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.8234, 0.0512], ← câu 1 & 2 rất giống
# [0.8234, 1.0000, 0.0389],
# [0.0512, 0.0389, 1.0000]]) ← câu 3 rất khác
Popular models
| Model | Dimensions | Speed | Quality | Use cases |
|---|---|---|---|---|
| all-MiniLM-L6-v2 | 384 | Fast | Good | General purpose |
| all-mpnet-base-v2 | 768 | Average | Very good | When you need high quality |
| intfloat/e5-large-v2 | 1024 | Slow | Excellent | Production search |
| BAAI/bge-m3 | 1024 | Slow | Excellent | Multilingual (Vietnamese) |
3. Application: Semantic Search
from sentence_transformers import SentenceTransformer, util
import torch
model = SentenceTransformer('all-MiniLM-L6-v2')
# "Database" tài liệu
documents = [
"Python là ngôn ngữ lập trình phổ biến nhất cho AI",
"JavaScript thống trị lĩnh vực web development",
"Docker giúp đóng gói ứng dụng thành containers",
"Kubernetes orchestrate containers trên production",
"NLP giúp máy tính hiểu ngôn ngữ con người",
"Computer Vision xử lý hình ảnh và video",
]
doc_embeddings = model.encode(documents, convert_to_tensor=True)
# Search
query = "làm sao để deploy ứng dụng?"
query_embedding = model.encode(query, convert_to_tensor=True)
# Tìm top-3 tài liệu liên quan nhất
scores = util.cos_sim(query_embedding, doc_embeddings)[0]
top_results = torch.topk(scores, k=3)
print("Query:", query)
for score, idx in zip(top_results.values, top_results.indices):
print(f" [{score:.4f}] {documents[idx]}")
# [0.6234] Docker giúp đóng gói ứng dụng thành containers
# [0.5891] Kubernetes orchestrate containers trên production
# [0.1234] JavaScript thống trị lĩnh vực web development
4. Application: Document Clustering
from sklearn.cluster import KMeans
import numpy as np
# Encode documents
embeddings = model.encode(documents)
# K-Means clustering
kmeans = KMeans(n_clusters=3, random_state=42)
clusters = kmeans.fit_predict(embeddings)
for i, (doc, cluster) in enumerate(zip(documents, clusters)):
print(f" Cluster {cluster}: {doc}")
5. Application: Deduplication
from sentence_transformers import util
texts = [
"NLP xử lý ngôn ngữ tự nhiên",
"Xử lý ngôn ngữ tự nhiên bằng NLP",
"Machine learning rất thú vị",
"ML là một lĩnh vực hấp dẫn",
"Hôm nay trời đẹp",
]
embeddings = model.encode(texts)
cosine_scores = util.cos_sim(embeddings, embeddings)
# Tìm cặp trùng lặp (similarity > 0.8)
threshold = 0.8
duplicates = []
for i in range(len(texts)):
for j in range(i + 1, len(texts)):
if cosine_scores[i][j] > threshold:
duplicates.append((i, j, cosine_scores[i][j].item()))
print(f" Duplicate: [{cosine_scores[i][j]:.4f}]")
print(f" A: {texts[i]}")
print(f" B: {texts[j]}")
Summary
| Method | Quality | Speed | OOV | Multilingual |
|---|---|---|---|---|
| Average Word2Vec | Low | Fast | No | No |
| Doc2Vec | Average | Fast | No | No |
| Sentence-BERT | Cao | Average | Yes | Depending on model |
| E5/BGE | Very high | Slower | Yes | Yes |
Next article
Lesson 7: RNN & LSTM — Enter the world of deep learning for NLP: sequential processing with recurrent neural networks.