Chuyển đến nội dung chính

Lesson 6: Sentence & Document Embeddings — From Doc2Vec to Sentence-BERT

Doc2Vec and Paragraph Vectors. Sentence embeddings: average pooling, Sentence-BERT, E5, BGE. Semantic similarity and cosine distance. Applications: semantic search, clustering, deduplication. Demo with Sentence-Transformers library.

🧠 AI & ML — Lesson 5 Lesson 6: Sentence & Document Embeddings — Words Doc2Vec to Sentence-BERT

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 2: Language Representation — From BoW to Word Embeddings

xdev.asia

Introduction

Word embeddings represent each word — but in reality, we need to compare sentences or documents with each other. Sentence embeddings turn each sentence into a unique vector, allowing for the calculation of semantic similarity — the foundation of semantic search, RAG, and many modern NLP applications.


1. From Word to Sentence Embeddings

Simple method: Average Pooling

import numpy as np

def average_embedding(words, model):
    """Trung bình vector của các từ trong câu."""
    vectors = [model[w] for w in words if w in model]
    if not vectors:
        return np.zeros(model.vector_size)
    return np.mean(vectors, axis=0)

sentence = ["xử", "lý", "ngôn", "ngữ", "tự", "nhiên"]
sent_vec = average_embedding(sentence, word2vec_model)
# Shape: (100,) — một vector duy nhất cho cả câu

⚠️ Average pooling loses order information and context — "not good" will be closer to "good"!


2. Sentence-BERT (SBERT)

Ideas

Fine-tune BERT with Siamese network to create high quality sentence embeddings.

from sentence_transformers import SentenceTransformer, util

# Load pre-trained model
model = SentenceTransformer('all-MiniLM-L6-v2')

# Encode sentences
sentences = [
    "NLP là lĩnh vực xử lý ngôn ngữ tự nhiên",
    "Xử lý ngôn ngữ tự nhiên thuộc về AI",
    "Hôm nay thời tiết đẹp quá",
]
embeddings = model.encode(sentences)
print(embeddings.shape)  # (3, 384)

# Tính cosine similarity
similarities = util.cos_sim(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.8234, 0.0512],  ← câu 1 & 2 rất giống
#         [0.8234, 1.0000, 0.0389],
#         [0.0512, 0.0389, 1.0000]])  ← câu 3 rất khác

Popular models

ModelDimensionsSpeed ​​QualityUse cases
all-MiniLM-L6-v2384FastGoodGeneral purpose
all-mpnet-base-v2768AverageVery goodWhen you need high quality
intfloat/e5-large-v21024SlowExcellentProduction search
BAAI/bge-m31024SlowExcellentMultilingual (Vietnamese)

3. Application: Semantic Search

from sentence_transformers import SentenceTransformer, util
import torch

model = SentenceTransformer('all-MiniLM-L6-v2')

# "Database" tài liệu
documents = [
    "Python là ngôn ngữ lập trình phổ biến nhất cho AI",
    "JavaScript thống trị lĩnh vực web development",
    "Docker giúp đóng gói ứng dụng thành containers",
    "Kubernetes orchestrate containers trên production",
    "NLP giúp máy tính hiểu ngôn ngữ con người",
    "Computer Vision xử lý hình ảnh và video",
]
doc_embeddings = model.encode(documents, convert_to_tensor=True)

# Search
query = "làm sao để deploy ứng dụng?"
query_embedding = model.encode(query, convert_to_tensor=True)

# Tìm top-3 tài liệu liên quan nhất
scores = util.cos_sim(query_embedding, doc_embeddings)[0]
top_results = torch.topk(scores, k=3)

print("Query:", query)
for score, idx in zip(top_results.values, top_results.indices):
    print(f"  [{score:.4f}] {documents[idx]}")
# [0.6234] Docker giúp đóng gói ứng dụng thành containers
# [0.5891] Kubernetes orchestrate containers trên production
# [0.1234] JavaScript thống trị lĩnh vực web development

4. Application: Document Clustering

from sklearn.cluster import KMeans
import numpy as np

# Encode documents
embeddings = model.encode(documents)

# K-Means clustering
kmeans = KMeans(n_clusters=3, random_state=42)
clusters = kmeans.fit_predict(embeddings)

for i, (doc, cluster) in enumerate(zip(documents, clusters)):
    print(f"  Cluster {cluster}: {doc}")

5. Application: Deduplication

from sentence_transformers import util

texts = [
    "NLP xử lý ngôn ngữ tự nhiên",
    "Xử lý ngôn ngữ tự nhiên bằng NLP",
    "Machine learning rất thú vị",
    "ML là một lĩnh vực hấp dẫn",
    "Hôm nay trời đẹp",
]

embeddings = model.encode(texts)
cosine_scores = util.cos_sim(embeddings, embeddings)

# Tìm cặp trùng lặp (similarity > 0.8)
threshold = 0.8
duplicates = []
for i in range(len(texts)):
    for j in range(i + 1, len(texts)):
        if cosine_scores[i][j] > threshold:
            duplicates.append((i, j, cosine_scores[i][j].item()))
            print(f"  Duplicate: [{cosine_scores[i][j]:.4f}]")
            print(f"    A: {texts[i]}")
            print(f"    B: {texts[j]}")

Summary

MethodQualitySpeed ​​OOVMultilingual
Average Word2VecLowFastNoNo
Doc2VecAverageFastNoNo
Sentence-BERTCaoAverageYesDepending on model
E5/BGEVery highSlowerYesYes

Next article

Lesson 7: RNN & LSTM — Enter the world of deep learning for NLP: sequential processing with recurrent neural networks.