Chuyển đến nội dung chính

Lesson 17: Vector Databases — Embeddings & Semantic Search

Deep understanding of embeddings and vector search algorithms. Compare FAISS, ChromaDB, Qdrant, Pinecone, Weaviate — when to use which. Build a practical semantic search engine and explore multimodal embeddings with CLIP.

🧠 AI & ML — Lesson 16 Lesson 17: Vector Databases — Embeddings & Semantic Search

AI & LLM: From Basics to Advanced

Part 4: Prompting & RAG

xdev.asia

Lesson 17: Vector Databases — Embeddings & Semantic Search

1. What are Embeddings?

Embedding is the process of converting data (text, image, audio) into arithmetic vectors in a high-dimensional space, so that things that are meaningfully close to each other are located close to each other in that space.

Evolutionary history

Word2Vec (2013 — Google): Word-level embedding. "king - man + woman ≈ queen". Problem: each word has only one vector, regardless of context ("bank" is a riverbank or a bank?).

ELMo / BERT (2018): Contextual embeddings — the same word has different vectors depending on the context. Big breakthrough.

Sentence Transformers (2019): Siamese networks are fine-tuned to create high-quality sentence-level embedding for similarity tasks.

LLM Embeddings (2022+): OpenAI text-embedding-3, Cohere embed-v3, Google text-embedding-004 — embedding from large models, understanding deeper semantics.

import numpy as np
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("BAAI/bge-m3")  # hỗ trợ tiếng Việt

sentences = [
    "Hà Nội là thủ đô của Việt Nam",
    "Thành phố Hà Nội nằm ở miền Bắc",
    "Python là ngôn ngữ lập trình phổ biến",
]

embeddings = model.encode(sentences, normalize_embeddings=True)
print(f"Shape: {embeddings.shape}")  # (3, 1024) — 3 câu, 1024 chiều

2. Cosine Similarity vs Euclidean Distance vs Dot Product

Three popular metrics to measure similarity between two vectors:

import numpy as np

def cosine_similarity(a, b):
    """Đo góc giữa hai vector — không phụ thuộc độ dài"""
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

def euclidean_distance(a, b):
    """Đo khoảng cách thực trong không gian"""
    return np.linalg.norm(a - b)

def dot_product(a, b):
    """Tích vô hướng — nhanh nhất, nhưng phụ thuộc độ dài vector"""
    return np.dot(a, b)

# Ví dụ minh họa
v1 = np.array([1.0, 0.0, 0.0])
v2 = np.array([0.9, 0.1, 0.0])
v3 = np.array([0.0, 0.0, 1.0])

print(f"cos(v1,v2) = {cosine_similarity(v1, v2):.4f}")   # ~0.994 — rất giống
print(f"cos(v1,v3) = {cosine_similarity(v1, v3):.4f}")   # 0.000 — vuông góc
print(f"euc(v1,v2) = {euclidean_distance(v1, v2):.4f}")  # ~0.141

# Nếu vectors đã normalize (độ dài = 1), cosine và dot product tương đương
# text-embedding-3-small trả về vectors đã normalize
MetricsWhen to useFeatures
Cosine SimilarityText embeddingsIgnore magnitude, only care about direction
Euclidean DistanceImage features, no normalizeCalculate absolute distance
Dot ProductNormalizedFastest, result = cosine when normalized

3. ANN — Approximate Nearest Neighbor

Finding the exact nearest neighbor in 1 billion vectors with 1536 dimensions is not feasible (O(n)). ANN algorithms trade a little accuracy for extreme speed.

HNSW — Hierarchical Navigable Small World

The most popular algorithm. Build a multi-layer hierarchical graph:

  • Upper layer: Few nodes, long connections — quick navigation to approximate area
  • Lower layer: Many nodes, short connections — accurate search in small areas
Search complexity: O(log n) thay vì O(n)
Recall@10: ~99% với cài đặt đúng

IVF — Inverted File Index

Divide the vector space into clusters (Voronoi cells). When searching, only search in the nearest few clusters:

Build: K-means clustering → assign vectors to clusters
Search: Find nearest clusters → search only in those clusters

4. FAISS — Facebook AI Similarity Search

FAISS is a C++ library (with Python binding) from Meta — fastest for local search, no built-in persistence.

import faiss
import numpy as np

# Tạo index HNSW
dimension = 1536  # text-embedding-3-small
index = faiss.IndexHNSWFlat(dimension, 32)  # 32 = số kết nối mỗi node
index.hnsw.efConstruction = 40  # trade-off: build time vs quality

# Thêm vectors
num_vectors = 10000
vectors = np.random.random((num_vectors, dimension)).astype("float32")
faiss.normalize_L2(vectors)  # normalize cho cosine similarity
index.add(vectors)

# Search
query = np.random.random((1, dimension)).astype("float32")
faiss.normalize_L2(query)

k = 5  # top-5
distances, indices = index.search(query, k)
print(f"Top-5 indices: {indices[0]}")
print(f"Distances: {distances[0]}")

# Lưu và load
faiss.write_index(index, "my_index.faiss")
loaded_index = faiss.read_index("my_index.faiss")

# IVF index cho datasets lớn (>1M vectors)
nlist = 100  # số clusters
quantizer = faiss.IndexFlatL2(dimension)
ivf_index = faiss.IndexIVFFlat(quantizer, dimension, nlist)
ivf_index.train(vectors)  # phải train trước
ivf_index.add(vectors)
ivf_index.nprobe = 10  # search trong 10 clusters gần nhất

When to use FAISS: Small-medium dataset (<10M vectors), needs extremely high speed, environment without network (edge, on-prem).

5. ChromaDB — Embedded, Easy to Use

ChromaDB is the "Energizer battery" vector database for prototyping — running embedded in the process, no need for a separate server.

import chromadb
from chromadb.utils import embedding_functions

# Embedded mode (lưu xuống disk)
client = chromadb.PersistentClient(path="./chroma_data")

# Dùng OpenAI embeddings
openai_ef = embedding_functions.OpenAIEmbeddingFunction(
    api_key="your-key",
    model_name="text-embedding-3-small",
)

collection = client.get_or_create_collection(
    name="documents",
    embedding_function=openai_ef,
    metadata={"hnsw:space": "cosine"},
)

# Thêm documents (ChromaDB tự embed nếu có embedding function)
collection.add(
    documents=[
        "RAG là kỹ thuật kết hợp retrieval với generation",
        "Vector database lưu embeddings và hỗ trợ similarity search",
        "Fine-tuning điều chỉnh weights của pre-trained model",
    ],
    metadatas=[
        {"category": "rag", "difficulty": "intermediate"},
        {"category": "database", "difficulty": "beginner"},
        {"category": "training", "difficulty": "advanced"},
    ],
    ids=["doc1", "doc2", "doc3"],
)

# Query với metadata filtering
results = collection.query(
    query_texts=["làm thế nào để tìm kiếm thông tin?"],
    n_results=2,
    where={"difficulty": {"$ne": "advanced"}},  # loại trừ advanced
)
print(results["documents"])
print(results["distances"])

6. Qdrant — Production-Ready

Qdrant (written in Rust) is a powerful production option with advanced filtering, payloads, and many enterprise features.

from qdrant_client import QdrantClient
from qdrant_client.models import (
    Distance, VectorParams, PointStruct, Filter, FieldCondition, MatchValue
)
import uuid

# Kết nối (local hoặc cloud)
client = QdrantClient(url="http://localhost:6333")
# client = QdrantClient(url="https://xxx.qdrant.io", api_key="your-key")

# Tạo collection
client.create_collection(
    collection_name="articles",
    vectors_config=VectorParams(size=1536, distance=Distance.COSINE),
)

# Upsert points với payload phong phú
points = [
    PointStruct(
        id=str(uuid.uuid4()),
        vector=[0.1] * 1536,  # thay bằng embedding thực
        payload={
            "title": "Giới thiệu RAG",
            "author": "Nguyen Van A",
            "category": "AI",
            "published_year": 2024,
            "tags": ["rag", "llm", "retrieval"],
        },
    )
]
client.upsert(collection_name="articles", points=points)

# Search với filter phức tạp
results = client.search(
    collection_name="articles",
    query_vector=[0.1] * 1536,
    limit=5,
    query_filter=Filter(
        must=[
            FieldCondition(key="category", match=MatchValue(value="AI")),
            FieldCondition(key="published_year", range={"gte": 2023}),
        ]
    ),
    with_payload=True,
)

for r in results:
    print(f"Score: {r.score:.4f} | Title: {r.payload['title']}")

7. Pinecone — Managed Cloud

Pinecone does not need infrastructure management or automatic scaling, but there is a cost.

from pinecone import Pinecone, ServerlessSpec

pc = Pinecone(api_key="your-pinecone-key")

# Tạo index serverless
pc.create_index(
    name="my-rag-index",
    dimension=1536,
    metric="cosine",
    spec=ServerlessSpec(cloud="aws", region="us-east-1"),
)

index = pc.Index("my-rag-index")

# Upsert với namespaces (phân chia logical)
index.upsert(
    vectors=[
        {"id": "v1", "values": [0.1] * 1536, "metadata": {"text": "...", "source": "doc1.pdf"}},
        {"id": "v2", "values": [0.2] * 1536, "metadata": {"text": "...", "source": "doc2.pdf"}},
    ],
    namespace="production",
)

# Query
results = index.query(
    vector=[0.15] * 1536,
    top_k=5,
    namespace="production",
    include_metadata=True,
)

8. Weaviate — Multimodal, GraphQL

Weaviate stands out with its multimodal capabilities (text + image) and GraphQL interface.

import weaviate
from weaviate.classes.init import Auth

client = weaviate.connect_to_weaviate_cloud(
    cluster_url="https://xxx.weaviate.network",
    auth_credentials=Auth.api_key("your-key"),
)

# Schema được define qua Collections
from weaviate.classes.config import Configure, Property, DataType

client.collections.create(
    name="Article",
    vectorizer_config=Configure.Vectorizer.text2vec_openai(),
    properties=[
        Property(name="title", data_type=DataType.TEXT),
        Property(name="content", data_type=DataType.TEXT),
        Property(name="category", data_type=DataType.TEXT),
    ],
)

# Query bằng near_text (semantic search)
articles = client.collections.get("Article")
response = articles.query.near_text(
    query="machine learning applications",
    limit=5,
    return_metadata=["distance"],
)

9. General Comparison

FAISSChromaDBQdrantPineconeWeaviate
DeploymentEmbeddedEmbedded/ServerSelf-hosted/CloudManaged CloudSelf-hosted/Cloud
SpeedExtremely fastFastVery fastFastAverage
FilteringNoneBasicAdvancedMetadata filterGraphQL
MultimodalNoNoLimitationsNoYes (CLIP)
ScaleUp to ~100MUp to ~1MBillionsAutomaticBillions
CostFreeFreeFree tierChargeFree tier
SuitableResearch, localPrototype, devProductionManaged prodMultimodal

Recommendation according to use case:

  • Simple Prototype / RAG: ChromaDB
  • Production on-premise: Qdrant
  • Managed, don't want ops: Pinecone
  • Multimodal search: Weaviate
  • Maximum speed, local: FAISS

10. Code: Build Semantic Search Engine

# pip install sentence-transformers chromadb pandas

import chromadb
from sentence_transformers import SentenceTransformer
import pandas as pd
import json

class SemanticSearchEngine:
    def __init__(self, collection_name: str = "knowledge_base"):
        self.model = SentenceTransformer("BAAI/bge-m3")
        self.client = chromadb.PersistentClient(path="./search_db")
        self.collection = self.client.get_or_create_collection(
            name=collection_name,
            metadata={"hnsw:space": "cosine"},
        )

    def index_documents(self, documents: list[dict]):
        """
        documents: [{"id": "1", "text": "...", "title": "...", "url": "..."}]
        """
        texts = [d["text"] for d in documents]
        print(f"Embedding {len(texts)} documents...")
        embeddings = self.model.encode(texts, normalize_embeddings=True, show_progress_bar=True)

        self.collection.add(
            embeddings=embeddings.tolist(),
            documents=texts,
            metadatas=[{"title": d.get("title", ""), "url": d.get("url", "")} for d in documents],
            ids=[d["id"] for d in documents],
        )
        print(f"Indexed {len(documents)} documents!")

    def search(self, query: str, top_k: int = 5) -> list[dict]:
        query_embedding = self.model.encode([query], normalize_embeddings=True)
        results = self.collection.query(
            query_embeddings=query_embedding.tolist(),
            n_results=top_k,
        )

        hits = []
        for i in range(len(results["ids"][0])):
            hits.append({
                "id": results["ids"][0][i],
                "text": results["documents"][0][i],
                "metadata": results["metadatas"][0][i],
                "score": 1 - results["distances"][0][i], # convert distance to similarity
            })
        return hits

    def stats(self):
        return {"total_documents": self.collection.count()}


# Demo
engine = SemanticSearchEngine()

# Index some sample documents
sample_docs = [
    {"id": "1", "title": "Introducing RAG", "text": "RAG combines retrieval with language generation to reduce hallucination"},
    {"id": "2", "title": "Vector Database", "text": "Vector database stores embeddings and supports semantic similarity search"},
    {"id": "3", "title": "Fine-tuning", "text": "Fine-tuning is the process of continuing to train the model on a specialized dataset"},
    {"id": "4", "title": "Prompt Engineering", "text": "Prompt engineering is a command design technique to optimize the output of LLM"},
    {"id": "5", "title": "LangChain", "text": "LangChain is a Python framework that makes building LLM-based applications easier"},
]

engine.index_documents(sample_docs)

# Search
query = "how to improve the accuracy of AI chatbot?"
results = engine.search(query, top_k=3)

print(f"\nQuery: {query}")
print("\nTop results:")
for r in results:
    print(f" [{r['score']:.4f}] {r['metadata']['title']}: {r['text'][:80]}...")

11. Multimodal Embeddings with CLIP

OpenAI's CLIP (Contrastive Language-Image Pre-training) allows embedding both text and images into the same vector space — the text and related images will be close to each other.

# pip install transformers pillow torch
from transformers import CLIPProcessor, CLIPModel
from PIL import Image
import torch
import requests

model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")

# Embed photos
image = Image.open("product.jpg")
image_inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
    image_features = model.get_image_features(**image_inputs)
    image_features = image_features / image_features.norm(dim=-1, keepdim=True)

# Embed text queries
texts = ["a sleeping cat", "red car", "blue sky"]
text_inputs = processor(text=texts, return_tensors="pt", padding=True)
with torch.no_grad():
    text_features = model.get_text_features(**text_inputs)
    text_features = text_features / text_features.norm(dim=-1, keepdim=True)

# Calculate similarity between images and text queries
similar = (image_features @ text_features.T).squeeze()
for text, sim in zip(texts, similar):
    print(f"{sim:.4f} | {text}")
# Output: text that correctly describes the image will have the highest similarity

CLIP allows building an image search system using text (Google Images DIY style) or RAG with documents containing both text and images.

Summary

Vector databases are the heart of every RAG system. Understanding similarity metrics, ANN algorithms (HNSW, IVF), and each database's characteristics will help you choose the right tool for each situation. With prototype: ChromaDB. Production on-prem: Qdrant. Managed cloud: Pinecone. Multimodal: Weaviate. Requires maximum local speed: FAISS.