Chuyển đến nội dung chính

Lesson 3: Vector Databases — Chroma, Qdrant, Pinecone, Weaviate

Compare the 4 most popular DB vectors: setup, API, performance, pricing. HNSW index, IVF, PQ. Hybrid search (vector + keyword). Metadata filtering. Hands-on with ChromaDB and Qdrant.

🧠 AI & ML — Lesson 2 Lesson 3: Vector Databases — Chroma, Qdrant, Pinecone, Weaver

Real Battle RAG: From Basic to Advanced

Part 1: RAG Platform

xdev.asia

Compare Vector Databases: Chroma, Qdrant, Pinecone, Weaviate

Introduction

In the previous lesson, we turned text into vectors. Next question: where to save vectors and how to search?

Traditional databases (MySQL, PostgreSQL) are designed for exact match ("SELECT WHERE id = 5"). But with vectors, we need to find nearest neighbors — "which vector is closest to the query vector?" That's why we need Vector Database.


1. What is Vector Database?

1.1 Quick comparison

SQL DatabaseVector Database
Save whatRows & columnsVectors (sequence of numbers) + metadata
QueryWHERE name = "Minh"NEAREST(query_vector, k=5)
SearchExact matchApproximate nearest neighbor (ANN)
Use caseTraditional CRUDSemantic search, RAG, recommendation

1.2 Four most popular Vector Databases

DBTypeHostingMain advantagesPricing (free tier)
ChromaDBEmbeddedLocalExtremely simple, perfect for prototyping100% free
QdrantClient-ServerSelf-hosted / CloudHigh performance, strong filteringFree (self-hosted)
PineconeManaged CloudCloud onlyZero-ops, easy scalingFree (100K vectors)
WeaviateClient-ServerSelf-hosted / CloudGraphQL API, multimodalFree (self-hosted)

2. ChromaDB — Get started in 5 minutes

ChromaDB is the best choice for newbies — runs right in Python, no server setup required.

2.1 Basic Setup & CRUD

pip install chromadb
"""ChromaDB: Vector DB đơn giản nhất"""
import chromadb

# === 1. Khởi tạo ===
# In-memory (tạm thời, mất khi tắt)
client = chromadb.Client()

# Persistent (lưu disk, giữ lại)
# client = chromadb.PersistentClient(path="./chroma_data")

# === 2. Tạo Collection (giống "table" trong SQL) ===
collection = client.create_collection(
    name="company_docs",
    metadata={"description": "Internal company documents"}
)

# === 3. Thêm documents ===
collection.add(
    documents=[
        "Chính sách nghỉ phép: 15 ngày/năm cho nhân viên full-time",
        "Quy trình xin phép: gửi đơn trước 3 ngày làm việc",
        "Lương thưởng: review mỗi 6 tháng, KPI-based",
        "Chế độ bảo hiểm: BHXH + bảo hiểm tư nhân Bảo Việt",
        "Giờ làm việc: 8:30-17:30, thứ 2-6, flexible ±1 giờ",
    ],
    ids=["doc1", "doc2", "doc3", "doc4", "doc5"],
    metadatas=[
        {"category": "leave", "department": "HR"},
        {"category": "leave", "department": "HR"},
        {"category": "compensation", "department": "HR"},
        {"category": "benefits", "department": "HR"},
        {"category": "policy", "department": "Admin"},
    ]
)
print(f"Added {collection.count()} documents")

# === 4. Tìm kiếm (semantic search) ===
results = collection.query(
    query_texts=["nghỉ phép bao nhiêu ngày?"],
    n_results=3  # Top 3 gần nhất
)

print("\n🔍 Query: 'nghỉ phép bao nhiêu ngày?'")
for doc, distance in zip(results["documents"][0], results["distances"][0]):
    relevance = 1 - distance  # Chuyển distance → relevance
    print(f"  [{relevance:.2%}] {doc[:60]}...")

# === 5. Filtering theo metadata ===
filtered = collection.query(
    query_texts=["chế độ cho nhân viên"],
    n_results=3,
    where={"category": "benefits"}  # Chỉ tìm trong "benefits"
)
print(f"\n🔍 Filtered (benefits only): {filtered['documents'][0]}")

2.2 When to use ChromaDB?

✅ Use when❌ Do not use when
Prototype / POCProduction 1M+ vectors
< 100K documentsMulti-user concurrent
Used in Jupyter/localNeed monitoring/metrics
1 developerTeam needs shared DB

3. Qdrant — Production-Ready Vector DB

3.1 Setup

# Chạy Qdrant server bằng Docker
docker run -p 6333:6333 qdrant/qdrant

# Hoặc dùng Qdrant Cloud (free tier)
pip install qdrant-client

3.2 Hands-on

"""Qdrant: Production vector DB"""
from qdrant_client import QdrantClient
from qdrant_client.models import (
    PointStruct, VectorParams, Distance, Filter,
    FieldCondition, MatchValue
)
from openai import OpenAI

openai = OpenAI()
qdrant = QdrantClient(url="http://localhost:6333")

# === 1. Tạo collection ===
qdrant.create_collection(
    collection_name="company_docs",
    vectors_config=VectorParams(
        size=1536,              # text-embedding-3-small dimension
        distance=Distance.COSINE
    )
)

# === 2. Embed + Insert ===
documents = [
    {"text": "Chính sách nghỉ phép: 15 ngày/năm", "category": "leave"},
    {"text": "Lương review mỗi 6 tháng theo KPI", "category": "salary"},
    {"text": "Bảo hiểm Bảo Việt cho toàn bộ nhân viên", "category": "insurance"},
]

# Embed tất cả cùng lúc (batch)
texts = [d["text"] for d in documents]
response = openai.embeddings.create(
    model="text-embedding-3-small",
    input=texts
)
embeddings = [r.embedding for r in response.data]

# Insert vào Qdrant
points = [
    PointStruct(
        id=i,
        vector=emb,
        payload={"text": doc["text"], "category": doc["category"]}
    )
    for i, (doc, emb) in enumerate(zip(documents, embeddings))
]
qdrant.upsert(collection_name="company_docs", points=points)

# === 3. Search ===
query_emb = openai.embeddings.create(
    model="text-embedding-3-small",
    input=["bảo hiểm sức khỏe"]
).data[0].embedding

results = qdrant.search(
    collection_name="company_docs",
    query_vector=query_emb,
    limit=3
)

print("🔍 Search: 'bảo hiểm sức khỏe'")
for r in results:
    print(f"  [{r.score:.4f}] {r.payload['text']}")

# === 4. Filtered Search ===
filtered = qdrant.search(
    collection_name="company_docs",
    query_vector=query_emb,
    query_filter=Filter(
        must=[FieldCondition(key="category", match=MatchValue(value="insurance"))]
    ),
    limit=3
)

3.3 Qdrant vs ChromaDB

FeatureChromaDBQdrant
Setup1 line of PythonDocker / Cloud
PerformanceGood (< 100K)Excellent (millions)
FilteringBasicVery strong (nested, range)
Multi-tenancy❌✅
REST API❌✅
Monitoring❌✅ Dashboard
ProductionPOC✅ Production-ready

💡 Exercise 3: Using ChromaDB, add 20 documents (from FAQ, docs, or articles). Query 10 different questions. Count: how many times is the top-1 result correct? → This is retrieval accuracy.


4. Detailed comparison of 4 Vector DBs

4.1 Comprehensive comparison table

FeaturesChromaDBQdrantPineconeWeaviate
LanguagePythonRustManagedGo
HostingEmbeddedSelf/CloudCloud onlySelf/Cloud
Index typeHNSWHNSWProprietaryHNSW
Max vectors~1MMillions1B+Millions
FilteringBasicAdvancedAdvancedGraphQL
Hybrid search❌✅✅✅
Multimodal❌❌❌✅
Free tierUnlimitedSelf-hosted100K vecsSelf-hosted
Good forPrototypeProductionEnterpriseMultimodal

4.2 Recommendation

Bạn đang ở giai đoạn nào?

├── Prototype / Học tập
│   └── → ChromaDB ✅ (đơn giản, miễn phí)
│
├── Production (startup / team nhỏ)
│   ├── Self-hosted OK → Qdrant ✅ (mạnh, miễn phí)
│   └── Không muốn ops → Pinecone (managed, trả phí)
│
└── Enterprise (scale lớn)
    ├── Cần multimodal → Weaviate
    └── Cần đơn giản → Pinecone

5. Hybrid Search — Combining Vector + Keyword

5.1 Why is Hybrid needed?

Semantic search (vector) is good at finding similar meanings but weak at:

  • Personal name: "Nguyen Van Binh" (need exact match)
  • Code: "ERR_404" (need keyword)
  • Data: "Q3 2025" (needs to be exact)

Hybrid search = Vector search + BM25 keyword search, get the advantages of both!

5.2 Reciprocal Rank Fusion (RRF)

"""Hybrid Search: Vector + Keyword"""
from rank_bm25 import BM25Okapi
import numpy as np

documents = [
    "Chính sách nghỉ phép 2026: 15 ngày/năm cho full-time",
    "Ticket ERR_404: Server timeout lúc 3AM ngày 15/3",
    "Meeting Q3-2025: doanh thu tăng 23% so với Q2",
    "Nhân viên Nguyễn Văn Bình: đánh giá KPI quý 4",
]

# BM25 keyword search
tokenized = [doc.lower().split() for doc in documents]
bm25 = BM25Okapi(tokenized)

query = "ticket ERR_404"

# Keyword scores
keyword_scores = bm25.get_scores(query.lower().split())

# Vector scores (giả lập — thực tế dùng embedding)
vector_scores = np.array([0.3, 0.7, 0.2, 0.1])  # Từ vector search

# RRF fusion
def rrf_score(rank, k=60):
    return 1 / (k + rank)

# Combine rankings
keyword_ranks = np.argsort(-keyword_scores) + 1
vector_ranks = np.argsort(-vector_scores) + 1

final_scores = []
for i in range(len(documents)):
    kr = np.where(keyword_ranks == i+1)[0][0] + 1
    vr = np.where(vector_ranks == i+1)[0][0] + 1
    score = rrf_score(kr) + rrf_score(vr)
    final_scores.append(score)

# Sort by combined score
ranked = sorted(enumerate(final_scores), key=lambda x: -x[1])

print(f"Query: '{query}'\n")
for rank, (idx, score) in enumerate(ranked, 1):
    print(f"  #{rank} [{score:.4f}] {documents[idx][:50]}...")

💡 Exercise 5: Create 10 documents mixed Vietnamese + code + personal name. Test: (a) vector search only, (b) keyword search only, (c) hybrid. Accuracy of each method?


Summary

ConceptsRemember
Vector DBStore vectors + find nearest neighbors quickly
ChromaDBPrototype, simple, embedded, free
QdrantProduction, Rust, strong, self-hosted
PineconeManaged cloud, zero-ops, enterprise
Hybrid SearchVector + Keyword = best for most cases
Metadata FilteringFind the NEAREST vector + filter by category/date

General exercises

  1. ✅ Complete small exercises (3, 5)
  2. ChromaDB RAG: Using ChromaDB + OpenAI, build a RAG for a real PDF file. Test 10 questions.
  3. Qdrant Setup: Run Qdrant using Docker, migrate data from ChromaDB. Compare search speed.
  4. Metadata Design: Given a knowledge base (FAQ, docs), design the metadata schema: which fields need filtering? (category, date, author, department...)

Next article: Document Loading — handle PDF, DOCX, Web, YouTube, code repos to prepare data for RAG.