Chuyển đến nội dung chính

Lesson 6: Metadata, Filtering & Hybrid Search

Attach metadata to chunks, filter by field, combine vector search + keyword search (BM25) for more accurate retrieval. Self-query retriever.

🧠 AI & ML — Lesson 5 Lesson 6: Metadata, Filtering & Hybrid Search

Real Battle RAG: From Basic to Advanced

Part 2: Document Processing Pipeline

xdev.asia

Introduction

In the previous lesson, you learned how to chunk documents. But pure vector search has one major drawback: it only searches by meaning (semantic), cannot filter by attributes (creation date, author, document type...).

Example: User asks "2025 leave policy". Vector search may return a 2023 policy for similar content — but the wrong year! Metadata filtering solves: year == 2025 AND category == "HR".

This article covers 3 techniques for upgrading retrieval:

  1. Metadata — attaches additional information to each chunk
  2. Filtering — filtering chunks before/after searching
  3. Hybrid Search — vector + keyword combination (BM25)

1. Metadata — Attach information to Chunks

1.1 What is Metadata?

Each chunk in the vector store consists of 3 parts:

┌─────────────────────────────────────────┐
│  Chunk                                   │
│  ├── content: "Nghỉ phép 15 ngày..."    │
│  ├── embedding: [0.12, -0.34, ...]      │  ← vector search dùng
│  └── metadata: {                         │  ← filtering dùng
│        source: "hr-policy.pdf",          │
│        page: 5,                          │
│        year: 2025,                       │
│        department: "HR",                 │
│        author: "Nguyen Van A"            │
│      }                                   │
└─────────────────────────────────────────┘

1.2 Automatically extract metadata

"""Gắn metadata khi chunk tài liệu"""
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.document_loaders import PyPDFLoader
from datetime import datetime

# Load PDF — tự động có metadata page number
loader = PyPDFLoader("hr-policy-2025.pdf")
pages = loader.load()

# Thêm metadata custom
for page in pages:
    page.metadata.update({
        "source_type": "pdf",
        "department": "HR",
        "year": 2025,
        "language": "vi",
        "last_updated": "2025-01-15",
    })

# Chunk — metadata được kế thừa cho mỗi chunk
splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50)
chunks = splitter.split_documents(pages)

print(chunks[0].metadata)
# {'source': 'hr-policy-2025.pdf', 'page': 0,
#  'source_type': 'pdf', 'department': 'HR', 'year': 2025, ...}

1.3 What should Metadata be attached to?

Metadata fieldsExampleUse cases
source"hr-policy.pdf"Retrieve source
page5User verify
year / date2025Filter by time
category"HR", "Finance"Filter by department
language"vi", "en"Multi-language RAG
author"Nguyen Van A"Filter by author
chunk_index3Sort order
doc_type"policy", "faq"Document classification

💡 Exercise 1: Load a folder containing 5 different files (PDF, TXT, DOCX). Automatically attach metadata including: source, file_type, file_size, created_date.


2. Metadata Filtering

2.1 Filter when querying

"""Filter metadata trong Chroma"""
from langchain_community.vectorstores import Chroma
from langchain_openai import OpenAIEmbeddings

# Index chunks (đã có metadata)
vectorstore = Chroma.from_documents(
    chunks,
    OpenAIEmbeddings(model="text-embedding-3-small"),
    collection_name="company_docs"
)

# Search KHÔNG filter — trả về kết quả từ mọi phòng ban
results = vectorstore.similarity_search("nghỉ phép bao nhiêu ngày?", k=5)

# Search CÓ filter — chỉ tìm trong tài liệu HR năm 2025
results = vectorstore.similarity_search(
    "nghỉ phép bao nhiêu ngày?",
    k=5,
    filter={"year": 2025, "department": "HR"}  # Exact match
)

# Filter phức tạp: $and, $or, $gt, $lt, $in
results = vectorstore.similarity_search(
    "chính sách lương",
    k=5,
    filter={
        "$and": [
            {"year": {"$gte": 2024}},           # Năm >= 2024
            {"department": {"$in": ["HR", "Finance"]}},  # HR hoặc Finance
        ]
    }
)

2.2 Self-Query Retriever — Self-parse filter from the query

"""Self-Query: AI tự tách query thành search + filter"""
from langchain.retrievers import SelfQueryRetriever
from langchain.chains.query_constructor.base import AttributeInfo
from langchain_openai import ChatOpenAI

# Mô tả metadata fields cho LLM hiểu
metadata_field_info = [
    AttributeInfo(name="year", description="Năm ban hành", type="integer"),
    AttributeInfo(name="department", description="Phòng ban: HR, Finance, IT", type="string"),
    AttributeInfo(name="doc_type", description="Loại: policy, faq, guide", type="string"),
]

retriever = SelfQueryRetriever.from_llm(
    llm=ChatOpenAI(model="gpt-4o-mini", temperature=0),
    vectorstore=vectorstore,
    document_contents="Tài liệu nội bộ công ty về chính sách và quy trình",
    metadata_field_info=metadata_field_info,
)

# User hỏi: "Chính sách HR năm 2025 về nghỉ phép"
# → LLM tự parse:
#   search_query = "chính sách nghỉ phép"
#   filter = {"year": 2025, "department": "HR"}
results = retriever.invoke("Chính sách HR năm 2025 về nghỉ phép")
Flow:
User query: "Chính sách HR năm 2025 về nghỉ phép"
                    │
          ┌─────────┴─────────┐
          │   Self-Query LLM  │
          │   (parse intent)  │
          └─────────┬─────────┘
                    │
    ┌───────────────┼───────────────┐
    │               │               │
search_query    filter_year    filter_dept
"nghỉ phép"      2025           "HR"
    │               │               │
    └───────────────┼───────────────┘
                    │
          ┌─────────┴─────────┐
          │   Vector Store    │
          │  (search+filter)  │
          └─────────┬─────────┘
                    │
              Filtered results

💡 Exercise 2: Create a Self-Query Retriever for a document set with at least 3 metadata fields. Test with 5 natural questions. Check if LLM parse filter is correct.


3. Hybrid Search — Vector + Keyword

3.1 Problems of pure Vector Search

Query: "Nghị định 168/2024/NĐ-CP"

Vector search: tìm theo ý nghĩa → có thể trả về
               Nghị định 150/2023 (nội dung tương tự nhưng SAI số!)

Keyword search (BM25): tìm đúng "168/2024/NĐ-CP" → ĐÚNG

→ Kết hợp cả 2 = Hybrid Search
Search typeStrongWeak
VectorUnderstanding meaning, synonym, contextWrong when needing exact match (code, number, name)
Keyword (BM25)Exact match, code, proper nounsDon't understand synonym, context
HybridCombining both advantagesNeed to tune weight

3.2 Implement Hybrid Search

"""Hybrid search với BM25 + Vector"""
from langchain_community.retrievers import BM25Retriever
from langchain.retrievers import EnsembleRetriever
from langchain_community.vectorstores import Chroma
from langchain_openai import OpenAIEmbeddings

# Chuẩn bị documents (đã chunk)
# chunks = [Document(...), Document(...), ...]

# 1. Vector retriever
vectorstore = Chroma.from_documents(chunks, OpenAIEmbeddings())
vector_retriever = vectorstore.as_retriever(search_kwargs={"k": 5})

# 2. BM25 retriever (keyword-based)
bm25_retriever = BM25Retriever.from_documents(chunks, k=5)

# 3. Ensemble (hybrid) — weight 50/50
hybrid_retriever = EnsembleRetriever(
    retrievers=[vector_retriever, bm25_retriever],
    weights=[0.5, 0.5],  # Tùy chỉnh: 0.7/0.3 nếu ưu tiên vector
)

# Query
results = hybrid_retriever.invoke("Nghị định 168/2024/NĐ-CP")

3.3 Reciprocal Rank Fusion (RRF)

When combining results from 2 retrievers, a merge + rank method is needed:

Vector results:        BM25 results:
1. Doc A (score 0.95)  1. Doc C (score 8.2)
2. Doc B (score 0.88)  2. Doc A (score 7.1)
3. Doc C (score 0.82)  3. Doc D (score 6.5)

RRF formula: score(d) = Σ 1/(k + rank(d))  (k=60 default)

Doc A: 1/(60+1) + 1/(60+2) = 0.0164 + 0.0161 = 0.0325  ← Top 1!
Doc C: 1/(60+3) + 1/(60+1) = 0.0159 + 0.0164 = 0.0323  ← Top 2
Doc B: 1/(60+2) + 0       = 0.0161                      ← Top 3
Doc D: 0       + 1/(60+3) = 0.0159                      ← Top 4

Doc A appears on both retrievers → highest rank!

3.4 Pinecone Hybrid Search (Production-ready)

"""Pinecone native hybrid search — sparse + dense vectors"""
from pinecone import Pinecone
from pinecone_text.sparse import BM25Encoder

# Sparse encoder (BM25)
bm25 = BM25Encoder()
bm25.fit([chunk.page_content for chunk in chunks])

# Dense encoder (embedding)
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")

# Index với cả 2 loại vector
pc = Pinecone(api_key="your-key")
index = pc.Index("hybrid-rag")

for chunk in chunks:
    dense = embeddings.embed_query(chunk.page_content)
    sparse = bm25.encode_queries(chunk.page_content)
    
    index.upsert(vectors=[{
        "id": chunk.metadata.get("id", str(hash(chunk.page_content))),
        "values": dense,          # Dense vector
        "sparse_values": sparse,  # Sparse vector (BM25)
        "metadata": chunk.metadata
    }])

# Query hybrid
query = "Nghị định 168/2024"
results = index.query(
    vector=embeddings.embed_query(query),
    sparse_vector=bm25.encode_queries(query),
    top_k=5,
    alpha=0.5,  # 0=pure sparse, 1=pure dense, 0.5=hybrid
)

💡 Exercise 3: Implement hybrid search on a set of documents. Compare results: (a) vector only, (b) BM25 only, (c) hybrid. Use 10 test questions, record the accuracy of each type.


4. Tuning Hybrid Weights

4.1 When to prioritize Vector vs Keyword?

Use casesVectorweightBM25 weightReason
FAQ / General Q&A0.70.3Users ask in many different ways
Law / Code0.30.7Need exact match rule code
Technical documents0.50.5Need both keyword and semantic
Multi-language0.80.2Vector is better for cross-language
Code documentation0.40.6Function names = keyword

4.2 Auto-tune weights

"""Benchmark hybrid weights trên golden test set"""
test_queries = [
    {"q": "nghỉ phép bao nhiêu ngày", "expected_doc": "hr-policy.pdf"},
    {"q": "Nghị định 168/2024", "expected_doc": "legal/nd168.pdf"},
    # ... 10+ câu test
]

weight_configs = [
    (0.3, 0.7), (0.4, 0.6), (0.5, 0.5),
    (0.6, 0.4), (0.7, 0.3), (0.8, 0.2),
]

best_config = None
best_accuracy = 0

for vec_w, bm25_w in weight_configs:
    hybrid = EnsembleRetriever(
        retrievers=[vector_retriever, bm25_retriever],
        weights=[vec_w, bm25_w],
    )
    
    correct = 0
    for test in test_queries:
        results = hybrid.invoke(test["q"])
        sources = [r.metadata["source"] for r in results[:3]]
        if test["expected_doc"] in sources:
            correct += 1
    
    accuracy = correct / len(test_queries)
    print(f"Vector={vec_w}, BM25={bm25_w}: {accuracy:.0%}")
    
    if accuracy > best_accuracy:
        best_accuracy = accuracy
        best_config = (vec_w, bm25_w)

print(f"\nBest: Vector={best_config[0]}, BM25={best_config[1]} ({best_accuracy:.0%})")

Summary

ConceptsRemember
MetadataAdditional information attached to the chunk (source, year, category...)
FilteringFilter chunks by metadata before/after search
Self-QueryLLM automatically parses the question into search + filter
BM25Keyword search, strong with exact match
Hybrid SearchVector + BM25, combining the advantages of both
RRFReciprocal Rank Fusion — merge 2 retriever results
Weight tuningBenchmark on the golden test set to choose the ratio

General exercises

  1. ✅ Complete 3 small exercises (1, 2, 3)
  2. Full Pipeline: Load 10+ documents → attach full metadata → index into Chroma → implement hybrid search → compare accuracy vector vs hybrid on 20 test questions.
  3. Self-Query + Hybrid: Combine Self-Query Retriever with Hybrid Search. User asked "HR policy in 2025 on salary" → self filter department + year + hybrid search content.
  4. Dashboard: Create Streamlit app: upload documents → attach metadata → search with UI filter (dropdown select year, department...).

Next article: Query Transformation — HyDE, Multi-Query, Step-Back — turns 1 question into many variations for more accurate searching.