Chuyển đến nội dung chính

Lesson 6: Hybrid Retrieval - BM25 + Vector + Reranker

Combine lexical and semantic retrieval with RRF, add a reranker to improve precision, groundedness, and citation accuracy.

🧠 AI & ML — L1 Lesson 6: Hybrid Retrieval - BM25 + Vector + Reranker Gemma 4 Local AI Engineering on Mac Part 3: RAG Engineering for Internal Data xdev.asia

Introduction

Vector-only search is usually sufficient for demos, but as internal data grows, you need hybrid retrieval to maintain consistent quality.

1. Role of Each Retriever

  • BM25: strong for rare keywords, error codes, table names
  • Vector: strong for semantics and paraphrasing
  • Reranker: fine-tunes the final candidate list

2. Fusion with RRF

Formula:

$$ score(d) = \sum_{r \in R} \frac{1}{k + rank_r(d)} $$

With $k=60$, RRF is generally stable without complex tuning.

Pseudo-code:

def rrf_fuse(rankings, k=60):
    score = {}
    for ranking in rankings:
        for i, doc_id in enumerate(ranking, start=1):
            score[doc_id] = score.get(doc_id, 0) + 1.0 / (k + i)
    return sorted(score.items(), key=lambda x: x[1], reverse=True)

3. Recommended Pipeline

  1. BM25 top 20
  2. Vector top 20
  3. RRF -> top 25 candidates
  4. Reranker -> top 5 contexts
  5. Build prompt and call Gemma 4

4. Practical Tuning

  • Increase top_k if recall is low
  • Decrease top_k if latency is too high
  • Enable reranker conditionally for complex queries

You can use a heuristic: enable the reranker when the query is long or the score gap is small.

5. Hallucination Control

Mandatory practices:

  • Prompt guardrail to use only context
  • Return a fallback response when data is insufficient
  • Require citations by doc_id - section

Demo Code

Hybrid retrieval combining BM25 + Vector search with RRF fusion:

Hybrid Retrieval

Source code: 05-hybrid-retrieval

Summary