Introduction
Vector-only search is usually sufficient for demos, but as internal data grows, you need hybrid retrieval to maintain consistent quality.
1. Role of Each Retriever
- BM25: strong for rare keywords, error codes, table names
- Vector: strong for semantics and paraphrasing
- Reranker: fine-tunes the final candidate list
2. Fusion with RRF
Formula:
$$ score(d) = \sum_{r \in R} \frac{1}{k + rank_r(d)} $$
With $k=60$, RRF is generally stable without complex tuning.
Pseudo-code:
def rrf_fuse(rankings, k=60):
score = {}
for ranking in rankings:
for i, doc_id in enumerate(ranking, start=1):
score[doc_id] = score.get(doc_id, 0) + 1.0 / (k + i)
return sorted(score.items(), key=lambda x: x[1], reverse=True)
3. Recommended Pipeline
- BM25 top 20
- Vector top 20
- RRF -> top 25 candidates
- Reranker -> top 5 contexts
- Build prompt and call Gemma 4
4. Practical Tuning
- Increase top_k if recall is low
- Decrease top_k if latency is too high
- Enable reranker conditionally for complex queries
You can use a heuristic: enable the reranker when the query is long or the score gap is small.
5. Hallucination Control
Mandatory practices:
- Prompt guardrail to use only context
- Return a fallback response when data is insufficient
- Require citations by
doc_id - section
Demo Code
Hybrid retrieval combining BM25 + Vector search with RRF fusion:

Source code: 05-hybrid-retrieval