Introduction
High-quality RAG starts with proper ingestion. If you chunk with the wrong structure, even the best retrieval won't produce stable answers.
1. Internal Data Sources
Common sources:
- Markdown docs
- Operations runbooks
- PDF processes or compliance documents
- Technical notes from the team
Each source needs its own parser, but they must all output to a unified schema.
2. Data Normalization Before Chunking
Recommended steps:
- Normalize encoding and unicode
- Clean up repeating headers/footers
- Mark code blocks to preserve them
- Split structure by headings
Don't blindly strip Vietnamese diacritics — it reduces retrieval quality.
3. Chunking Strategy
Reference parameters:
- Chunk size: 600-1000 tokens
- Overlap: 80-150 tokens
- Prefer splitting by section rather than hard character cuts
The goal is to maintain complete semantics for each chunk.
4. Metadata Schema
Recommended payload:
{
"doc_id": "pg-backup-v2",
"title": "Backup PostgreSQL",
"section": "3. PITR",
"source": "docs/backup.md",
"language": "vi",
"updated_at": "2026-04-03"
}
Metadata enables accurate filtering by topic, source, and update time.
5. Embedding Pipeline
Best practices:
- Batch embedding in batches
- Cache by content hash
- Only re-embed changed chunks
For large datasets, incremental ingestion saves significant time.
6. Index Lifecycle
Use 2 collections:
active: serves queriesstaging: ingests new data
When staging passes eval, swap to active to reduce downtime risk.
Demo Code
RAG endpoint query result with citations:

Source code: 04-ingestion