Chuyển đến nội dung chính

Lesson 2: Fine-tuning vs RAG — The biggest AI debate of 2025

Detailed comparison of Fine-tuning vs RAG: Knowledge gap vs Behavior gap. Practical decision checklist. Hybrid approach. Actual case studies: when RAG wins, when Fine-tuning wins.

🧠 AI & ML — Lesson 1 Lesson 2: Fine-tuning vs RAG — Competition Biggest discussion AI 2025

Fine-tuning LLM: The Art of AI Tuning

Part 1: Overview & Strategy — When to Fine-tune?

xdev.asia

Introduction

"Should I use Fine-tuning or RAG?" — this is the most asked question in every AI meetup, forum, and interview in 2025–2026. Correct answer: Depends on the problem you are solving. This article gives you a framework to answer correctly.


1. Diagnosis: Knowledge Gap vs Behavior Gap

Core principles

┌─────────────────────────────────────────────────────┐
│                                                     │
│   Model KHÔNG BIẾT thông tin bạn cần?               │
│   → Knowledge Gap → RAG 📚                          │
│                                                     │
│   Model BIẾT nhưng KHÔNG LÀM ĐÚNG cách bạn muốn?   │
│   → Behavior Gap → Fine-tuning 🎯                   │
│                                                     │
│   Cả hai?                                           │
│   → Fine-tuning + RAG (Hybrid) 🔀                   │
│                                                     │
└─────────────────────────────────────────────────────┘

2. Compare details

2.1 Comprehensive comparison table

CriteriaRAGFine-tuning
ResolvedKnowledge gap (lack of information)Behavior gap (behavior)
Data subject to changeRegular → Strong RAGLittle change → FT suitable
UpdateInstant (update DB)Slow (retrain model)
ExplainableCao (source cite)Low (black box)
Setup costs$50–$500$50–$10,000+
Maintenance costsLow (update data only)High (re-train when needed)
LatencySlower (more retrieval step)Faster (no retrieval needed)
AccuracyDepends on quality retrievalDepends on training data
HallucinationReduce (with source)It's still possible (if the data is bad)
Large scaleRetrieval cost/timeCosts 1 training session

2.2 Practical example

Case 1: Chatbot hỗ trợ khách hàng cần biết chính sách công ty
→ Chính sách thay đổi thường xuyên
→ Cần cite nguồn cho customer
→ RAG THẮNG ✅

Case 2: Model phải trả lời bằng tiếng Việt, formal, format markdown cụ thể
→ Đây là "hành vi" không phải "kiến thức"
→ Prompt engineering không ổn định
→ FINE-TUNING THẮNG ✅

Case 3: Model y khoa cần biết thuật ngữ chuyên ngành VÀ access medical records
→ Thuật ngữ = behavior (fine-tune)
→ Medical records = knowledge (RAG)
→ HYBRID THẮNG ✅

Case 4: Model cần trả lời giá sản phẩm real-time
→ Giá thay đổi liên tục
→ Fine-tune sẽ bị outdated ngay lập tức
→ RAG (hoặc Tool Use) THẮNG ✅

Case 5: Model nhỏ (Flash/Mini) cần perform như model lớn (Pro/4o)
→ "Chắt lọc" kiến thức từ model lớn xuống nhỏ
→ Distillation = một dạng fine-tuning
→ FINE-TUNING THẮNG ✅

3. Decision Flowchart

                    ┌─────────────────────┐
                    │  Bạn cần gì từ LLM? │
                    └──────────┬──────────┘
                               │
              ┌────────────────┼────────────────┐
              ▼                ▼                ▼
    ┌─────────────┐  ┌──────────────┐  ┌──────────────┐
    │ Kiến thức   │  │ Hành vi      │  │ Cả hai       │
    │ mới/riêng   │  │ /Style/Format│  │              │
    └──────┬──────┘  └──────┬───────┘  └──────┬───────┘
           │                │                  │
           ▼                ▼                  ▼
    ┌──────────┐    ┌─────────────┐    ┌──────────────┐
    │ Data thay│    │Prompt eng.  │    │ FT cho style │
    │ đổi nhiều│    │ đã thử?     │    │ + RAG cho    │
    │ không?   │    │             │    │   knowledge  │
    └─────┬────┘    └──────┬──────┘    └──────────────┘
     Yes  │  No        No  │  Yes
      │   │             │  │
      ▼   ▼             ▼  ▼
    ┌───┐┌────┐    ┌───┐┌─────────┐
    │RAG││Cả 2│    │Thử││Fine-tune│
    │   ││    │    │PE ││         │
    └───┘└────┘    └───┘└─────────┘

4. Hybrid Approach — Best of Both Worlds

4.1 Hybrid Architecture

# Fine-tune model cho: style, format, domain terminology
# RAG cho: factual data, recent information

class HybridAI:
    def __init__(self):
        self.model = "ft:gpt-4o-mini:xdev:customer-support:abc123"  # Fine-tuned
        self.rag = RAGPipeline(collection="company_docs")           # RAG
    
    def answer(self, question):
        # Step 1: Retrieve relevant context
        context = self.rag.search(question, top_k=3)
        
        # Step 2: Use fine-tuned model with context
        response = openai.chat.completions.create(
            model=self.model,  # Fine-tuned model → đúng style/format
            messages=[
                {"role": "system", "content": f"Context:\n{context}"},
                {"role": "user", "content": question}
            ]
        )
        return response.choices[0].message.content

4.2 When to use Hybrid?

  • Needs both separate style AND separate data
  • Large enterprise system
  • Specialized domains (medical, legal, financial)
  • Budget is enough for both

5. Practical Case Studies

Case Study 1: Customer Support Bot — RAG wins

Problem: Chatbot needs to answer questions about 500+ products, policies change weekly.

Try Fine-tuning: The model is under the old policy, each update requires training → costs $200/time × 4 times/month = $800/month.

Try RAG: Update database in 5 minutes, retrieval cost ~$0.001/query. Monthly cost: ~$50.

Conclusion: RAG is 16x cheaper and always up-to-date.

Case Study 2: Code Review Bot — Fine-tuning wins

Problem: Model needs to review code according to the team's own coding standards (naming conventions, architectural patterns, very specific error handling style).

Try Prompt: System prompt is too long (3000 tokens), still not consistent.

Try RAG: Coding standards document does not have enough context, output is too generic.

Try Fine-tuning: 200 examples (code + review comments) → model consistency 95%+, system prompt reduced from 3000 → 200 tokens.

Conclusion: Fine-tuning reduces token cost by 93% + increases consistency.

Case Study 3: Medical Q&A — Hybrid wins

Problem: Medical chatbots need to understand specialized terminology AND respond based on patient records.

Solution: Fine-tune for medical Vietnamese terminology + RAG for patient records.


6. Cost Comparison: Concrete Numbers

Scenario: 10,000 queries/day, 30 days

ApproachSetup costMonthly inferenceTotal/month
Base model + Prompt$0~$300$300
RAG ​​$200 (1 time)~$400 (retrieval overhead)$400
Fine-tuning$100–$500 (1 time)~$250 (shorter prompts)$250
Hybrid$500~$350$350

💡 Fine-tuning can be cheaper than the base model if you can shorten the system prompt (less tokens = less money). But including maintenance cost!


Lesson summary

  • Knowledge gap → RAG | Behavior gap → Fine-tuning | Both → Hybrid
  • Data changes frequently → RAG (instant update)
  • Need high consistency in style/format → Fine-tuning
  • 80% of cases → Prompt Engineering or RAG is enough
  • Hybrid approach is the industry standard for enterprises
  • Always calculate total cost (training + inference + maintenance)

Exercises

  1. Analyze 5 use cases in your company → classify Knowledge vs Behavior gap
  2. Draw a decision flowchart for a specific use case
  3. Calculate estimated costs: RAG vs Fine-tuning for that use case
  4. Hybrid architecture design for a practical system