Chuyển đến nội dung chính

Lesson 4: Collect & Design Dataset for Fine-tuning

Types of datasets: instruction-following, conversation, classification. Standard JSONL format. Collect data from logs, documents, user feedback. Synthetic data generation. Quality vs Quantity.

🧠 AI & ML — Lesson 3 Lesson 4: Collect & Design Dataset for Fine tuning

Fine-tuning LLM: The Art of AI Tuning

Part 2: Data Preparation — The foundation of all success

xdev.asia

Introduction

"Garbage in, garbage out" — it's never been truer than when it comes to fine-tuning. Quality data set is 90% success.


1. JSONL standard Dataset format

1.1 Instruction-following format

{"messages": [
  {"role": "system", "content": "Bạn là trợ lý y khoa tiếng Việt."},
  {"role": "user", "content": "Triệu chứng sốt xuất huyết?"},
  {"role": "assistant", "content": "Sốt xuất huyết dengue có các triệu chứng chính:\n1. Sốt cao đột ngột 39-40°C\n2. Đau đầu dữ dội..."}
]}

1.2 Multi-turn conversation format

{"messages": [
  {"role": "system", "content": "..."},
  {"role": "user", "content": "Câu hỏi 1"},
  {"role": "assistant", "content": "Trả lời 1"},
  {"role": "user", "content": "Follow-up"},
  {"role": "assistant", "content": "Trả lời follow-up"}
]}

2. Data source

2.1 From production logs

# Extract từ customer support logs
def extract_training_data(support_logs):
    training_data = []
    for log in support_logs:
        if log["customer_rating"] >= 4:  # Chỉ lấy conversations tốt
            training_data.append({
                "messages": [
                    {"role": "system", "content": SYSTEM_PROMPT},
                    {"role": "user", "content": log["customer_question"]},
                    {"role": "assistant", "content": log["agent_response"]}
                ]
            })
    return training_data

2.2 Synthetic Data Generation

def generate_synthetic_data(seed_examples, n=100):
    prompt = f"Given these examples, generate {n} similar but diverse examples..."
    # Dùng GPT-4o/Claude để generate training data cho model nhỏ hơn

3. How much data is enough?

Use casesMinimumRecommendedExcellent
Style/tone change50200500+
Domain-specific1005002,000+
Classification50/class200/class1,000+/class
Complex reasoning2001,0005,000+

💡 Quality >> Quantity: 100 perfect examples > 1,000 average examples


Summary

  • JSONL format with messages array is the most popular standard
  • Data sources: production logs, manual creation, synthetic generation
  • Quality > Quantity — invest time in data with the highest ROI
  • Start with 100–200 high quality examples

Exercises

  1. Create 50 training examples for the use case you choose
  2. Try synthetic data generation — compare manual vs synthetic quality
  3. Validate dataset: check format, handle edge cases
  4. Split into train (80%) / validation (10%) / test (10%)