簡介
「垃圾輸入,垃圾輸出」 — 在微調方面,這一點最為真實。品質數據整合功率達到 90%。
1. JSONL 標準資料集格式
1.1 指令遵循格式
{"messages": [
{"role": "system", "content": "Bạn là trợ lý y khoa tiếng Việt."},
{"role": "user", "content": "Triệu chứng sốt xuất huyết?"},
{"role": "assistant", "content": "Sốt xuất huyết dengue có các triệu chứng chính:\n1. Sốt cao đột ngột 39-40°C\n2. Đau đầu dữ dội..."}
]}
1.2 多輪對話格式
{"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "Câu hỏi 1"},
{"role": "assistant", "content": "Trả lời 1"},
{"role": "user", "content": "Follow-up"},
{"role": "assistant", "content": "Trả lời follow-up"}
]}
2.資料來源
2.1 來自生產日誌
# Extract từ customer support logs
def extract_training_data(support_logs):
training_data = []
for log in support_logs:
if log["customer_rating"] >= 4: # Chỉ lấy conversations tốt
training_data.append({
"messages": [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": log["customer_question"]},
{"role": "assistant", "content": log["agent_response"]}
]
})
return training_data
2.2 合成資料生成
def generate_synthetic_data(seed_examples, n=100):
prompt = f"Given these examples, generate {n} similar but diverse examples..."
# Dùng GPT-4o/Claude để generate training data cho model nhỏ hơn
3.多少數據才夠?
| 使用案例 | 最低 | 推薦 | 優秀 |
|---|---|---|---|
| 風格/語氣變化 | 50 | 50 200 | 200 500+ |
| 特定領域 | 100 | 100 500 | 500 2,000+ |
| 分類 | 50人/班 | 200人/班 | 1,000+/班 |
| 複雜推理 | 200 | 200 1,000 | 5,000+ |
💡 品質 >> 數量:100 個完美範例 > 1,000 個平均範例
總結
- 帶有訊息數組的 JSONL 格式是最受歡迎的標準
- 資料來源:生產日誌、手動建立、合成生成
- 品質 > 數量 — 將時間投入到具有最高投資報酬率的數據上
- 從 100–200 個高品質範例開始
練習
- 為您選擇的用例建立 50 個訓練範例
- 嘗試產生合成數據-比較手動與合成的質量
- 驗證資料集:檢查格式,處理邊緣狀況
- 分為訓練(80%)/驗證(10%)/測試(10%)