簡介
原始數據很少可以用於微調。本文建構了一個專業的資料處理管道。
1. 資料清洗管道
class DataCleaner:
def __init__(self):
self.stats = {"total": 0, "removed": 0, "cleaned": 0}
def clean(self, examples):
results = []
for ex in examples:
self.stats["total"] += 1
# Step 1: Remove duplicates
if self.is_duplicate(ex): continue
# Step 2: Filter too short/long
if not self.valid_length(ex, min_tokens=20, max_tokens=4096): continue
# Step 3: Quality scoring
if self.quality_score(ex) < 0.7: continue
# Step 4: Format validation
ex = self.normalize_format(ex)
results.append(ex)
self.stats["cleaned"] += 1
return results
2. 代幣化深入研究
import tiktoken
enc = tiktoken.encoding_for_model("gpt-4o-mini")
def analyze_dataset(examples):
token_counts = []
for ex in examples:
text = json.dumps(ex, ensure_ascii=False)
tokens = len(enc.encode(text))
token_counts.append(tokens)
print(f"Total examples: {len(token_counts)}")
print(f"Total tokens: {sum(token_counts):,}")
print(f"Avg tokens/example: {sum(token_counts)/len(token_counts):.0f}")
print(f"Training cost (3 epochs, $3/1M): ${sum(token_counts)*3*3/1_000_000:.2f}")
3. 資料增強
- 釋義:重寫問題/答案,保持意義完整
- 語言混合:新增英語-越南語混合的範例
- 邊緣情況:為罕見情況建立範例
- 反面範例:教模型“不回答什麼”
總結
- 資料清理管道:去重→過濾→品質評分→標準化
- 代幣化分析有助於計算準確的成本
- 增強可以增加多樣性,而無需收集更多
- 訓練/驗證/測試分割:80/10/10 或 90/5/5
練習
- 為您的資料集建立資料清理管道
- 分析代幣分佈-找出異常值
- 從 5 個種子範例建立 20 個增強範例
- 進行品質評分並刪除不良範例