Chuyển đến nội dung chính

第 5 課:資料清理與增強-從“垃圾”到“黃金”

資料清理管道:重複資料刪除、篩選、品質評分。數據增強技術。處理不平衡和邊緣情況。代幣化深入研究。訓練/驗證/測試分割。

🧠 人工智慧與機器學習 — 第 4 課 第 5 課:資料清理與增強 — 來自 “垃圾”變成“黃金”

微調 LLM:AI 調優的藝術

第 2 部分:資料準備 — 所有成功的基礎

亞洲開發網

簡介

原始數據很少可以用於微調。本文建構了一個專業的資料處理管道。


1. 資料清洗管道

class DataCleaner:
    def __init__(self):
        self.stats = {"total": 0, "removed": 0, "cleaned": 0}
    
    def clean(self, examples):
        results = []
        for ex in examples:
            self.stats["total"] += 1
            
            # Step 1: Remove duplicates
            if self.is_duplicate(ex): continue
            
            # Step 2: Filter too short/long
            if not self.valid_length(ex, min_tokens=20, max_tokens=4096): continue
            
            # Step 3: Quality scoring
            if self.quality_score(ex) < 0.7: continue
            
            # Step 4: Format validation
            ex = self.normalize_format(ex)
            
            results.append(ex)
            self.stats["cleaned"] += 1
        
        return results

2. 代幣化深入研究

import tiktoken

enc = tiktoken.encoding_for_model("gpt-4o-mini")

def analyze_dataset(examples):
    token_counts = []
    for ex in examples:
        text = json.dumps(ex, ensure_ascii=False)
        tokens = len(enc.encode(text))
        token_counts.append(tokens)
    
    print(f"Total examples: {len(token_counts)}")
    print(f"Total tokens: {sum(token_counts):,}")
    print(f"Avg tokens/example: {sum(token_counts)/len(token_counts):.0f}")
    print(f"Training cost (3 epochs, $3/1M): ${sum(token_counts)*3*3/1_000_000:.2f}")

3. 資料增強

  • 釋義:重寫問題/答案,保持意義完整
  • 語言混合:新增英語-越南語混合的範例
  • 邊緣情況:為罕見情況建立範例
  • 反面範例:教模型“不回答什麼”

總結

  • 資料清理管道:去重→過濾→品質評分→標準化
  • 代幣化分析有助於計算準確的成本
  • 增強可以增加多樣性,而無需收集更多
  • 訓練/驗證/測試分割:80/10/10 或 90/5/5

練習

  1. 為您的資料集建立資料清理管道
  2. 分析代幣分佈-找出異常值
  3. 從 5 個種子範例建立 20 個增強範例
  4. 進行品質評分並刪除不良範例