はじめに
生データを微調整できる状態になることはほとんどありません。この記事では、プロフェッショナルな データ処理パイプラインを構築します。
1. データ クリーニング パイプライン
class DataCleaner:
def __init__(self):
self.stats = {"total": 0, "removed": 0, "cleaned": 0}
def clean(self, examples):
results = []
for ex in examples:
self.stats["total"] += 1
# Step 1: Remove duplicates
if self.is_duplicate(ex): continue
# Step 2: Filter too short/long
if not self.valid_length(ex, min_tokens=20, max_tokens=4096): continue
# Step 3: Quality scoring
if self.quality_score(ex) < 0.7: continue
# Step 4: Format validation
ex = self.normalize_format(ex)
results.append(ex)
self.stats["cleaned"] += 1
return results
2. トークン化の詳細
import tiktoken
enc = tiktoken.encoding_for_model("gpt-4o-mini")
def analyze_dataset(examples):
token_counts = []
for ex in examples:
text = json.dumps(ex, ensure_ascii=False)
tokens = len(enc.encode(text))
token_counts.append(tokens)
print(f"Total examples: {len(token_counts)}")
print(f"Total tokens: {sum(token_counts):,}")
print(f"Avg tokens/example: {sum(token_counts)/len(token_counts):.0f}")
print(f"Training cost (3 epochs, $3/1M): ${sum(token_counts)*3*3/1_000_000:.2f}")
3. データの拡張
- 言い換え: 意味を保ったまま質問/回答を書き直します。
- 言語の混合: 英語とベトナム語の混合の例を追加
- 特殊なケース: まれな状況の例を作成します
- 否定的な例: モデルに「答えてはいけないもの」を教える
概要
- データ クリーニング パイプライン: 重複排除 → フィルター → 品質スコア → 正規化
- トークン化分析は正確なコストの計算に役立ちます
- 拡張により、さらに収集する必要なく多様性が増加します
- トレーニング/ヴァル/テスト分割: 80/10/10 または 90/5/5
演習
- データセットのデータ クリーニング パイプラインを構築する
- トークンの分布を分析する — 外れ値を見つける
- 5 つのシード サンプルから 20 の拡張サンプルを作成する
- 品質スコアリングを実行し、不適切な例を削除する