Chuyển đến nội dung chính

第 21 課:ML 管道和特徵儲存 — 訓練、服務和 A/B 測試

Fashion POD 的 ML 平台 — 特徵儲存、訓練管道、模型服務、A/B 測試、模型監控、偏差檢測、MLOps 堆疊(MLflow、註冊表、實驗追蹤)。

🏗️ 建築 — 第 21 課 第 21 課:ML 管道和特徵儲存 — 培訓、服務和 A/B 測試

時裝設計與按需印刷系統架構-從領域分析到生產

第 6 部分:資料平台與分析

亞洲開發網

1. 機器學習平台概述

Raw Data -> Feature Pipeline -> Feature Store
                     │
                     ├-> Training Jobs -> Model Registry
                     │
                     └-> Online Features -> Model Serving -> Predictions

2. 特徵庫設計

interface FeatureDefinition {
  name: string;
  entity: 'user' | 'product' | 'shop';
  type: 'float' | 'int' | 'string' | 'vector';
  source: string;
  ttlHours?: number;
}

const features: FeatureDefinition[] = [
  { name: 'user_30d_click_count', entity: 'user', type: 'int', source: 'events' },
  { name: 'product_ctr_7d', entity: 'product', type: 'float', source: 'analytics' },
  { name: 'product_clip_embedding', entity: 'product', type: 'vector', source: 'ai' },
  { name: 'shop_return_rate_30d', entity: 'shop', type: 'float', source: 'orders' },
];
  • 線下商店:訓練資料集
  • 網上商店:低延遲推理功能
  • 時間點正確性:避免資料外洩

3. 培訓流程

Schedule trigger (daily/weekly)
  -> Build training dataset
  -> Train model
  -> Evaluate metrics
  -> Register candidate model
  -> Optional shadow deployment
class TrainingOrchestrator {
  async run(job: TrainingJob) {
    const dataset = await this.datasetBuilder.build(job.featureSet, job.timeWindow);
    const model = await this.trainer.train(job.algorithm, dataset);
    const metrics = await this.evaluator.evaluate(model, dataset.validation);

    await this.mlflow.logRun({ job, metrics });

    if (metrics.auc >= job.minAuc && metrics.calibrationError <= job.maxCalibrationError) {
      await this.registry.register(model, metrics);
    }
  }
}

4. 模型服務

圖案使用時
在線推理推薦、即時個人化
批量推理每晚排名預計算,趨勢預測
流式推理按事件進行詐欺/風險評分
interface PredictionRequest {
  modelName: string;
  entityId: string;
  features: Record<string, unknown>;
}

class ModelServingGateway {
  async predict(req: PredictionRequest) {
    const onlineFeatures = await this.featureStore.getOnline(req.entityId);
    const merged = { ...onlineFeatures, ...req.features };
    return this.runtime.predict(req.modelName, merged);
  }
}

5.A/B測試框架

interface Experiment {
  id: string;
  name: string;
  variants: Array<{ name: 'control' | 'treatment'; weight: number }>;
  primaryMetric: 'ctr' | 'conversion' | 'revenue_per_session';
  guardrails: string[];
}

function assignVariant(userId: string, experimentId: string): string {
  const bucket = hash(userId + experimentId) % 100;
  return bucket < 50 ? 'control' : 'treatment';
}
  • 在運行測試之前主要指標已經明確
  • Guardrail:延遲、錯誤率、退款率
  • 停止標準:意義+實際影響

6. 監控和漂移檢測

Monitors:
- Data drift: PSI / KS distance
- Prediction drift: distribution shift
- Performance drift: CTR/conversion decay
- Operational: latency/error/timeout
if (psi(featureDistTrain, featureDistLive) > 0.2) {
  alert('Feature drift high');
  triggerRetraining('recommendation_model');
}

7. MLOps 治理

  • 具有版本控制+批准工作流程的模型註冊表
  • 實驗追蹤(MLflow)
  • 可重複的訓練(代碼+資料快照)
  • 模型降級時的回滾策略

八、總結

  • 特色商店 是同步訓練和服務的中心

  • 培訓管道 推廣模型之前需要品質標準

  • A/B 測試 是安全生產決策機制

  • 漂移監測 幫助及早發現模型惡化

  • MLOps 治理 確保機器學習能夠持續運作