1. 機器學習平台概述
Raw Data -> Feature Pipeline -> Feature Store
│
├-> Training Jobs -> Model Registry
│
└-> Online Features -> Model Serving -> Predictions
2. 特徵庫設計
interface FeatureDefinition {
name: string;
entity: 'user' | 'product' | 'shop';
type: 'float' | 'int' | 'string' | 'vector';
source: string;
ttlHours?: number;
}
const features: FeatureDefinition[] = [
{ name: 'user_30d_click_count', entity: 'user', type: 'int', source: 'events' },
{ name: 'product_ctr_7d', entity: 'product', type: 'float', source: 'analytics' },
{ name: 'product_clip_embedding', entity: 'product', type: 'vector', source: 'ai' },
{ name: 'shop_return_rate_30d', entity: 'shop', type: 'float', source: 'orders' },
];
- 線下商店:訓練資料集
- 網上商店:低延遲推理功能
- 時間點正確性:避免資料外洩
3. 培訓流程
Schedule trigger (daily/weekly)
-> Build training dataset
-> Train model
-> Evaluate metrics
-> Register candidate model
-> Optional shadow deployment
class TrainingOrchestrator {
async run(job: TrainingJob) {
const dataset = await this.datasetBuilder.build(job.featureSet, job.timeWindow);
const model = await this.trainer.train(job.algorithm, dataset);
const metrics = await this.evaluator.evaluate(model, dataset.validation);
await this.mlflow.logRun({ job, metrics });
if (metrics.auc >= job.minAuc && metrics.calibrationError <= job.maxCalibrationError) {
await this.registry.register(model, metrics);
}
}
}
4. 模型服務
| 圖案 | 使用時 |
|---|---|
| 在線推理 | 推薦、即時個人化 |
| 批量推理 | 每晚排名預計算,趨勢預測 |
| 流式推理 | 按事件進行詐欺/風險評分 |
interface PredictionRequest {
modelName: string;
entityId: string;
features: Record<string, unknown>;
}
class ModelServingGateway {
async predict(req: PredictionRequest) {
const onlineFeatures = await this.featureStore.getOnline(req.entityId);
const merged = { ...onlineFeatures, ...req.features };
return this.runtime.predict(req.modelName, merged);
}
}
5.A/B測試框架
interface Experiment {
id: string;
name: string;
variants: Array<{ name: 'control' | 'treatment'; weight: number }>;
primaryMetric: 'ctr' | 'conversion' | 'revenue_per_session';
guardrails: string[];
}
function assignVariant(userId: string, experimentId: string): string {
const bucket = hash(userId + experimentId) % 100;
return bucket < 50 ? 'control' : 'treatment';
}
- 在運行測試之前主要指標已經明確
- Guardrail:延遲、錯誤率、退款率
- 停止標準:意義+實際影響
6. 監控和漂移檢測
Monitors:
- Data drift: PSI / KS distance
- Prediction drift: distribution shift
- Performance drift: CTR/conversion decay
- Operational: latency/error/timeout
if (psi(featureDistTrain, featureDistLive) > 0.2) {
alert('Feature drift high');
triggerRetraining('recommendation_model');
}
7. MLOps 治理
- 具有版本控制+批准工作流程的模型註冊表
- 實驗追蹤(MLflow)
- 可重複的訓練(代碼+資料快照)
- 模型降級時的回滾策略
八、總結
特色商店 是同步訓練和服務的中心
培訓管道 推廣模型之前需要品質標準
A/B 測試 是安全生產決策機制
漂移監測 幫助及早發現模型惡化
MLOps 治理 確保機器學習能夠持續運作