1. ML Platform Overview
Raw Data -> Feature Pipeline -> Feature Store
│
├-> Training Jobs -> Model Registry
│
└-> Online Features -> Model Serving -> Predictions
2. Feature Store Design
interface FeatureDefinition {
name: string;
entity: 'user' | 'product' | 'shop';
type: 'float' | 'int' | 'string' | 'vector';
source: string;
ttlHours?: number;
}
const features: FeatureDefinition[] = [
{ name: 'user_30d_click_count', entity: 'user', type: 'int', source: 'events' },
{ name: 'product_ctr_7d', entity: 'product', type: 'float', source: 'analytics' },
{ name: 'product_clip_embedding', entity: 'product', type: 'vector', source: 'ai' },
{ name: 'shop_return_rate_30d', entity: 'shop', type: 'float', source: 'orders' },
];
- Offline store: training datasets
- Online store: low-latency inference features
- Point-in-time correctness: tránh data leakage
3. Training Pipeline
Schedule trigger (daily/weekly)
-> Build training dataset
-> Train model
-> Evaluate metrics
-> Register candidate model
-> Optional shadow deployment
class TrainingOrchestrator {
async run(job: TrainingJob) {
const dataset = await this.datasetBuilder.build(job.featureSet, job.timeWindow);
const model = await this.trainer.train(job.algorithm, dataset);
const metrics = await this.evaluator.evaluate(model, dataset.validation);
await this.mlflow.logRun({ job, metrics });
if (metrics.auc >= job.minAuc && metrics.calibrationError <= job.maxCalibrationError) {
await this.registry.register(model, metrics);
}
}
}
4. Model Serving
| Pattern | Khi dùng |
|---|---|
| Online inference | Recommendation, personalization thời gian thực |
| Batch inference | Nightly ranking precompute, trend prediction |
| Streaming inference | Fraud/risk scoring theo event |
interface PredictionRequest {
modelName: string;
entityId: string;
features: Record<string, unknown>;
}
class ModelServingGateway {
async predict(req: PredictionRequest) {
const onlineFeatures = await this.featureStore.getOnline(req.entityId);
const merged = { ...onlineFeatures, ...req.features };
return this.runtime.predict(req.modelName, merged);
}
}
5. A/B Testing Framework
interface Experiment {
id: string;
name: string;
variants: Array<{ name: 'control' | 'treatment'; weight: number }>;
primaryMetric: 'ctr' | 'conversion' | 'revenue_per_session';
guardrails: string[];
}
function assignVariant(userId: string, experimentId: string): string {
const bucket = hash(userId + experimentId) % 100;
return bucket < 50 ? 'control' : 'treatment';
}
- Primary metric rõ ràng trước khi chạy test
- Guardrail: latency, error rate, refund rate
- Stop criteria: significance + practical impact
6. Monitoring & Drift Detection
Monitors:
- Data drift: PSI / KS distance
- Prediction drift: distribution shift
- Performance drift: CTR/conversion decay
- Operational: latency/error/timeout
if (psi(featureDistTrain, featureDistLive) > 0.2) {
alert('Feature drift high');
triggerRetraining('recommendation_model');
}
7. MLOps Governance
- Model registry with versioning + approval workflow
- Experiment tracking (MLflow)
- Reproducible training (code + data snapshot)
- Rollback strategy khi model degrade
8. Tổng kết
Feature store là trung tâm để đồng bộ training và serving
Training pipeline cần tiêu chí chất lượng trước khi promote model
A/B testing là cơ chế ra quyết định production an toàn
Drift monitoring giúp phát hiện suy giảm model sớm
MLOps governance đảm bảo ML có thể vận hành bền vững