Introduction
The model after deployment does not stand still. User data changes, market behavior changes, products change. Therefore, an ML system has a responsibility to monitor quality after deployment instead of considering training as finished.
Lesson objectives
- Understand prediction drift, data drift and concept drift.
- Know the minimum signals that need to be monitored.
- Design a basic retraining loop.
Things to keep track of
- Distributing input features.
- Prediction rate for each class.
- Model quality when there is a feedback label.
- Latency, request errors and service availability.
Types of drifting
- Data drift: input distribution changes.
- Concept drift: the relationship between input and target changes.
- Prediction drift: model output changes abnormally.
When to retrain?
Not every time you see a drift, you immediately retrain. You need to answer whether drifting really affects performance, whether there is enough new reliable data to retrain, and whether retraining requires human review before release.
Minimum monitoring for newbies
- Dashboard tracks number of requests and errors.
- Distribution chart of some important features.
- Track key metrics by week or month.
- Warning when the distribution deviates beyond the threshold.
Common mistakes
- Only monitors infrastructure but not model quality.
- Automatic retrain does not check for regression.
- Do not version data and models.
Practice exercises
- Design a monitoring checklist for model churn or housing.
- Identify 5 indicators that must be monitored.
- Write a simple retrain policy: when to retrain, who approves, how to rollback.
Completion criteria
- Distinguish between data drift and concept drift.
- Recommended minimum set of monitor indicators.
- Have a basic retrain and rollback plan.
Practice step by step (advanced)
- Select a set of monitoring indicators for quality and performance.
- Set data drift warning thresholds for 5 main features.
- Simulate a drift scenario and observe the warning.
- Design a retrain process with an approval step.
- Write a rollback playbook when the new model is inferior.
Artifact should be submitted
- Monitoring checklist weekly.
- Retraining and release model process.
- Incident response playbook for the model.
Self-test questions
- What is the difference between data drift and concept drift?
- When should you retrain periodically, when should you retrain according to events?
- If drift increases but metric has not decreased, what action should be taken first?