Lesson objectives
- Complete end-to-end ML projects according to clear rubrics
- Present results from a technical and business perspective
- Prepare portfolio ready for interview
Submission checklist
- Problem description, metric and baseline
- Preprocessing pipeline + reproducible training model
- Validation and error analysis results
- Meaningful explainability (SHAP/permutation)
- The inferencing API is active and has run instructions
- Monitoring plan (drift, retraining, alert)
- 1-page report for stakeholders
Capstone scoring rubric (100 points)
- Correct problem definition and metric: 10
- Data quality and preprocessing: 15
- Baseline and systematic improvement: 15
- Model evaluation + CV + tuning: 20
- Error analysis + result interpretation: 15
- Serving + reset + run instructions: 15
- Monitoring/retraining + business reporting: 10
Implementation instructions
- Choose a specific use case (churn, fraud, demand forecasting, pricing).
- Use a simple baseline before using a complex model.
- Only optimize the set metrics from the beginning, don't change the metrics midway.
- Pack all preprocessing into the pipeline to avoid leakage.
- Write a short README: how to train, eval, serve and monitor after deployment.
Expected output
You have a complete project to include in your CV/portfolio and demo during the interview.
Suggested capstone submission route
- Conclusion and measurable success criteria.
- Baseline pin, primary metric set and secondary metric set.
- Complete pipeline + error analysis report.
- Has a minimal inference demo (batch or API).
- Submit the final report according to the 100-point rubric.
Artifact is required
- Dataset card describes data and risks.
- Model card describes the scope of use, limitations and fairness notes.
- Repo has instructions to rerun the entire process.
Self-assessment before submission
- Are the results reproducible on another machine?
- Has the risk of leakage or bias been stated transparently?
- Is there a plan for monitoring if put into production?