Introduction
This is the first mini-project classification of the series. The goal is not just to train the churn model, but to know how to ask product questions: predict who is leaving for what, what actions will be triggered, and what metrics are really worth paying attention to.
Lesson objectives
- Make a relatively complete classification workflow.
- Build baseline, choose metrics, try multiple thresholds.
- Present results from a business perspective.
Problem context
A subscription company wants to predict which customers are at risk of leaving in the next 30 days. If you know early, the CS or marketing team can send retention incentives.
Recommended procedure
- Read and check the data.
- See if the churn rate is out of class.
- Create a simple baseline.
- Train logistic regression first.
- Evaluation by confusion matrix, precision, recall, F1, ROC-AUC.
- Test threshold according to business goals.
Feature suggestion
- tenure
- monthly_charges
- contract_type
- support_tickets
- payment_method
- is_auto_renew
Frame code
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
How to present to stakeholders
Don't just say "model achieves F1 = 0.71". Let's say what percentage of customers about to churn is caught by the model, how many customers are falsely warned, and whether the retention cost is reasonable.
Common mistakes
- Accuracy indicator when data is out of class.
- Does not specify the threshold being used.
- Using features that arise after the churn time, causing leakage.
Practice exercises
- Make a complete notebook for churn prediction.
- Choose 2 different thresholds and compare hypothetical costs/benefits.
- Write the conclusion as a short email to PM.
Completion criteria
- Has clear baseline, main model and metrics.
- There is an explanation about threshold.
- Conclude from a business perspective, not just a technical one.
Practice step by step (advanced)
- Design the churn target variable with a clear forecasting window time.
- Build baseline according to business rules (rule-based).
- Train at least 2 models: Logistic + Tree-based.
- Optimize threshold according to assumed retention costs.
- Write a short executive summary for the business team.
Artifact should be submitted
- End-to-end notebook with pipeline.
- Markdown file summarizes cost assumptions.
- An action list table: which group needs to intervene first.
Self-test questions
- At what step can the churn label leak?
- Why do we need a baseline rule-based before the ML model?
- How to choose which threshold is closer to the profit target?