Introduction
Regression predicts a continuous number, while classification predicts a label. Logistic regression is the best introductory classification model because it is simple, fast, easy to explain, and still very useful in practice.
Lesson objectives
- Understand how logistic regression differs from linear regression.
- Read the model's output probability.
- Know how to use threshold to turn probabilities into prediction labels.
From straight lines to probabilities
Linear regression can produce any value from negative infinity to positive infinity. Binary classification requires probabilities between 0 and 1. Logistic regression solves this problem with the sigmoid function:
$$ \sigma(z) = \frac{1}{1 + e^{-z}} $$
Where $z = w_1x_1 + ... + w_nx_n + b$.
Threshold is not a fixed truth
Many new users default threshold = 0.5. This is only true when the costs of false positive and false negative are equal.
For example, disease prediction often prioritizes recall; Customer churn may accept redundant calls rather than missing out on customers who are about to leave.
Code example
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report
model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
proba = model.predict_proba(X_test)[:, 1]
preds = (proba >= 0.5).astype(int)
print(classification_report(y_test, preds))
Advantages and disadvantages
Advantages: fast, easy to baseline, easy to explain.
Disadvantages: ineffective when class boundaries are too nonlinear, sensitive to feature engineering and scale.
Common mistakes
- Only look at accuracy when the data is out of class.
- Do not test different thresholds.
- Causes leakage in the encoding or scaling step.
Practice exercises
- Train logistic regression for a churn or spam problem.
- Compare results at thresholds 0.3, 0.5 and 0.7.
- Write comments: which threshold suits the problem better and why.
Completion criteria
- Explain the role of sigmoid.
- Understand the probability that the output differs from the predicted label.
- Know how to change threshold according to business goals.
Practice step by step (advanced)
- Use dataset classification with slight class bias.
- Train Logistic Regression with default class_weight and balanced.
- Compare precision, recall, F1 at thresholds 0.3, 0.5, 0.7.
- Draw precision-recall tradeoff.
- Choose threshold according to a specific product scenario.
Artifact should be submitted
- Metric table according to threshold.
- A confusion matrix with clearly annotated FP/FN.
- 1 page explanation sent to PM for simulation.
Self-test questions
- Why is threshold 0.5 not always correct?
- In what cases should recall be prioritized over precision?
- High ROC-AUC but still bad business outcomes, what could be the reason?