Chuyển đến nội dung chính

Lesson 7: Logistic Regression & probability for classification

Logistic regression, sigmoid, decision boundary, threshold and how to properly read prediction probabilities.

🧠 AI & ML — Lesson 6 Lesson 7: Logistic Regression & probability for classification

Machine Learning: From Basics to Advanced

Part 1: Supervised Learning foundation

xdev.asia

Introduction

Regression predicts a continuous number, while classification predicts a label. Logistic regression is the best introductory classification model because it is simple, fast, easy to explain, and still very useful in practice.

Lesson objectives

  • Understand how logistic regression differs from linear regression.
  • Read the model's output probability.
  • Know how to use threshold to turn probabilities into prediction labels.

From straight lines to probabilities

Linear regression can produce any value from negative infinity to positive infinity. Binary classification requires probabilities between 0 and 1. Logistic regression solves this problem with the sigmoid function:

$$ \sigma(z) = \frac{1}{1 + e^{-z}} $$

Where $z = w_1x_1 + ... + w_nx_n + b$.

Threshold is not a fixed truth

Many new users default threshold = 0.5. This is only true when the costs of false positive and false negative are equal.

For example, disease prediction often prioritizes recall; Customer churn may accept redundant calls rather than missing out on customers who are about to leave.

Code example

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report

model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
proba = model.predict_proba(X_test)[:, 1]
preds = (proba >= 0.5).astype(int)

print(classification_report(y_test, preds))

Advantages and disadvantages

Advantages: fast, easy to baseline, easy to explain.

Disadvantages: ineffective when class boundaries are too nonlinear, sensitive to feature engineering and scale.

Common mistakes

  • Only look at accuracy when the data is out of class.
  • Do not test different thresholds.
  • Causes leakage in the encoding or scaling step.

Practice exercises

  • Train logistic regression for a churn or spam problem.
  • Compare results at thresholds 0.3, 0.5 and 0.7.
  • Write comments: which threshold suits the problem better and why.

Completion criteria

  • Explain the role of sigmoid.
  • Understand the probability that the output differs from the predicted label.
  • Know how to change threshold according to business goals.

Practice step by step (advanced)

  1. Use dataset classification with slight class bias.
  2. Train Logistic Regression with default class_weight and balanced.
  3. Compare precision, recall, F1 at thresholds 0.3, 0.5, 0.7.
  4. Draw precision-recall tradeoff.
  5. Choose threshold according to a specific product scenario.

Artifact should be submitted

  • Metric table according to threshold.
  • A confusion matrix with clearly annotated FP/FN.
  • 1 page explanation sent to PM for simulation.

Self-test questions

  • Why is threshold 0.5 not always correct?
  • In what cases should recall be prioritized over precision?
  • High ROC-AUC but still bad business outcomes, what could be the reason?