Chuyển đến nội dung chính

Lesson 8: Important metrics: Accuracy, Precision, Recall, F1, AUC

Choose the right metric according to the business problem; when to use PR-AUC instead of ROC-AUC; Avoid optimizing for the wrong goal.

🧠 AI & ML — Lesson 7 Lesson 8: Important Metrics: Accuracy, Precision, Recall, F1, AUC

Machine Learning: From Basics to Advanced

Part 1: Supervised Learning foundation

xdev.asia

Introduction

Whether the model is good or not depends a lot on the metric you choose. The same model can look very good in terms of accuracy but very bad in terms of recall. This article helps you choose metrics according to the right problem instead of out of habit.

Lesson objectives

  • Distinguishing metrics for regression and classification.
  • Understand accuracy, precision, recall, F1, ROC-AUC and when to use them.
  • Know why metrics must stick to product goals.

Regression metrics

  • MAE: easy to understand, measures the average absolute error.
  • RMSE: penalty for serious errors is greater than MAE.
  • R-squared: measures the amount of variance explained, but don't sanctify it.

Classification metrics

  • Accuracy: overall correct prediction rate.
  • Precision: among the samples predicted to be positive, how many are actually positive.
  • Recall: among the true positive samples, how many can the model capture?
  • F1-score: balance between precision and recall.
  • ROC-AUC: measures the ability to rank the probability between two classes.

Example of selecting metrics according to context

  • Detect fraud: prioritize high recall so you don't miss it.
  • Spam emails: need to be accurate enough to avoid mistakenly blocking regular emails.
  • Customer churn: often cares about recall and the business value of retention actions.

Sample code

from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score

print('Accuracy:', accuracy_score(y_test, preds))
print('Precision:', precision_score(y_test, preds))
print('Recall:', recall_score(y_test, preds))
print('F1:', f1_score(y_test, preds))
print('ROC-AUC:', roc_auc_score(y_test, proba))

Common mistakes

  • Use accuracy for strongly imbalanced data.
  • Compare models using different metrics between tests.
  • Optimize technical metrics but forget business costs.

Practice exercises

  • Take an imbalanced classification problem.
  • Calculate confusion matrix and 5 main metrics.
  • Write a short paragraph: if you were a PM, which metric would you choose as the main KPI?

Completion criteria

  • Choose the appropriate metric for at least 3 different problems.
  • Explain precision and recall with real-life examples.
  • Can read the confusion matrix without confusion.

Practice step by step (advanced)

  1. Choose a churn or fraud problem with class imbalance.
  2. Measure 5 metrics: Accuracy, Precision, Recall, F1, ROC-AUC.
  3. Add PR-AUC for comparison in class difference data.
  4. Simulate FP/FN cost and calculate expected cost.
  5. Key metrics used for weekly reporting.

Artifact should be submitted

  • Metric table + PR curve and ROC curve graphs.
  • Simple cost matrix model.
  • Conclusion of main metrics and secondary metrics for monitoring.

Self-test questions

  • Why is accuracy misleading on out-of-class data?
  • How does PR-AUC differ from ROC-AUC in terms of meaning?
  • When do you need to monitor multiple metrics at the same time?