Chuyển đến nội dung chính

Lesson 16: Decision Tree, Random Forest, XGBoost

Compare tree-based models, understand feature importance, overfitting control, and how to choose a model according to the data.

🧠 AI & ML — Lesson 15 Lesson 16: Decision Tree, Random Forest, XGBoost

Machine Learning: From Basics to Advanced

Part 3: Advanced algorithms just enough to use

xdev.asia

Introduction

At this point you have a good foundation with linear models. This article introduces the most powerful group of algorithms for panel data in practice: decision tree, random forest and boosting.

Lesson objectives

  • Understand the intuition of decision trees, random forests and boosting.
  • Know the trade-off between explainability, speed and strength.
  • Choose the correct tree-based model for table data.

Decision tree

The decision tree divides the data by questions like age > 35 or num_tickets > 3. The advantage is that it is easy to understand, does not need to scale features and captures nonlinear relationships. The downside is that it's easy to overfit if the tree is too deep.

Random forest

Random forest is many decision trees voting together. It often reduces overfitting compared to a single tree and is a very strong baseline for tabular data.

XGBoost and boosting

Boosting trains multiple trees sequentially; Each new tree focuses on correcting the errors of the previous tree. Therefore, boosting is often very powerful on the leaderboard, but can also be easily abused if validation is not well controlled.

When to use which model?

  • Decision tree: to learn intuitively or need a very easy-to-explain model.
  • Random forest: strong baseline for tabular data.
  • XGBoost, LightGBM, CatBoost: when you want to seriously optimize performance.

Common mistakes

  • Tuning too many parameters from the beginning.
  • Use feature importance as evidence of cause and effect.
  • Trust leaderboard without viewing leakage or error analysis.

Practice exercises

  • Compare Logistic Regression, Random Forest and XGBoost on the same dataset.
  • Record metrics, training time, explainability.
  • Conclude which model is most suitable for the small and medium enterprise environment.

Completion criteria

  • Explain the difference between bagging and boosting.
  • Know when a tree-based model is stronger than a linear model.
  • Compare at least 3 models on the same problem.

Practice step by step (advanced)

  1. Run 3 models: Decision Tree, Random Forest, XGBoost.
  2. Keep the same train/validation split for fair comparison.
  3. Light tuning for each model (2-3 main parameters).
  4. Compare metrics, training time and ease of interpretation.
  5. Write a guideline for selecting models according to data size.

Artifact should be submitted

  • Benchmark table for three models.
  • The feature importance chart has careful annotations.
  • Rules for choosing models according to each project context.

Self-test questions

  • How are Bagging and Boosting different in terms of error reduction mechanism?
  • When is Random Forest better than XGBoost?
  • Why is feature importance not synonymous with cause and effect?