Chuyển đến nội dung chính

Lesson 17: Clustering (K-Means, DBSCAN, Hierarchical)

Unsupervised learning for customer segmentation and data structure discovery in the absence of labels.

🧠 AI & ML — Lesson 16 Lesson 17: Clustering (K-Means, DBSCAN, Hierarchical)

Machine Learning: From Basics to Advanced

Part 3: Advanced algorithms just enough to use

xdev.asia

Introduction

You don't always have labels. Clustering is a group of techniques that help find hidden structures in unlabeled data. In practice, it is often used for customer segmentation, behavioral grouping, and data discovery.

Lesson objectives

  • Understand how clustering is different from supervised learning.
  • Know how to use K-Means at a basic level.
  • Understand the limitations of clustering evaluation.

How does K-Means work?

  1. Choose number of clusters k.
  2. Assign each point to the nearest cluster center.
  3. Update the cluster center again.
  4. Repeat until stable.

Appropriate practical problem

  • Grouping customers according to purchasing behavior.
  • Group posts or users by embedding.
  • Create segments for business teams to act differently.

How to choose the number of clusters?

There is no absolute answer. Can refer to elbow method, silhouette score and most importantly, business interpretation.

Common mistakes

  • Think that the generated cluster is always a natural truth.
  • Choose k just because the chart looks nice.
  • Do not check the business significance of the cluster.

Practice exercises

  • Run K-Means on a dataset segmentation.
  • Describe each cluster in business language.
  • Suggest a different action for each cluster.

Completion criteria

  • Understand clustering without labels.
  • Can run K-Means with scaled data.
  • Interpret the cluster from a business perspective.

Practice step by step (advanced)

  1. Normalize features before clustering.
  2. Run K-Means with multiple k values.
  3. Evaluate using silhouette score and business interpretation.
  4. Try adding DBSCAN or Hierarchical for comparison.
  5. Name the cluster according to business language.

Artifact should be submitted

  • Comparison table of clustering algorithms.
  • Profile describing each cluster (cluster profile).
  • List of recommended actions for each cluster.

Self-test questions

  • Why doesn't clustering have an absolutely correct answer?
  • When is DBSCAN more beneficial than K-Means?
  • How to evaluate whether a cluster is useful for business or not?