Chuyển đến nội dung chính

Lesson 18: PCA, t-SNE, UMAP for visualization

Reduce data dimensionality to understand clusters, detect anomalies, and increase downstream model performance.

🧠 AI & ML — Lesson 17 Lesson 18: PCA, t-SNE, UMAP for visualization

Machine Learning: From Basics to Advanced

Part 3: Advanced algorithms just enough to use

xdev.asia

Introduction

When data has too many dimensions, it is difficult to see, difficult to draw, and sometimes difficult to model. PCA, t-SNE and UMAP help reduce data dimensionality but serve different purposes.

Lesson objectives

  • Distinguish PCA from t-SNE and UMAP.
  • Know when to use it for dimensional compression and when to use it for visualization.
  • Avoid misinterpreting dimensionality reduction plots.

PCA

PCA finds new axes that retain the most variance of the data. This is a fast, linear technique that is relatively easy to explain and useful when it comes to reducing input dimensionality.

t-SNE and UMAP

t-SNE and UMAP are mainly useful for visualizing local structure in high-dimensional data. UMAP is generally faster and quite suitable for embedding; Robust t-SNE for visualizing local clusters.

Interpretation warning

The distance on the 2D chart after dimension reduction does not always reflect the true distance in the original space. Don't use pretty graphs as too strong evidence for a conclusion.

Common mistakes

  • Using t-SNE or UMAP as the main feature input without checking carefully.
  • Interpret the diagram as a true class boundary.
  • Do not scale data before PCA when needed.

Practice exercises

  • Run PCA and t-SNE on the same dataset.
  • Draw 2D diagrams for both.
  • Write comments: which tool is more suitable for visualization, which tool is more suitable for preprocessing.

Completion criteria

  • Distinguish the purposes of PCA, t-SNE, UMAP.
  • Do not over-interpret the dimensionality reduction diagram.
  • Know how to choose tools for the right purpose.

Practice step by step (advanced)

  1. Run PCA retaining 90% of the variance.
  2. Data visualization using 2D PCA.
  3. Run t-SNE and UMAP on the same input embedding.
  4. Comparison of runtime and cluster image stability.
  5. Write an explanatory warning for dashboard viewers.

Artifact should be submitted

  • Visual set from 3 techniques.
  • Table comparing the goals of using each technique.
  • Instructions for choosing dimension reduction tools according to use-case.

Self-test questions

  • What advantages does PCA have in terms of reproducibility?
  • Why is t-SNE not suitable as the main feature?
  • When is UMAP preferable to t-SNE?