Chuyển đến nội dung chính

Lesson 19: Anomaly Detection in real systems

Isolation Forest, One-Class SVM and design warning rules for fraud, log monitoring, quality control.

🧠 AI & ML — Lesson 18 Lesson 19: Anomaly Detection in the system really

Machine Learning: From Basics to Advanced

Part 3: Advanced algorithms just enough to use

xdev.asia

Introduction

There are problems where positive classes are so rare that there are almost not enough labels to train a standard classification, for example detecting fraud, operational abnormalities, sensor errors. Then anomaly detection is a direction worth considering.

Lesson objectives

  • Understand how anomaly detection differs from classification.
  • Know some introductory techniques like Isolation Forest.
  • Know how to evaluate abnormalities in the business context.

Core intuition

An outlier is one that differs from the rest of the data in a meaningful way. The point is that different is not always bad. Therefore, anomaly detection always needs to be tied to the operational context.

Isolation Forest

Intuitive idea: outliers are often isolated faster in random splits of the tree. The easier it is to separate a point, the more likely it is to be an anomaly.

Evaluate the model

You may need a set of limited labels, review the top alarms with a business expert, and measure the cost of false alarms versus the cost of missing them.

Common mistakes

  • Call every outlier an important anomaly.
  • Not confirmed with domain expert.
  • Use threshold arbitrarily without considering the operational impact.

Practice exercises

  • Run Isolation Forest on a transaction dataset.
  • Take the top 20 most unusual points to review.
  • Write comments: which warnings are reasonable, which warnings could be false alarms.

Completion criteria

  • Understanding anomaly detection is not just about finding outliers with your eyes.
  • Know how to use a basic model like Isolation Forest.
  • Link the assessment with actual operating costs.

Practice step by step (advanced)

  1. Clearly define what an anomaly is in a specific problem.
  2. Run Isolation Forest with several levels of contamination.
  3. Check the top abnormal scores by manual review.
  4. Compare false alarms between configurations.
  5. Propose practical operating warning thresholds.

Artifact should be submitted

  • List of top anomalies with explanations.
  • Operational impact table of false positives.
  • Proposed warning triage procedure.

Self-test questions

  • What is the difference between statistical outlier and business anomaly?
  • Why does contamination need to be adjusted according to context?
  • When is human-in-the-loop needed for anomaly review?