Chuyển đến nội dung chính

Bài 1: AI, ML & Deep Learning — Concepts and Terminology

AI vs ML vs DL. Supervised, Unsupervised, Reinforcement Learning. Classification, Regression, Clustering. Neural Networks basics. Training, Validation, Test sets. Bias-Variance tradeoff.

AI, ML và Deep Learning Hierarchy

AI, ML và Deep Learning — quan hệ lồng nhau và ba paradigm học máy

Tổng quan Domain 1

Domain 1 chiếm 20% đề thi AIF-C01. Bạn cần hiểu rõ các khái niệm nền tảng về AI, ML, và Deep Learning — không cần code, nhưng phải phân biệt được khi nào dùng approach nào.

Exam tip: Domain này thường có các câu hỏi dạng "Which type of machine learning is BEST suited for..." — yêu cầu bạn chọn đúng paradigm cho use case.

1. AI vs Machine Learning vs Deep Learning

Ba khái niệm này có quan hệ lồng nhau (nested relationship):

┌─────────────────────────────────────────────┐
│  Artificial Intelligence (AI)               │
│  "Machines that mimic human intelligence"   │
│  ┌───────────────────────────────────────┐   │
│  │  Machine Learning (ML)               │   │
│  │  "Learning from data without         │   │
│  │   explicit programming"              │   │
│  │  ┌─────────────────────────────────┐  │   │
│  │  │  Deep Learning (DL)             │  │   │
│  │  │  "Neural networks with many     │  │   │
│  │  │   layers"                       │  │   │
│  │  └─────────────────────────────────┘  │   │
│  └───────────────────────────────────────┘   │
└─────────────────────────────────────────────┘
ConceptDefinitionExample
AIBroad field — machines performing tasks that typically require human intelligenceChatbot, self-driving car, chess engine
MLSubset of AI — algorithms learn patterns from dataSpam filter, recommendation engine
DLSubset of ML — neural networks with multiple layersImage recognition, language translation

Key Differences for the Exam

  • Traditional Programming: Rules + Data → Output
  • Machine Learning: Data + Output → Rules (model learns the rules)
  • Deep Learning: Tự động extract features từ raw data (không cần manual feature engineering)

2. Three ML Paradigms

2.1. Supervised Learning

Model học từ labeled data — mỗi input đi kèm output đúng (label/target).

Task TypeOutputUse CaseAlgorithms
ClassificationDiscrete categorySpam vs Not Spam, Fraud detectionLogistic Regression, Random Forest, SVM
RegressionContinuous numberHouse price prediction, Stock forecastLinear Regression, XGBoost

Exam tip: Nếu đề bài nói "predict a category" hoặc "classify" → Classification. Nếu nói "predict a number/value" → Regression.

2.2. Unsupervised Learning

Model học từ unlabeled data — tự tìm patterns, structure trong dữ liệu.

Task TypeWhat it doesUse Case
ClusteringGroup similar data pointsCustomer segmentation, Document grouping
Dimensionality ReductionReduce features while preserving infoData visualization, noise reduction
Anomaly DetectionFind unusual data pointsFraud detection, equipment failure
AssociationFind rules between items"Customers who bought X also bought Y"

2.3. Reinforcement Learning (RL)

Agent học bằng cách trial-and-error trong một environment. Nhận reward (positive) hoặc penalty (negative) cho mỗi action.

Agent → Action → Environment → State + Reward → Agent (loop)

Use cases:

  • Game AI (AlphaGo)
  • Robotics navigation
  • Autonomous driving
  • AWS DeepRacer (self-driving car simulation)

2.4. Choosing the Right Paradigm — Exam Decision Tree

Do you have labeled data?
├── YES → Supervised Learning
│   ├── Predicting a category? → Classification
│   └── Predicting a number? → Regression
├── NO →
│   ├── Want to find groups/patterns? → Unsupervised (Clustering)
│   └── Learning through trial & error? → Reinforcement Learning

3. Data Concepts for ML

3.1. Data Types

TypeDescriptionExamples
StructuredOrganized in rows & columns (tabular)CSV, database tables, spreadsheets
Semi-structuredHas some organization but flexibleJSON, XML, log files
UnstructuredNo predefined formatImages, videos, audio, free text
Time-seriesData points indexed by timeStock prices, IoT sensor readings

3.2. Labeled vs Unlabeled Data

  • Labeled data: Mỗi data point có kèm answer (label). Ví dụ: email + tag "spam"/"not spam". Dùng cho Supervised Learning.
  • Unlabeled data: Chỉ có data, không có label. Dùng cho Unsupervised Learning.
  • Amazon SageMaker Ground Truth: Dịch vụ AWS giúp label data (human + ML-assisted labeling).

3.3. Training, Validation, Test Sets

┌────────────────────────────────────────────────┐
│              Full Dataset (100%)               │
├──────────────────┬──────────┬──────────────────┤
│  Training (70%)  │ Val(15%) │   Test (15%)     │
│  Model learns    │ Tune     │ Final evaluation │
│  from this data  │ hyper-   │ (never seen      │
│                  │ params   │  during training) │
└──────────────────┴──────────┴──────────────────┘
  • Training set: Model học patterns từ đây
  • Validation set: Tune hyperparameters, chống overfitting
  • Test set: Đánh giá cuối cùng — model chưa bao giờ thấy data này

4. Neural Networks Basics

4.1. Architecture

Input Layer → Hidden Layer(s) → Output Layer
    x₁ ──┐     ┌── h₁ ──┐
    x₂ ──┼─────┼── h₂ ──┼──── ŷ (prediction)
    x₃ ──┘     └── h₃ ──┘

Each connection has a weight (w)
Each neuron applies an activation function

Key components:

  • Weights: Parameters the model learns during training
  • Bias: Additional parameter to shift the activation function
  • Activation Function: ReLU, Sigmoid, Softmax — introduces non-linearity
  • Loss Function: Measures how wrong the model's predictions are
  • Optimizer: Updates weights to minimize loss (e.g., SGD, Adam)

4.2. Types of Neural Networks

TypeBest ForAWS Service
CNN (Convolutional NN)Images, videoAmazon Rekognition
RNN/LSTM (Recurrent NN)Sequential data, time seriesAmazon Forecast
TransformerNLP, text generationAmazon Bedrock (LLMs)
GAN (Generative Adversarial)Generate new data (images)—

5. Model Evaluation Concepts

5.1. Overfitting vs Underfitting

ProblemTraining AccuracyTest AccuracyCauseSolution
OverfittingVery HighLowModel memorizes training dataMore data, regularization, dropout, early stopping
UnderfittingLowLowModel too simpleMore features, more complex model, longer training
Good FitHighHighBalanced complexity—

5.2. Bias-Variance Tradeoff

  • High Bias = Underfitting (model quá đơn giản, bỏ qua patterns)
  • High Variance = Overfitting (model quá phức tạp, nhạy cảm với noise)
  • Mục tiêu: tìm sweet spot giữa bias và variance

5.3. Common Metrics

Classification metrics:

MetricFormulaWhen to use
Accuracy(TP + TN) / TotalBalanced classes
PrecisionTP / (TP + FP)"Don't flag innocent as spam"
RecallTP / (TP + FN)"Don't miss any fraud"
F1 Score2 × (P × R) / (P + R)Imbalanced classes
AUC-ROCArea under ROC curveBinary classification overall

Regression metrics:

  • RMSE (Root Mean Square Error): Penalizes large errors
  • MAE (Mean Absolute Error): Average error magnitude
  • R²: How well model explains variance (1.0 = perfect)

6. Key Terms Cheat Sheet

TermDefinition (for exam)
FeatureInput variable used for prediction (column in data)
Label / TargetThe answer we want the model to predict
HyperparameterSettings configured BEFORE training (learning rate, epochs)
ParameterValues the model learns DURING training (weights, biases)
EpochOne complete pass through the entire training dataset
Batch SizeNumber of samples processed before updating weights
InferenceUsing a trained model to make predictions on new data
Transfer LearningUsing a pre-trained model and adapting it for a new task

7. Practice Questions

Q1: A company wants to predict whether customers will cancel their subscription (yes/no). Which ML approach is most appropriate?

  • A) Unsupervised Learning — Clustering
  • B) Supervised Learning — Regression
  • C) Supervised Learning — Classification ✓
  • D) Reinforcement Learning

Explanation: Predicting a binary outcome (yes/no) with labeled historical data = supervised classification.

Q2: A retail company has customer purchase data but NO predefined groups. They want to segment customers into groups for targeted marketing. Which approach should they use?

  • A) Supervised Learning — Classification
  • B) Unsupervised Learning — Clustering ✓
  • C) Reinforcement Learning
  • D) Supervised Learning — Regression

Explanation: No labels + finding natural groups in data = unsupervised clustering.

Q3: A model performs extremely well on training data (99% accuracy) but poorly on new data (65% accuracy). What is this called?

  • A) Underfitting
  • B) Overfitting ✓
  • C) High bias
  • D) Regularization

Explanation: High training accuracy + low test accuracy = overfitting (model memorized training data).