Chuyển đến nội dung chính

Lesson 1: AI, ML & Deep Learning — Concepts and Terminology

AI vs ML vs DL. Supervised, Unsupervised, Reinforcement Learning. Classification, Regression, Clustering. Neural Networks basics. Training, Validation, Test sets. Bias-Variance tradeoff.

AI, ML and Deep Learning Hierarchy

AI, ML and Deep Learning — nested relationship and the three machine learning paradigms

Domain 1 Overview

Domain 1 accounts for 20% of the AIF-C01 exam. You need to understand the foundational concepts of AI, ML, and Deep Learning — no coding is required, but you must be able to distinguish when to use which approach.

Exam tip: This domain often features questions like "Which type of machine learning is BEST suited for..." — requiring you to select the right paradigm for a given use case.

1. AI vs Machine Learning vs Deep Learning

These three concepts have a nested relationship:

┌─────────────────────────────────────────────┐
│  Artificial Intelligence (AI)               │
│  "Machines that mimic human intelligence"   │
│  ┌───────────────────────────────────────┐   │
│  │  Machine Learning (ML)               │   │
│  │  "Learning from data without         │   │
│  │   explicit programming"              │   │
│  │  ┌─────────────────────────────────┐  │   │
│  │  │  Deep Learning (DL)             │  │   │
│  │  │  "Neural networks with many     │  │   │
│  │  │   layers"                       │  │   │
│  │  └─────────────────────────────────┘  │   │
│  └───────────────────────────────────────┘   │
└─────────────────────────────────────────────┘
ConceptDefinitionExample
AIBroad field — machines performing tasks that typically require human intelligenceChatbot, self-driving car, chess engine
MLSubset of AI — algorithms learn patterns from dataSpam filter, recommendation engine
DLSubset of ML — neural networks with multiple layersImage recognition, language translation

Key Differences for the Exam

  • Traditional Programming: Rules + Data → Output
  • Machine Learning: Data + Output → Rules (model learns the rules)
  • Deep Learning: Automatically extracts features from raw data (no manual feature engineering needed)

2. Three ML Paradigms

2.1. Supervised Learning

The model learns from labeled data — each input comes with the correct output (label/target).

Task TypeOutputUse CaseAlgorithms
ClassificationDiscrete categorySpam vs Not Spam, Fraud detectionLogistic Regression, Random Forest, SVM
RegressionContinuous numberHouse price prediction, Stock forecastLinear Regression, XGBoost

Exam tip: If the question says "predict a category" or "classify" → Classification. If it says "predict a number/value" → Regression.

2.2. Unsupervised Learning

The model learns from unlabeled data — it finds patterns and structure in the data on its own.

Task TypeWhat it doesUse Case
ClusteringGroup similar data pointsCustomer segmentation, Document grouping
Dimensionality ReductionReduce features while preserving infoData visualization, noise reduction
Anomaly DetectionFind unusual data pointsFraud detection, equipment failure
AssociationFind rules between items"Customers who bought X also bought Y"

2.3. Reinforcement Learning (RL)

An agent learns through trial-and-error in an environment. It receives a reward (positive) or penalty (negative) for each action.

Agent → Action → Environment → State + Reward → Agent (loop)

Use cases:

  • Game AI (AlphaGo)
  • Robotics navigation
  • Autonomous driving
  • AWS DeepRacer (self-driving car simulation)

2.4. Choosing the Right Paradigm — Exam Decision Tree

Do you have labeled data?
├── YES → Supervised Learning
│   ├── Predicting a category? → Classification
│   └── Predicting a number? → Regression
├── NO →
│   ├── Want to find groups/patterns? → Unsupervised (Clustering)
│   └── Learning through trial & error? → Reinforcement Learning

3. Data Concepts for ML

3.1. Data Types

TypeDescriptionExamples
StructuredOrganized in rows & columns (tabular)CSV, database tables, spreadsheets
Semi-structuredHas some organization but flexibleJSON, XML, log files
UnstructuredNo predefined formatImages, videos, audio, free text
Time-seriesData points indexed by timeStock prices, IoT sensor readings

3.2. Labeled vs Unlabeled Data

  • Labeled data: Each data point has an accompanying answer (label). Example: email + tag "spam"/"not spam". Used for Supervised Learning.
  • Unlabeled data: Only data, no labels. Used for Unsupervised Learning.
  • Amazon SageMaker Ground Truth: An AWS service that helps label data (human + ML-assisted labeling).

3.3. Training, Validation, Test Sets

┌────────────────────────────────────────────────┐
│              Full Dataset (100%)               │
├──────────────────┬──────────┬──────────────────┤
│  Training (70%)  │ Val(15%) │   Test (15%)     │
│  Model learns    │ Tune     │ Final evaluation │
│  from this data  │ hyper-   │ (never seen      │
│                  │ params   │  during training) │
└──────────────────┴──────────┴──────────────────┘
  • Training set: The model learns patterns from this data
  • Validation set: Used to tune hyperparameters and prevent overfitting
  • Test set: Final evaluation — the model has never seen this data

4. Neural Networks Basics

4.1. Architecture

Input Layer → Hidden Layer(s) → Output Layer
    x₁ ──┐     ┌── h₁ ──┐
    x₂ ──┼─────┼── h₂ ──┼──── ŷ (prediction)
    x₃ ──┘     └── h₃ ──┘

Each connection has a weight (w)
Each neuron applies an activation function

Key components:

  • Weights: Parameters the model learns during training
  • Bias: Additional parameter to shift the activation function
  • Activation Function: ReLU, Sigmoid, Softmax — introduces non-linearity
  • Loss Function: Measures how wrong the model's predictions are
  • Optimizer: Updates weights to minimize loss (e.g., SGD, Adam)

4.2. Types of Neural Networks

TypeBest ForAWS Service
CNN (Convolutional NN)Images, videoAmazon Rekognition
RNN/LSTM (Recurrent NN)Sequential data, time seriesAmazon Forecast
TransformerNLP, text generationAmazon Bedrock (LLMs)
GAN (Generative Adversarial)Generate new data (images)—

5. Model Evaluation Concepts

5.1. Overfitting vs Underfitting

ProblemTraining AccuracyTest AccuracyCauseSolution
OverfittingVery HighLowModel memorizes training dataMore data, regularization, dropout, early stopping
UnderfittingLowLowModel too simpleMore features, more complex model, longer training
Good FitHighHighBalanced complexity—

5.2. Bias-Variance Tradeoff

  • High Bias = Underfitting (model too simple, misses patterns)
  • High Variance = Overfitting (model too complex, sensitive to noise)
  • The goal: find the sweet spot between bias and variance

5.3. Common Metrics

Classification metrics:

MetricFormulaWhen to use
Accuracy(TP + TN) / TotalBalanced classes
PrecisionTP / (TP + FP)"Don't flag innocent as spam"
RecallTP / (TP + FN)"Don't miss any fraud"
F1 Score2 × (P × R) / (P + R)Imbalanced classes
AUC-ROCArea under ROC curveBinary classification overall

Regression metrics:

  • RMSE (Root Mean Square Error): Penalizes large errors
  • MAE (Mean Absolute Error): Average error magnitude
  • R²: How well model explains variance (1.0 = perfect)

6. Key Terms Cheat Sheet

TermDefinition (for exam)
FeatureInput variable used for prediction (column in data)
Label / TargetThe answer we want the model to predict
HyperparameterSettings configured BEFORE training (learning rate, epochs)
ParameterValues the model learns DURING training (weights, biases)
EpochOne complete pass through the entire training dataset
Batch SizeNumber of samples processed before updating weights
InferenceUsing a trained model to make predictions on new data
Transfer LearningUsing a pre-trained model and adapting it for a new task

7. Practice Questions

Q1: A company wants to predict whether customers will cancel their subscription (yes/no). Which ML approach is most appropriate?

  • A) Unsupervised Learning — Clustering
  • B) Supervised Learning — Regression
  • C) Supervised Learning — Classification ✓
  • D) Reinforcement Learning

Explanation: Predicting a binary outcome (yes/no) with labeled historical data = supervised classification.

Q2: A retail company has customer purchase data but NO predefined groups. They want to segment customers into groups for targeted marketing. Which approach should they use?

  • A) Supervised Learning — Classification
  • B) Unsupervised Learning — Clustering ✓
  • C) Reinforcement Learning
  • D) Supervised Learning — Regression

Explanation: No labels + finding natural groups in data = unsupervised clustering.

Q3: A model performs extremely well on training data (99% accuracy) but poorly on new data (65% accuracy). What is this called?

  • A) Underfitting
  • B) Overfitting ✓
  • C) High bias
  • D) Regularization

Explanation: High training accuracy + low test accuracy = overfitting (model memorized training data).