Chuyển đến nội dung chính

Lesson 1: What is ML? How to study without being overwhelmed

Compare AI/ML/Deep Learning with real-life examples. Introducing end-to-end workflow and practice-oriented ML learning mindset.

🧠 AI & ML — Lesson 0 Lesson 1: What is ML? How to study without being overwhelmed

Machine Learning: From Basics to Advanced

Part 0: Getting started for newbies (Week 0)

xdev.asia

Introduction

If you're new to Machine Learning, the hardest thing is often not the code but the feeling that everything is too broad: models, data, metrics, train/test, then add AI, Deep Learning, LLM. This article's mission is to do one very clear thing: create an overall map for you to know what you're learning, what you're learning for, and in what order to learn so you don't get overwhelmed.

Lesson objectives

  • Distinguish between AI, Machine Learning, Deep Learning and Data Science
  • Understand the end-to-end ML process at an intuitive level
  • Know which problem should be solved using ML and which problem does not need ML

1. What is Machine Learning?

Machine Learning is a way for computers to learn rules from data instead of us writing all the rules by hand.

For example:

  • Email spam problem: instead of writing hundreds of rules like "if there is a word free then it is spam", we give the model thousands of emails labeled as spam/not spam so it can learn the pattern on its own.
  • House price prediction problem: model learns the relationship between area, location, number of rooms, house age and selling price.

Important point: ML is not “intelligent” in the same way humans are. It's only good at finding statistical regularities if the data is good enough and the goal is clear enough.

2. Distinguish commonly confused concepts

ConceptShort meaningExample
AIThe broadest concept: making machines behave like intelligencechatbots, AI games
Machine LearningBranch of AI that learns from datachurn prediction, fraud detection
Deep LearningML uses multi-layer neural networksimage recognition, speech recognition
Data ScienceMining data for insight and decision supportdashboard, cohort analysis

Quick way to remember:

  • AI is the "big box".
  • ML is a popular way to do AI.
  • Deep Learning is a group of techniques within ML.
  • Data Science is not exactly the same as ML, but uses ML as a powerful tool.

3. End-to-end ML pipeline

A real ML project usually follows the following flow:

  1. Identify the business problem.
  2. Transform a business problem into a prediction problem.
  3. Collect and understand data.
  4. Choose an evaluation metric.
  5. Create baseline.
  6. Train the model better than baseline.
  7. Check for errors and explain the results.
  8. Put the model into use.
  9. Monitor drift and retrain when needed.

Example churn problem:

  • Business question: Which customers are about to leave the service?
  • ML formulation: predict the probability that a customer will cancel in the next 30 days.
  • Input features: number of login times, number of support tickets, service packages, usage time.
  • Output: churn probability or churn/no churn label.

4. When should you use ML, when not?

ML is suitable when:

  • There is enough historical data.
  • There is a clear output for learning.
  • Difficult rules are written in hard rules.
  • Has practical value if the prediction is better than the current one.

You should not use ML when:

  • The problem only needs a few fixed rules to solve.
  • No data or too dirty data.
  • There is no right/wrong way to measure the model.
  • The cost of implementing ML is higher than the benefits.

Example without ML:

  • Calculate VAT according to fixed law.
  • Automatically number invoices.
  • Check if a field is empty or not.

5. Basic types of ML problems

Regression

Predict a real number.

For example:

  • House price
  • Revenue next week
  • Temperature

Classification

Predict a label.

For example:

  • Spam / no spam
  • Churn / no churn
  • Cheating/not cheating

Clustering

There are no built-in labels, the model groups the data itself.

For example:

  • Customer segmentation
  • Group products according to purchasing behavior

Anomaly Detection

Find abnormalities.

For example:

  • Unusual transactions
  • The device generates an error signal

6. What is Baseline and why is it required?

Baseline is the simplest model or strategy for comparison.

For example:

  • Predict house price using the average price of the training set.
  • Predict all customers to have "no churn".

If the new model does not beat the baseline, then the project has no value. Newbies often make the mistake of jumping straight into XGBoost or a neural network without knowing if the complex model actually improves anything.

7. A very short example with scikit-learn

from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

model = LogisticRegression(max_iter=300)
model.fit(X_train, y_train)

pred = model.predict(X_test)
print("Accuracy:", accuracy_score(y_test, pred))

The important thing in this code is not to remember each function, but to understand the structure:

  • has data X, y
  • separate train/test
  • fit the model on the train
  • evaluated on test

It is the backbone of most ML workflows.

8. In what order should newbies learn?

The reasonable order for this series is:

  1. Understand the big picture.
  2. Know how to manipulate data with Pandas.
  3. Make the first model quickly.
  4. Learn metrics and overfitting early.
  5. Then go into pipeline, tuning and production.

If you learn backwards, for example jumping into tuning before understanding the baseline, you will easily "learn by rote from a notebook".

Practice exercises

  1. Write in your own words the difference between AI, ML and Deep Learning.
  2. Choose 3 problems in work or life and classify whether they are regression, classification, clustering or should not use ML.
  3. Find an example problem that can be solved better using rule-based than ML.

Common mistakes

  • Learn algorithms before understanding business problems.
  • Using ML for too simple problems.
  • Only care about the model and ignore data and metrics.

Completion criteria

  • Explain what ML is in everyday language
  • Distinguish between 4 basic types of problems
  • Describe the end-to-end ML workflow at a high level