Chuyển đến nội dung chính

Lesson 4: First model in 30 minutes + baseline

Create your first model with scikit-learn, understand what a baseline is and why you always need a baseline before optimizing.

🧠 AI & ML — Lesson 3 Lesson 4: First model in 30 minutes + baseline. baseline

Machine Learning: From Basics to Advanced

Part 0: Getting started for newbies (Week 0)

xdev.asia

Introduction

This is a very important article because it helps you overcome the biggest psychological barrier: "know the theory but have never trained any model". The goal of this article is for you to manually fit your first model, measure the results, and understand why baseline is always a mandatory step in every serious ML project.

Lesson objectives

  • Train is first modeled using scikit-learn
  • Understand how train/test split works
  • Know how to create baselines and compare with machine learning models

1. Sample problem: predicting house prices

We assume we have house data consisting of the following columns:

  • area
  • room number
  • age of house
  • district
  • selling price

In this problem:

  • X is the feature set
  • y is the house price

This is a regression problem because the output is a real number.

2. What is Baseline?

Baseline is the simplest option to use as a landmark.

House price prediction example:

  • baseline 1: always guess the average price of the training set
  • Baseline 2: always guess at the median price

If the ML model is not better than the baseline, then there is no reason to use a complex model.

3. Train/test split

We divide the data into two parts:

  • train: to model learning
  • test: to evaluate on unseen data
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

Meaning of random_state=42 is for you and someone else to run again and get the same results.

4. First model with Linear Regression

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error

df = pd.read_csv('data/raw/houses.csv')

features = ['dien_tich', 'so_phong', 'tuoi_nha']
X = df[features]
y = df['gia']

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

model = LinearRegression()
model.fit(X_train, y_train)

pred = model.predict(X_test)
mae = mean_absolute_error(y_test, pred)
print('MAE:', mae)

The most important thing to understand:

  • fit() is when the model learns from train data
  • predict() is when the model predicts on new data
  • metric tells us how good the model is

5. Measure baseline for comparison

baseline_pred = [y_train.mean()] * len(y_test)
baseline_mae = mean_absolute_error(y_test, baseline_pred)

print('Baseline MAE:', baseline_mae)
print('Model MAE:', mae)

If Model MAE < Baseline MAE, the model created value.

6. Why choose MAE?

With regression, common metrics are:

  • MAE
  • MSE
  • RMSE
  • $R^2$

For beginners, MAE is the easiest metric to understand because it represents the average error in the correct units of the problem.

For example:

  • MAE = 0.25 billion means that the average model deviation is about 250 million.

7. The first model does not need to be strong

Many people just find Linear Regression simple so they want to skip it and switch to XGBoost. That is the wrong learning rhythm.

Reasons to start simple:

  • easy to debug
  • easy to explain
  • easy to detect data errors
  • Have a thinking baseline for the following models

8. Checklist reads results properly

After training, don't just look at a metric number and make conclusions. Ask:

  1. Is the model better than the baseline?
  2. Is the metric correct for the problem?
  3. Is the test data representative of real data?
  4. Are there any columns leaking?

9. A very short classification variant

If the problem is classification, the code structure is almost identical, only the model and metrics are changed.

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

clf = LogisticRegression(max_iter=300)
clf.fit(X_train, y_train)
pred = clf.predict(X_test)

print('Accuracy:', accuracy_score(y_test, pred))

This helps you see how the same workflow can be applied to many different problems.

Practice exercises

  1. Create a baseline for a regression problem.
  2. Train a Linear Regression model.
  3. Compare metrics between baseline and model.
  4. Write 3 concluding sentences: is the model worth keeping, and why.

Common mistakes

  • Evaluate the model on the train set itself.
  • Do not compare with baseline.
  • Using metrics without understanding its actual meaning.

Completion criteria

  • Self-train the first model
  • Create a reasonable baseline
  • Explain why the model is better or worse than the baseline