Introduction
This is a very important article because it helps you overcome the biggest psychological barrier: "know the theory but have never trained any model". The goal of this article is for you to manually fit your first model, measure the results, and understand why baseline is always a mandatory step in every serious ML project.
Lesson objectives
- Train is first modeled using scikit-learn
- Understand how train/test split works
- Know how to create baselines and compare with machine learning models
1. Sample problem: predicting house prices
We assume we have house data consisting of the following columns:
- area
- room number
- age of house
- district
- selling price
In this problem:
Xis the feature setyis the house price
This is a regression problem because the output is a real number.
2. What is Baseline?
Baseline is the simplest option to use as a landmark.
House price prediction example:
- baseline 1: always guess the average price of the training set
- Baseline 2: always guess at the median price
If the ML model is not better than the baseline, then there is no reason to use a complex model.
3. Train/test split
We divide the data into two parts:
train: to model learningtest: to evaluate on unseen data
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
Meaning of random_state=42 is for you and someone else to run again and get the same results.
4. First model with Linear Regression
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error
df = pd.read_csv('data/raw/houses.csv')
features = ['dien_tich', 'so_phong', 'tuoi_nha']
X = df[features]
y = df['gia']
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
model = LinearRegression()
model.fit(X_train, y_train)
pred = model.predict(X_test)
mae = mean_absolute_error(y_test, pred)
print('MAE:', mae)
The most important thing to understand:
fit()is when the model learns from train datapredict()is when the model predicts on new data- metric tells us how good the model is
5. Measure baseline for comparison
baseline_pred = [y_train.mean()] * len(y_test)
baseline_mae = mean_absolute_error(y_test, baseline_pred)
print('Baseline MAE:', baseline_mae)
print('Model MAE:', mae)
If Model MAE < Baseline MAE, the model created value.
6. Why choose MAE?
With regression, common metrics are:
- MAE
- MSE
- RMSE
- $R^2$
For beginners, MAE is the easiest metric to understand because it represents the average error in the correct units of the problem.
For example:
- MAE = 0.25 billion means that the average model deviation is about 250 million.
7. The first model does not need to be strong
Many people just find Linear Regression simple so they want to skip it and switch to XGBoost. That is the wrong learning rhythm.
Reasons to start simple:
- easy to debug
- easy to explain
- easy to detect data errors
- Have a thinking baseline for the following models
8. Checklist reads results properly
After training, don't just look at a metric number and make conclusions. Ask:
- Is the model better than the baseline?
- Is the metric correct for the problem?
- Is the test data representative of real data?
- Are there any columns leaking?
9. A very short classification variant
If the problem is classification, the code structure is almost identical, only the model and metrics are changed.
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
clf = LogisticRegression(max_iter=300)
clf.fit(X_train, y_train)
pred = clf.predict(X_test)
print('Accuracy:', accuracy_score(y_test, pred))
This helps you see how the same workflow can be applied to many different problems.
Practice exercises
- Create a baseline for a regression problem.
- Train a Linear Regression model.
- Compare metrics between baseline and model.
- Write 3 concluding sentences: is the model worth keeping, and why.
Common mistakes
- Evaluate the model on the train set itself.
- Do not compare with baseline.
- Using metrics without understanding its actual meaning.
Completion criteria
- Self-train the first model
- Create a reasonable baseline
- Explain why the model is better or worse than the baseline