Chuyển đến nội dung chính

Lesson 5: Mini-project 1 — Predicting house prices

The first complete practice session: simple EDA, train/test split, baseline model, evaluation and lessons learned.

🧠 AI & ML — Lesson 4 Lesson 5: Mini-project 1 — Predicting house prices

Machine Learning: From Basics to Advanced

Part 0: Getting started for newbies (Week 0)

xdev.asia

Introduction

This is the first mini-project of the series. The goal is not to optimize to the best level, but for you to go through the entire process from start to finish: understand the data, choose features, split train/test, build baseline, train model, evaluate and draw conclusions.

Problem context

You have housing data in a city with the following information:

  • area
  • number of bedrooms
  • toilet number
  • district
  • age of house
  • selling price

The goal is to predict the house price for a new home.

Mini-project goal

  • Do basic EDA to understand the data
  • Build the first baseline
  • Train at least one regression model
  • Give clear conclusions using metrics

1. Business question

A realtor or listing system wants to estimate a fair selling price for a new home.

Corresponding ML question:

With the input information of the house, what is the estimated selling price?

2. Minimum EDA required

import pandas as pd

df = pd.read_csv('data/raw/houses.csv')

print(df.head())
print(df.shape)
print(df.info())
print(df.isnull().sum())
print(df.describe())

You need to answer at least the following questions:

  1. How many rows of data are there?
  2. Are there any columns missing data?
  3. What is the target column?
  4. Are there any columns that don't make sense to include in the model?

3. Select the first feature

In the first round, don't choose too many columns. Just a few obvious features:

features = ['dien_tich', 'so_phong', 'so_toilet', 'tuoi_nha']
X = df[features]
y = df['gia']

Reasons for choosing few features in the first round:

  • easy to debug
  • easy to understand the influence of each variable
  • reduce errors in complex data processing

4. Baseline

from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

baseline_pred = [y_train.mean()] * len(y_test)
baseline_mae = mean_absolute_error(y_test, baseline_pred)

print('Baseline MAE:', baseline_mae)

5. First model

from sklearn.linear_model import LinearRegression

model = LinearRegression()
model.fit(X_train, y_train)

pred = model.predict(X_test)
model_mae = mean_absolute_error(y_test, pred)

print('Model MAE:', model_mae)

6. Read the results

For example:

  • Baseline MAE: 0.85 billion
  • Model MAE: 0.42 billion

Preliminary conclusion:

  • The model is quite clearly better than the baseline
  • Initial workflow is on track
  • Can be further improved with feature engineering or more powerful models

7. Expand one step further

Try adding a categorical variable like quan:

X = pd.get_dummies(df[['dien_tich', 'so_phong', 'so_toilet', 'tuoi_nha', 'quan']], drop_first=True)
y = df['gia']

Then train again and compare metrics. This is a simple way to test whether adding location information actually helps the model.

8. Short report sample after mini-project

You should write a summary like this:

I use 4 basic numerical features to predict house prices. Baseline prediction by average price gives MAE = X, while Linear Regression gives MAE = Y. This shows that the model has learned the relationship between feature and selling price. However, the model currently does not use the detailed position feature and has not processed the outlier.

This is a very important habit because ML is not just about coding, but also about communicating the results.

Extra challenge

  1. Compare Linear Regression with Random Forest Regressor.
  2. Try removing a feature to see how the metric changes.
  3. Create features gia_m2 and see if there is leakage if used incorrectly.

Common mistakes

  • Using every column without understanding the meaning.
  • Forgot to separate train/test before evaluating.
  • Seeing that the metric is better than the baseline and then drawing conclusions too early, without checking the data.

Completion criteria

  • Finished running the mini-project from start to finish
  • Has a baseline and at least 1 machine learning model
  • Can write a short, easy-to-understand conclusion about the results