Introduction
This is the first mini-project of the series. The goal is not to optimize to the best level, but for you to go through the entire process from start to finish: understand the data, choose features, split train/test, build baseline, train model, evaluate and draw conclusions.
Problem context
You have housing data in a city with the following information:
- area
- number of bedrooms
- toilet number
- district
- age of house
- selling price
The goal is to predict the house price for a new home.
Mini-project goal
- Do basic EDA to understand the data
- Build the first baseline
- Train at least one regression model
- Give clear conclusions using metrics
1. Business question
A realtor or listing system wants to estimate a fair selling price for a new home.
Corresponding ML question:
With the input information of the house, what is the estimated selling price?
2. Minimum EDA required
import pandas as pd
df = pd.read_csv('data/raw/houses.csv')
print(df.head())
print(df.shape)
print(df.info())
print(df.isnull().sum())
print(df.describe())
You need to answer at least the following questions:
- How many rows of data are there?
- Are there any columns missing data?
- What is the target column?
- Are there any columns that don't make sense to include in the model?
3. Select the first feature
In the first round, don't choose too many columns. Just a few obvious features:
features = ['dien_tich', 'so_phong', 'so_toilet', 'tuoi_nha']
X = df[features]
y = df['gia']
Reasons for choosing few features in the first round:
- easy to debug
- easy to understand the influence of each variable
- reduce errors in complex data processing
4. Baseline
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
baseline_pred = [y_train.mean()] * len(y_test)
baseline_mae = mean_absolute_error(y_test, baseline_pred)
print('Baseline MAE:', baseline_mae)
5. First model
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(X_train, y_train)
pred = model.predict(X_test)
model_mae = mean_absolute_error(y_test, pred)
print('Model MAE:', model_mae)
6. Read the results
For example:
- Baseline MAE: 0.85 billion
- Model MAE: 0.42 billion
Preliminary conclusion:
- The model is quite clearly better than the baseline
- Initial workflow is on track
- Can be further improved with feature engineering or more powerful models
7. Expand one step further
Try adding a categorical variable like quan:
X = pd.get_dummies(df[['dien_tich', 'so_phong', 'so_toilet', 'tuoi_nha', 'quan']], drop_first=True)
y = df['gia']
Then train again and compare metrics. This is a simple way to test whether adding location information actually helps the model.
8. Short report sample after mini-project
You should write a summary like this:
I use 4 basic numerical features to predict house prices. Baseline prediction by average price gives MAE = X, while Linear Regression gives MAE = Y. This shows that the model has learned the relationship between feature and selling price. However, the model currently does not use the detailed position feature and has not processed the outlier.
This is a very important habit because ML is not just about coding, but also about communicating the results.
Extra challenge
- Compare Linear Regression with Random Forest Regressor.
- Try removing a feature to see how the metric changes.
- Create features
gia_m2and see if there is leakage if used incorrectly.
Common mistakes
- Using every column without understanding the meaning.
- Forgot to separate train/test before evaluating.
- Seeing that the metric is better than the baseline and then drawing conclusions too early, without checking the data.
Completion criteria
- Finished running the mini-project from start to finish
- Has a baseline and at least 1 machine learning model
- Can write a short, easy-to-understand conclusion about the results