Introduction
After you have trained the first model, the next question is: why did the model learn? This article helps you understand linear regression at an intuitive level strong enough to read the loss, debug the results, and know when the model is learning correctly or incorrectly.
Lesson objectives
- Understand what linear regression is trying to learn.
- Intuitively grasp loss function, gradient descent and regularization.
- Know when to use linear regression and when not to force it.
What is Linear regression doing?
Imagine you want to predict house prices from area. The linear model tries to find the best straight line so that the error between the predicted price and the actual price is minimal.
Basic formula:
$$ \hat{y} = w_1x_1 + w_2x_2 + ... + w_nx_n + b $$
Loss function: a measure of how bad the model is
With regression, a familiar loss is MSE:
$$ MSE = \frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y_i})^2 $$
The meaning is very common: the more the predicted model deviates from reality, the greater the loss. Squaring helps severely penalize points that are too far away.
Gradient descent: how the model corrects errors
Gradient descent is an iterative process: predict, calculate loss, see how to increase or decrease weights, then update many times until the model is more stable.
If the learning rate is too large, the model easily jumps back and forth and does not converge. If it's too small, the model learns very slowly.
Minimal code example
from sklearn.linear_model import LinearRegression, Ridge
from sklearn.metrics import mean_absolute_error
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = LinearRegression()
model.fit(X_train, y_train)
preds = model.predict(X_test)
print('MAE:', mean_absolute_error(y_test, preds))
print('Intercept:', model.intercept_)
print('Coefficients:', model.coef_)
Regularization to prevent overlearning
- Ridge often helps smoother weighting.
- Lasso can push some weights to 0, suitable when you want to select features.
Common mistakes
- Data normalization is completed but not applied the same way to the test set.
- Conclusion that the coefficient is a cause and effect relationship.
- Only look at the train score without looking at the test error.
Practice exercises
- Train Linear Regression and Ridge on the same data set.
- Compare the MAE of the two models.
- Write in 5 short lines: when is Ridge better than pure Linear Regression?
Completion criteria
- Explain the loss function in simple language.
- Understand gradient descent used to update weights.
- Compare Linear Regression and Ridge on a real example.
Practice step by step (advanced)
- Choose a regression dataset with at least 6 features.
- Train baseline using Linear Regression without regularization.
- Test Ridge with 5 different alpha values.
- Draw a graph comparing MAE for each alpha.
- Write a comment: How does too large an alpha affect bias/variance?
Artifact should be submitted
- Notebook has a comparison of Linear vs Ridge.
- Table of MAE and RMSE results.
- Conclusion in 8-10 lines about choosing regularization.
Self-test questions
- Why is MSE more sensitive to outliers than MAE?
- How does learning rate affect convergence speed?
- When is Lasso worth trying more than Ridge?