Introduction
Learning ML isn't just about getting better grades. A model that is strong on the training set but weak on new data is a model that is not ready for practical use. This article helps you spot overfitting and underfitting with very specific signs.
Lesson objectives
- Distinguish between overfitting and underfitting.
- Know how to read train score and validation score together.
- There is a checklist to handle when the model has not learned properly.
What is Underfitting?
Underfitting occurs when the model is too simple or the features are too poor, causing it to not learn the signal well enough even on the training set.
What is overfitting?
Overfitting occurs when the model learns both the real signal and the unique noise of the training set.
Typical signs:
- Train score is very high.
- Validation score is clearly low.
- Each time you change the split, the results fluctuate drastically.
How to handle it in a pragmatic way
When underfitting: add useful features, use a stronger model, train longer if it is an iterative algorithm.
When overfitting: reduce model complexity, add regularization, increase data or use cross-validation.
Learning curves
Learning curve shows you if you increase the data, the model is likely to improve. This is a better way to diagnose than guessing.
Common mistakes
- Constantly increasing model complexity without proper validation.
- Run many times and then choose the best split.
- Fix features based on test set.
Practice exercises
- Train 3 models with increasing complexity.
- Record training score and validation score.
- Conclusion which model is underfit, which model is overfit, which model is most reasonable.
Completion criteria
- Explain two concepts in everyday language.
- Know how to read the difference between training and validation.
- Suggest at least 2 solutions for each situation.
Practice step by step (advanced)
- Train three models with increasing complexity.
- Collect training/validation scores for each model.
- Draw the learning curve according to the amount of training data.
- Try regularization or reducing model depth.
- Record changes before and after editing.
Artifact should be submitted
- Learning curve chart.
- Comparison table before/after adjustment.
- Checklist of overfitting decisions applied to the following project.
Self-test questions
- Which sign best distinguishes underfit and overfit?
- Why can more data reduce overfitting?
- When is reducing model complexity the most reasonable choice?