Introduction
A single train-test split can give you a false sense of model quality. Cross-validation provides a more stable evaluation, while hyperparameter tuning helps you find a reasonable configuration without relying on the luck of a data split.
Lesson objectives
- Understand why you should not absolutely believe in a single split.
- Use cross-validation to estimate more stable performance.
- Tuning hyperparameters according to a controlled process.
What is cross-validation?
The data is divided into many folds. Each time, one fold does validation, the remaining folds do training. The final result is the average of multiple evaluations.
What is hyperparameter?
Hyperparameters are values you choose before training, for example max_depth, n_estimators, C or learning_rate. They are different from parameters where the model learns itself from data.
Practical tuning
- Start with a simple baseline.
- Select the least important hyperparameters.
- Use GridSearchCV or RandomizedSearchCV.
- Track both mean score and standard deviation between folds.
Sample code
from sklearn.model_selection import RandomizedSearchCV
from sklearn.ensemble import RandomForestClassifier
Common mistakes
- Tuning is too wide when the baseline is not yet stable.
- Use test set to adjust hyperparameter.
- Run lots of tests but don't record the results.
Practice exercises
- Tuning a tree-based model with 3 hyperparameters.
- Compare scores before tuning and after tuning.
- Record comments: has the tuning really improved, or has it only improved very little?
Completion criteria
- Can use cross-validation in a complete pipeline.
- Distinguish between validation used for tuning and testing used for final evaluation.
- Record experiments systematically.
Practice step by step (advanced)
- Run baseline with a standard split.
- Run KFold or StratifiedKFold 5 folds.
- Tuning using RandomizedSearchCV with no more than 4 important parameters.
- Compare mean score and std score between configurations.
- Finalize configuration according to performance + stability.
Artifact should be submitted
- Table of top 10 configurations by score.
- Score distribution chart by fold.
- Rules for stopping tuning to avoid over-search.
Self-test questions
- When should RandomizedSearchCV be preferred over GridSearchCV?
- What does a high standard deviation between folds mean?
- Why is the test set not used for tuning?