Chuyển đến nội dung chính

Lesson 12: Pipelines & ColumnTransformer with scikit-learn

Build a pipeline that resists manual errors, has good reuse, and reduces the risk of leakage in training.

🧠 AI & ML — Lesson 11 Lesson 12: Pipelines & ColumnTransformer with scikit-learn

Machine Learning: From Basics to Advanced

Part 2: Industry standard workflow

xdev.asia

Introduction

When an ML project starts with many pre-processing steps, it's easy to go wrong by hand-writing each step. Pipeline and ColumnTransformer help you combine all preprocessing and models into a unified flow, easy to reproduce, easy to debug and reduce the risk of leakage.

Lesson objectives

  • Understand why you should use Pipeline instead of discrete processing.
  • Know how to use ColumnTransformer for data of many column types.
  • Build a consistent train/predict workflow.

Why is pipeline important?

Pipelines help avoid common mistakes like fitting scalers on both train and test, forgetting to apply the same transform when predicting new data, or saving the model but forgetting the preprocessing logic.

Standard structure

  • Numeric: impute then scale.
  • Categorical: impute then one-hot.
  • Merge using ColumnTransformer.
  • Set classifier or regressor in the last step.

Code example

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.impute import SimpleImputer
from sklearn.ensemble import RandomForestClassifier

Benefits in practice

  • Easy to put into cross-validation.
  • Easy to save/load with joblib.
  • Fewer errors when deploying batch inference.

Common mistakes

  • Wrong use of column list.
  • Added new feature but forgot to update ColumnTransformer.
  • Call fit_transform outside the pipeline and then fit again in the pipeline.

Practice exercises

  • Build a complete pipeline for churn or housing data.
  • Compare code using pipeline and code processed manually.
  • Write 5 lines: which type of error does the pipeline help reduce the most?

Completion criteria

  • Can build Pipeline and ColumnTransformer yourself.
  • Understand how pipelines help avoid leakage.
  • Complete workflow can be saved/loaded.

Practice step by step (advanced)

  1. Write a complete pipeline including preprocessing + model.
  2. Split numeric/categorical into two transform branches.
  3. Use the same pipeline to train, validate and predict new samples.
  4. Save the pipeline with joblib and reload it for prediction.
  5. Write small tests to ensure the input schema is not broken.

Artifact should be submitted

  • File pipeline can be reused.
  • Minimum predict script for 1 record.
  • Checklist for leakage prevention based on pipeline.

Self-test questions

  • Why is manual fit_transform more error-prone than pipeline?
  • When adding new features, what needs to be updated in ColumnTransformer?
  • What is the biggest benefit of pipeline when deploying?