Chuyển đến nội dung chính

Lesson 11: Missing Values, Categorical Variables, Feature Engineering

Actual data processing process: missing, encoding, scaling, outlier handling and basic feature crosses.

🧠 AI & ML — Lesson 10 Lesson 11: Missing Values, Categorical Variables, Feature Engineering

Machine Learning: From Basics to Advanced

Part 2: Industry standard workflow

xdev.asia

Introduction

Real-life data is rarely clean. Columns are missing values, text has strange symbols, categories have too many levels, features are manually created inconsistently. This article helps you handle these tasks in a repeatable and error-free way.

Lesson objectives

  • Handle missing values for numeric and categorical.
  • Encode categorical variables properly.
  • Understand what feature engineering is and when to stop.

Missing values: does not mean missing them all

Missing data is sometimes a signal. Before filling, ask: why is the data missing, is it missing randomly or systematically, and should we create additional columns to mark missing or not?

Common treatment

  • Numeric: median is usually safer than mean when there is an outlier.
  • Categorical: fill with the most common value or label it Unknown.
  • Missing indicator: useful when the missing itself is a signal.

Feature engineering is practical

Prioritize features with clear business logic such as total spending in the last 30 days, number of support tickets per month or usage rate above limit.

Frame code

from sklearn.impute import SimpleImputer
from sklearn.preprocessing import OneHotEncoder

num_imputer = SimpleImputer(strategy='median')
cat_imputer = SimpleImputer(strategy='most_frequent')
encoder = OneHotEncoder(handle_unknown='ignore')

Common mistakes

  • Fill missing before splitting data.
  • Create a feature that sees the future.
  • One-hot has too many categories, making the feature space uselessly bloated.

Practice exercises

  • Choose a tabular dataset that has both numeric and categorical.
  • Try 2 ways to fill in missing numbers for numeric.
  • Create 3 new features with clear business explanations.

Completion criteria

  • Know how to choose the fill missing strategy for each column type.
  • Using one-hot encoding does not cause errors when encountering new categories.
  • Create new features without causing leakage.

Practice step by step (advanced)

  1. Prepare data profiling report (missing rate, cardinality, outlier).
  2. Create 2 versions of missing handling for comparison.
  3. Encode classification using one-hot and compare with target safe encoding.
  4. Add 3 business features with clearly explained origins.
  5. Evaluate the impact of each step on the final metric.

Artifact should be submitted

  • Data quality table before/after processing.
  • List of new features and reasons for their existence.
  • Experiment log following each preprocessing step.

Self-test questions

  • Non-random missingness needs to be handled differently?
  • When does one-hot become ineffective?
  • How to prove that the new feature has real value?