Chuyển đến nội dung chính

Lesson 3: Python/Pandas crash course for ML

DataFrame, filtering, groupby, merge, basic missing data handling and fast EDA for those not familiar with Python data.

🧠 AI & ML — Lesson 2 Lesson 3: Python/Pandas crash course for ML

Machine Learning: From Basics to Advanced

Part 0: Getting started for newbies (Week 0)

xdev.asia

Introduction

You don't need to be a Python expert to learn ML, but you must be comfortable enough with tabular data. This article is a "crash course" that focuses on exactly what is needed most in ML: reading data, filtering, transforming, checking for missing data, and creating simple features.

Lesson objectives

  • Read and explore data using Pandas
  • Perform the most common operations before training the model
  • Know when to use a notebook and when to separate code into separate files

1. Two libraries you will use constantly

NumPy

Used for array arithmetic and vectorization operations.

Pandas

Used for tabular data (DataFrame).

In most basic ML projects, you will spend more time in Pandas than in the model.

2. Create a first DataFrame

import pandas as pd

df = pd.DataFrame({
    'dien_tich': [45, 60, 80, 120],
    'so_phong': [1, 2, 3, 4],
    'gia': [1.2, 1.8, 2.6, 4.1]
})

print(df)

The result is a table with rows and columns. In ML:

  • each row is usually an observation
  • each column is a feature or target

3. 5 Pandas actions to remember immediately

Quick view of data

df.head()
df.shape
df.columns
df.info()
df.describe()

Select column

df['dien_tich']
df[['dien_tich', 'so_phong']]

Filter the stream

df[df['so_phong'] >= 2]

Create new column

df['gia_m2'] = df['gia'] / df['dien_tich']

Arrange

df.sort_values('gia', ascending=False)

4. Read data from CSV

df = pd.read_csv('data/raw/houses.csv')

The first thing after reading data should always be:

print(df.head())
print(df.shape)
print(df.isnull().sum())

The goal is to answer three questions:

  1. How many rows and columns does the data have?
  2. Which column is missing data?
  3. Which column is number, which column is text?

5. What are missing values?

Missing values ​​are missing data cells. For example, the customer does not declare their age, or the system does not record a certain field.

Check:

df.isnull().sum()

Basic handling:

  • skip rows/columns if too many are missing
  • fill in the mean, median or mode value
  • added flags is_missing

For example:

df['tuoi'] = df['tuoi'].fillna(df['tuoi'].median())

6. Categorical variables

Many non-numeric columns, for example:

  • city
  • service package type
  • gender

ML models often do not work directly with plain text, so they need to be transformed. The most basic way is one-hot encoding.

pd.get_dummies(df, columns=['thanh_pho'], drop_first=True)

7. Groupby and aggregation

This is a very powerful skill for understanding data and creating features.

df.groupby('thanh_pho')['gia'].mean()
df.groupby('loai_khach')['doanh_thu'].agg(['mean', 'count'])

Application example:

  • Average revenue by customer group
  • number of purchases by city
  • churn rate by service package

8. Merge data

In practice, data rarely resides in a single table.

customers = pd.read_csv('customers.csv')
orders = pd.read_csv('orders.csv')

df = customers.merge(orders, on='customer_id', how='left')

Important rule: after merging, always check the line number and newly generated null values.

9. Basic feature engineering

Feature engineering is creating additional columns to help the model learn better.

For example:

  • gia_m2 = gia / dien_tich
  • thoi_gian_su_dung = ngay_hien_tai - ngay_dang_ky
  • tong_chi_tieu_30_ngay

Good features often come from understanding the problem rather than from model tricks.

10. Notebook or script?

Use a notebook when:

  • explore data
  • draw graphs
  • quickly try out ideas

Use script / module when:

  • code is repeated many times
  • Needs reuse
  • want the pipeline to be clean and easy to maintain

Practical principles:

  • notebook to explore
  • src/ to produce code

Short example: First EDA

import pandas as pd

df = pd.read_csv('data/raw/customers.csv')

print(df.head())
print(df.shape)
print(df.isnull().sum())
print(df['plan'].value_counts())

df['monthly_spend'] = df['total_spend'] / df['months_active']
print(df[['monthly_spend', 'churn']].head())

Practice exercises

  1. Read any CSV file and run head, info, describe.
  2. Create at least 2 new features from the original data.
  3. Write 3 sentences describing what you learned from EDA.

Common mistakes

  • Edit data without understanding what that column means.
  • Create features using information from the future.
  • Merge finished but did not check the line number.

Completion criteria

  • Can read CSV into DataFrame
  • Can perform basic filter, groupby, and merge
  • Create a new meaningful feature for the problem