Introduction
You don't need to be a Python expert to learn ML, but you must be comfortable enough with tabular data. This article is a "crash course" that focuses on exactly what is needed most in ML: reading data, filtering, transforming, checking for missing data, and creating simple features.
Lesson objectives
- Read and explore data using Pandas
- Perform the most common operations before training the model
- Know when to use a notebook and when to separate code into separate files
1. Two libraries you will use constantly
NumPy
Used for array arithmetic and vectorization operations.
Pandas
Used for tabular data (DataFrame).
In most basic ML projects, you will spend more time in Pandas than in the model.
2. Create a first DataFrame
import pandas as pd
df = pd.DataFrame({
'dien_tich': [45, 60, 80, 120],
'so_phong': [1, 2, 3, 4],
'gia': [1.2, 1.8, 2.6, 4.1]
})
print(df)
The result is a table with rows and columns. In ML:
- each row is usually an observation
- each column is a feature or target
3. 5 Pandas actions to remember immediately
Quick view of data
df.head()
df.shape
df.columns
df.info()
df.describe()
Select column
df['dien_tich']
df[['dien_tich', 'so_phong']]
Filter the stream
df[df['so_phong'] >= 2]
Create new column
df['gia_m2'] = df['gia'] / df['dien_tich']
Arrange
df.sort_values('gia', ascending=False)
4. Read data from CSV
df = pd.read_csv('data/raw/houses.csv')
The first thing after reading data should always be:
print(df.head())
print(df.shape)
print(df.isnull().sum())
The goal is to answer three questions:
- How many rows and columns does the data have?
- Which column is missing data?
- Which column is number, which column is text?
5. What are missing values?
Missing values are missing data cells. For example, the customer does not declare their age, or the system does not record a certain field.
Check:
df.isnull().sum()
Basic handling:
- skip rows/columns if too many are missing
- fill in the mean, median or mode value
- added flags
is_missing
For example:
df['tuoi'] = df['tuoi'].fillna(df['tuoi'].median())
6. Categorical variables
Many non-numeric columns, for example:
- city
- service package type
- gender
ML models often do not work directly with plain text, so they need to be transformed. The most basic way is one-hot encoding.
pd.get_dummies(df, columns=['thanh_pho'], drop_first=True)
7. Groupby and aggregation
This is a very powerful skill for understanding data and creating features.
df.groupby('thanh_pho')['gia'].mean()
df.groupby('loai_khach')['doanh_thu'].agg(['mean', 'count'])
Application example:
- Average revenue by customer group
- number of purchases by city
- churn rate by service package
8. Merge data
In practice, data rarely resides in a single table.
customers = pd.read_csv('customers.csv')
orders = pd.read_csv('orders.csv')
df = customers.merge(orders, on='customer_id', how='left')
Important rule: after merging, always check the line number and newly generated null values.
9. Basic feature engineering
Feature engineering is creating additional columns to help the model learn better.
For example:
gia_m2 = gia / dien_tichthoi_gian_su_dung = ngay_hien_tai - ngay_dang_kytong_chi_tieu_30_ngay
Good features often come from understanding the problem rather than from model tricks.
10. Notebook or script?
Use a notebook when:
- explore data
- draw graphs
- quickly try out ideas
Use script / module when:
- code is repeated many times
- Needs reuse
- want the pipeline to be clean and easy to maintain
Practical principles:
- notebook to explore
src/to produce code
Short example: First EDA
import pandas as pd
df = pd.read_csv('data/raw/customers.csv')
print(df.head())
print(df.shape)
print(df.isnull().sum())
print(df['plan'].value_counts())
df['monthly_spend'] = df['total_spend'] / df['months_active']
print(df[['monthly_spend', 'churn']].head())
Practice exercises
- Read any CSV file and run
head,info,describe. - Create at least 2 new features from the original data.
- Write 3 sentences describing what you learned from EDA.
Common mistakes
- Edit data without understanding what that column means.
- Create features using information from the future.
- Merge finished but did not check the line number.
Completion criteria
- Can read CSV into DataFrame
- Can perform basic filter, groupby, and merge
- Create a new meaningful feature for the problem