EDA & Data Analysis: thống kê mô tả, phát hiện outliers, feature correlation trên AWS
1. Exploratory Data Analysis (EDA)
EDA là bước phân tích dữ liệu ban đầu để hiểu structure, patterns, và anomalies trước khi modeling. SageMaker cung cấp nhiều tools để thực hiện EDA ở scale lớn.
2. AWS Tools cho Data Analysis
| Tool | Use Case | Interface |
|---|---|---|
| SageMaker Studio Notebooks | Interactive EDA, Python/R analysis | JupyterLab-based IDE |
| SageMaker Data Wrangler | Visual data prep, 300+ transforms, auto-insights | Drag-and-drop GUI |
| Amazon Athena | SQL queries on S3 data | SQL console |
| Amazon QuickSight | BI dashboards, executive reports | Visual BI tool |
| Amazon Redshift | Large-scale data warehousing, SQL analytics | SQL |
| AWS Glue DataBrew | No-code data profiling và cleaning recipes | Visual tool |
Exam tip: Data Wrangler = visual data prep cho ML (generates SageMaker Processing code). DataBrew = data analyst/BI (no ML context). QuickSight = BI dashboards for business users, không phải ML.
3. Data Quality Issues
Đề thi thường hỏi về nhận biết và xử lý các vấn đề chất lượng data phổ biến.
| Issue | Detection Method | Impact on Model |
|---|---|---|
| Missing Values | Null counts, missing rate per column | Errors, biased results |
| Outliers | Box plots, Z-score > 3, IQR method | Skewed weights, poor generalization |
| Class Imbalance | Class distribution histogram | Biased toward majority class |
| Feature Correlation | Correlation matrix, VIF score | Multicollinearity → unstable coefficients |
| Data Leakage | Features with suspiciously high correlation to target | Over-optimistic eval, fails in production |
| Distribution Skew | Histogram, skewness metric | Violated model assumptions |
3.1. Data Leakage — Critical Concept
Data leakage là khi information từ outside the training set rò rỉ vào features, khiến model có accuracy cao trong training nhưng thất bại khi production.
Common Data Leakage Patterns:
❌ Target leakage:
Feature "loan_default_flag" → predicting "credit_risk"
(feature derived from target)
❌ Future data leakage:
Using tomorrow's stock price to predict today's trade
❌ Train/test contamination:
Scaling data BEFORE splitting (test mean leaks into train)
✅ Correct approach:
Split data FIRST → fit scaler on train only → transform both
Exam tip: Always split before transforming. StandardScaler.fit() chỉ được gọi trên training set. Sau đó transform() trên cả train và test. Fit+transform trên toàn bộ dataset là data leakage.
4. Amazon Athena
Athena cho phép chạy SQL queries directly trên S3 without loading data vào database. Pay per scan — tối ưu bằng cách dùng Parquet + partitioning.
Cost Optimization Tips:
┌────────────────────────────────────────────────┐
│ Partition data by date/region/category: │
│ s3://bucket/data/year=2024/month=01/ │
│ → Query chỉ scan the required partitions │
│ │
│ Use columnar formats (Parquet/ORC): │
│ → Read only needed columns │
│ │
│ Compress data (Snappy, Gzip): │
│ → Reduce scan size → reduce cost │
└────────────────────────────────────────────────┘
5. Amazon QuickSight
QuickSight là BI service, không phải ML tool. Key feature: SPICE (in-memory engine) cho fast dashboards.
| Feature | Description |
|---|---|
| SPICE | Super-fast Parallel In-memory Calculation Engine — cached dataset |
| ML Insights | Built-in anomaly detection, forecasting trên dashboards |
| Q (NLQ) | Natural language queries — "show me sales by region last month" |
6. Cheat Sheet — Analysis Tools
| Scenario | Tool |
|---|---|
| Interactive Python EDA on large data | SageMaker Studio Notebooks |
| Visual no-code ML data prep | SageMaker Data Wrangler |
| SQL on S3 data (serverless) | Amazon Athena |
| Business dashboards và reporting | Amazon QuickSight |
| Large data warehouse SQL | Amazon Redshift |
| No-code data profiling recipes | AWS Glue DataBrew |
7. Practice Questions
Q1: A data scientist standardized features using the mean and standard deviation of the ENTIRE dataset before splitting into train/test sets. What problem does this cause?
- A) Model underfitting
- B) Slow training convergence
- C) Data leakage from test set statistics into training ✓
- D) Class imbalance
Explanation: Fitting a scaler on the entire dataset causes data leakage — the test set statistics (mean, std) influence the training data transformation. Always fit transformers on training data only, then apply the fitted transformer to both train and test sets.
Q2: A business analyst needs to create executive dashboards from S3 data with fast interactive visualizations. Which AWS service is BEST suited?
- A) Amazon SageMaker Studio
- B) Amazon Athena
- C) Amazon QuickSight ✓
- D) AWS Glue DataBrew
Explanation: Amazon QuickSight is the AWS BI service designed for business dashboards and visualizations with SPICE in-memory engine for fast interactive queries. SageMaker Studio is for ML development, Athena is SQL querying, DataBrew is data preparation.
Q3: A model trained on customer churn data has 99% training accuracy but performs poorly on production data. Investigation shows "days_since_last_call" is more predictive than expected. What is the MOST likely cause?
- A) Overfitting due to too many features
- B) Underfitting due to low model complexity
- C) Data leakage — the feature is derived from post-churn activity ✓
- D) Class imbalance
Explanation: This is classic target leakage — "days_since_last_call" may reflect churn behavior after the fact (customers call to cancel). This future information isn't available in production, causing the model to fail.