Chuyển đến nội dung chính

Bài 3: Data Analysis & Visualization

EDA trên SageMaker notebooks. Amazon Athena cho SQL analytics. Amazon QuickSight cho BI dashboards. Phát hiện data quality issues. Detect class imbalance, outliers, correlations, data drift.

Exploratory Data Analysis trên AWS

EDA & Data Analysis: thống kê mô tả, phát hiện outliers, feature correlation trên AWS

1. Exploratory Data Analysis (EDA)

EDA là bước phân tích dữ liệu ban đầu để hiểu structure, patterns, và anomalies trước khi modeling. SageMaker cung cấp nhiều tools để thực hiện EDA ở scale lớn.

2. AWS Tools cho Data Analysis

ToolUse CaseInterface
SageMaker Studio NotebooksInteractive EDA, Python/R analysisJupyterLab-based IDE
SageMaker Data WranglerVisual data prep, 300+ transforms, auto-insightsDrag-and-drop GUI
Amazon AthenaSQL queries on S3 dataSQL console
Amazon QuickSightBI dashboards, executive reportsVisual BI tool
Amazon RedshiftLarge-scale data warehousing, SQL analyticsSQL
AWS Glue DataBrewNo-code data profiling và cleaning recipesVisual tool

Exam tip: Data Wrangler = visual data prep cho ML (generates SageMaker Processing code). DataBrew = data analyst/BI (no ML context). QuickSight = BI dashboards for business users, không phải ML.

3. Data Quality Issues

Đề thi thường hỏi về nhận biết và xử lý các vấn đề chất lượng data phổ biến.

IssueDetection MethodImpact on Model
Missing ValuesNull counts, missing rate per columnErrors, biased results
OutliersBox plots, Z-score > 3, IQR methodSkewed weights, poor generalization
Class ImbalanceClass distribution histogramBiased toward majority class
Feature CorrelationCorrelation matrix, VIF scoreMulticollinearity → unstable coefficients
Data LeakageFeatures with suspiciously high correlation to targetOver-optimistic eval, fails in production
Distribution SkewHistogram, skewness metricViolated model assumptions

3.1. Data Leakage — Critical Concept

Data leakage là khi information từ outside the training set rò rỉ vào features, khiến model có accuracy cao trong training nhưng thất bại khi production.

Common Data Leakage Patterns:

❌ Target leakage:
   Feature "loan_default_flag" → predicting "credit_risk"
   (feature derived from target)

❌ Future data leakage:
   Using tomorrow's stock price to predict today's trade

❌ Train/test contamination:
   Scaling data BEFORE splitting (test mean leaks into train)

✅ Correct approach:
   Split data FIRST → fit scaler on train only → transform both

Exam tip: Always split before transforming. StandardScaler.fit() chỉ được gọi trên training set. Sau đó transform() trên cả train và test. Fit+transform trên toàn bộ dataset là data leakage.

4. Amazon Athena

Athena cho phép chạy SQL queries directly trên S3 without loading data vào database. Pay per scan — tối ưu bằng cách dùng Parquet + partitioning.

Cost Optimization Tips:
┌────────────────────────────────────────────────┐
│  Partition data by date/region/category:       │
│  s3://bucket/data/year=2024/month=01/          │
│  → Query chỉ scan the required partitions      │
│                                                │
│  Use columnar formats (Parquet/ORC):           │
│  → Read only needed columns                   │
│                                               │
│  Compress data (Snappy, Gzip):                │
│  → Reduce scan size → reduce cost             │
└────────────────────────────────────────────────┘

5. Amazon QuickSight

QuickSight là BI service, không phải ML tool. Key feature: SPICE (in-memory engine) cho fast dashboards.

FeatureDescription
SPICESuper-fast Parallel In-memory Calculation Engine — cached dataset
ML InsightsBuilt-in anomaly detection, forecasting trên dashboards
Q (NLQ)Natural language queries — "show me sales by region last month"

6. Cheat Sheet — Analysis Tools

ScenarioTool
Interactive Python EDA on large dataSageMaker Studio Notebooks
Visual no-code ML data prepSageMaker Data Wrangler
SQL on S3 data (serverless)Amazon Athena
Business dashboards và reportingAmazon QuickSight
Large data warehouse SQLAmazon Redshift
No-code data profiling recipesAWS Glue DataBrew

7. Practice Questions

Q1: A data scientist standardized features using the mean and standard deviation of the ENTIRE dataset before splitting into train/test sets. What problem does this cause?

  • A) Model underfitting
  • B) Slow training convergence
  • C) Data leakage from test set statistics into training ✓
  • D) Class imbalance

Explanation: Fitting a scaler on the entire dataset causes data leakage — the test set statistics (mean, std) influence the training data transformation. Always fit transformers on training data only, then apply the fitted transformer to both train and test sets.

Q2: A business analyst needs to create executive dashboards from S3 data with fast interactive visualizations. Which AWS service is BEST suited?

  • A) Amazon SageMaker Studio
  • B) Amazon Athena
  • C) Amazon QuickSight ✓
  • D) AWS Glue DataBrew

Explanation: Amazon QuickSight is the AWS BI service designed for business dashboards and visualizations with SPICE in-memory engine for fast interactive queries. SageMaker Studio is for ML development, Athena is SQL querying, DataBrew is data preparation.

Q3: A model trained on customer churn data has 99% training accuracy but performs poorly on production data. Investigation shows "days_since_last_call" is more predictive than expected. What is the MOST likely cause?

  • A) Overfitting due to too many features
  • B) Underfitting due to low model complexity
  • C) Data leakage — the feature is derived from post-churn activity ✓
  • D) Class imbalance

Explanation: This is classic target leakage — "days_since_last_call" may reflect churn behavior after the fact (customers call to cancel). This future information isn't available in production, causing the model to fail.