Chuyển đến nội dung chính

Bài 1: Data Repositories & Ingestion — S3, Kinesis, Glue

S3 data lake cho ML. Kinesis Data Streams/Firehose cho streaming ingestion. AWS Glue ETL jobs và Data Catalog. Lake Formation. Data Wrangler. Chiến lược lưu trữ: Parquet, ORC, CSV, JSON.

AWS ML Data Repositories & Ingestion

Data Repositories & Ingestion: S3, Kinesis, Glue và Lake Formation trong ML pipeline

1. Tổng quan Data Engineering trong MLS-C01

Domain Data Engineering chiếm 20% đề thi MLS-C01. Đây là phần bắt buộc phải nắm vững — đề thi thường hỏi "Which service should be used to ingest/store/transform data for ML?"

Exam tip: Phần lớn câu hỏi Data Engineering sẽ cho một scenario và hỏi service phù hợp. Key pattern: batch → S3 + Glue; streaming → Kinesis; structured/SQL → Athena; catalog → Glue Data Catalog.

2. Amazon S3 — ML Data Lake

Amazon S3 là nền tảng lưu trữ dữ liệu ML trên AWS. Mọi pipeline ML đều bắt đầu và kết thúc từ S3: training data, model artifacts, predictions.

2.1. S3 Storage Classes cho ML

Storage ClassUse CaseCost
S3 StandardActive training data, frequent accessCao nhất
S3 Intelligent-TieringMixed access patterns (tự động tier)Tự động tối ưu
S3 Standard-IABackup datasets, infrequent accessThấp hơn Standard
S3 Glacier Instant RetrievalArchived datasets, occasional retrievalThấp
S3 Glacier Deep ArchiveLong-term compliance archivesThấp nhất

2.2. File Formats for ML

FormatTypeBest ForCompression
ParquetColumnarAnalytics, large datasets, feature storesExcellent
ORCColumnarHive/EMR workloadsExcellent
CSVRow-basedSimple, SageMaker training inputPoor
JSONSemi-structuredNested data, APIsPoor
RecordIOBinarySageMaker Pipe Mode trainingGood

Exam tip: Khi đề hỏi về performance optimization cho large-scale training, đáp án thường là chuyển sang Parquet (columnar, compressed) và dùng Pipe Mode thay vì File Mode trong SageMaker.

S3 Data Lake Architecture for ML:

┌─────────────────────────────────────────────────────────┐
│                    Amazon S3 Buckets                     │
├──────────────┬──────────────┬──────────────┬────────────┤
│  Raw Zone    │ Processed    │  Features    │  Models    │
│  (landing)   │  Zone        │  Zone        │  & Output  │
│              │              │              │            │
│  CSV/JSON    │  Parquet/ORC │  Feature     │  Model     │
│  original    │  cleaned     │  Store       │  Artifacts │
│  data        │  transformed │  snapshots   │  Predictions│
└──────────────┴──────────────┴──────────────┴────────────┘
       ↑                ↑                ↑
   Kinesis          AWS Glue         SageMaker
  (streaming)        (ETL)           Processing

3. Amazon Kinesis — Streaming Ingestion

Kinesis là họ dịch vụ cho real-time data streaming. Đây là topic quan trọng trong đề thi — cần phân biệt rõ 4 services.

ServiceFunctionDestinationML Use Case
Kinesis Data Streams (KDS)Custom real-time processingCustom consumersReal-time feature engineering
Kinesis Data FirehoseManaged delivery (no code)S3, Redshift, ES, SplunkBatch loading to data lake
Kinesis Data AnalyticsSQL/Flink on streamsS3, RedshiftReal-time aggregations, anomaly detect
Kinesis Video StreamsVideo ingestionRekognition, SageMakerComputer vision pipelines

Exam tip: Câu hỏi phổ biến: "IoT sensors gửi data liên tục, cần store vào S3 cho ML training mà không cần custom code?" → Kinesis Data Firehose (managed, no code). "Cần xử lý real-time với custom logic?" → Kinesis Data Streams.

3.1. KDS Shards & Capacity

Kinesis Data Streams Capacity:

┌─────────────────────────────────────────────┐
│  Each Shard:                                │
│  • Ingest:  1 MB/s OR 1,000 records/s       │
│  • Read:    2 MB/s                          │
│  • Retention: 24 hours (default) → 7 days  │
└─────────────────────────────────────────────┘

Stream with N shards:
• Total ingest: N × 1 MB/s
• Total read:   N × 2 MB/s

4. AWS Glue — ETL for ML

AWS Glue là fully managed ETL service. Trong ML pipeline, Glue dùng để transform và clean data trước khi đưa vào training.

4.1. Glue Components

ComponentFunction
Glue Data CatalogCentral metadata repository — schemas, tables, partitions
Glue CrawlersAuto-discover schema từ S3/RDS/Redshift và populate Data Catalog
Glue ETL JobsSpark-based transformation jobs (Python/Scala)
Glue DataBrewNo-code visual data preparation (250+ pre-built transforms)
Glue StudioVisual ETL job builder (drag-and-drop)

Exam tip: Glue Data Catalog là metadata store chung cho Athena, EMR, Redshift Spectrum. Khi đề hỏi "centralized schema management" → Glue Data Catalog. Khi hỏi "no-code data cleaning" → Glue DataBrew.

5. AWS Lake Formation

Lake Formation build trên S3 + Glue để management data lake security và governance. Key feature: column-level và row-level access control.

Lake Formation Architecture:

  IAM Users ──→ Lake Formation ──→ S3 Data Lake
  IAM Roles       (Security         (Raw/Processed)
                   & Governance)
                       ↓
                  Column/Row
                  Level Access
                  Control

6. Cheat Sheet — Data Ingestion Services

ScenarioService
Streaming → S3 với no-codeKinesis Data Firehose
Real-time processing với custom logicKinesis Data Streams
SQL on streaming dataKinesis Data Analytics (Flink)
Batch ETL Spark-basedAWS Glue ETL Jobs
No-code visual data prepGlue DataBrew
Schema discovery from S3Glue Crawlers + Data Catalog
SQL queries on S3Amazon Athena
Data lake governanceAWS Lake Formation
Large-scale Spark/HadoopAmazon EMR

7. Practice Questions

Q1: A company wants to ingest IoT sensor data into Amazon S3 for ML training. The data arrives continuously and no custom processing is required. Which service is the MOST cost-effective?

  • A) Amazon Kinesis Data Streams with a Lambda consumer
  • B) Amazon Kinesis Data Firehose ✓
  • C) Amazon EMR with Spark Streaming
  • D) AWS Glue ETL jobs on a schedule

Explanation: Kinesis Data Firehose is fully managed and requires no custom code — it directly delivers streaming data to S3, Redshift, or Elasticsearch. Data Streams requires custom consumers, EMR is heavy lift, and Glue is for batch ETL.

Q2: A data engineer wants to query raw CSV files in S3 using SQL without loading them into a database. Which service should be used?

  • A) Amazon RDS
  • B) Amazon DynamoDB
  • C) Amazon Athena ✓
  • D) Amazon Redshift

Explanation: Amazon Athena is serverless and allows SQL queries directly on S3 data without loading. It reads files in-place and supports formats like CSV, Parquet, ORC, JSON.

Q3: Which file format provides the BEST performance for columnar analytics queries on large ML datasets stored in Amazon S3?

  • A) CSV
  • B) JSON
  • C) XML
  • D) Apache Parquet ✓

Explanation: Parquet is a columnar format with excellent compression and predicate pushdown support. Columnar formats allow reading only the required columns, dramatically reducing I/O for analytical queries.