Chuyển đến nội dung chính

Lesson 1: Data Repositories & Ingestion — S3, Kinesis, Glue

S3 data lake for ML. Kinesis Data Streams/Firehose for streaming ingestion. AWS Glue ETL jobs and Data Catalog. Lake Formation. Data Wrangler. Storage strategies: Parquet, ORC, CSV, JSON.

AWS ML Data Repositories & Ingestion

Data Repositories & Ingestion: S3, Kinesis, Glue, and Lake Formation in the ML pipeline

1. Data Engineering Overview in MLS-C01

The Data Engineering domain accounts for 20% of the MLS-C01 exam. This is a must-know section — exam questions typically ask "Which service should be used to ingest/store/transform data for ML?"

Exam tip: Most Data Engineering questions present a scenario and ask for the appropriate service. Key pattern: batch → S3 + Glue; streaming → Kinesis; structured/SQL → Athena; catalog → Glue Data Catalog.

2. Amazon S3 — ML Data Lake

Amazon S3 is the foundational storage platform for ML data on AWS. Every ML pipeline starts and ends with S3: training data, model artifacts, predictions.

2.1. S3 Storage Classes for ML

Storage ClassUse CaseCost
S3 StandardActive training data, frequent accessHighest
S3 Intelligent-TieringMixed access patterns (auto-tiered)Auto-optimized
S3 Standard-IABackup datasets, infrequent accessLower than Standard
S3 Glacier Instant RetrievalArchived datasets, occasional retrievalLow
S3 Glacier Deep ArchiveLong-term compliance archivesLowest

2.2. File Formats for ML

FormatTypeBest ForCompression
ParquetColumnarAnalytics, large datasets, feature storesExcellent
ORCColumnarHive/EMR workloadsExcellent
CSVRow-basedSimple, SageMaker training inputPoor
JSONSemi-structuredNested data, APIsPoor
RecordIOBinarySageMaker Pipe Mode trainingGood

Exam tip: When the question asks about performance optimization for large-scale training, the answer is usually to switch to Parquet (columnar, compressed) and use Pipe Mode instead of File Mode in SageMaker.

S3 Data Lake Architecture for ML:

┌─────────────────────────────────────────────────────────┐
│                    Amazon S3 Buckets                     │
├──────────────┬──────────────┬──────────────┬────────────┤
│  Raw Zone    │ Processed    │  Features    │  Models    │
│  (landing)   │  Zone        │  Zone        │  & Output  │
│              │              │              │            │
│  CSV/JSON    │  Parquet/ORC │  Feature     │  Model     │
│  original    │  cleaned     │  Store       │  Artifacts │
│  data        │  transformed │  snapshots   │  Predictions│
└──────────────┴──────────────┴──────────────┴────────────┘
       ↑                ↑                ↑
   Kinesis          AWS Glue         SageMaker
  (streaming)        (ETL)           Processing

3. Amazon Kinesis — Streaming Ingestion

Kinesis is the family of services for real-time data streaming. This is an important exam topic — you need to clearly distinguish the 4 services.

ServiceFunctionDestinationML Use Case
Kinesis Data Streams (KDS)Custom real-time processingCustom consumersReal-time feature engineering
Kinesis Data FirehoseManaged delivery (no code)S3, Redshift, ES, SplunkBatch loading to data lake
Kinesis Data AnalyticsSQL/Flink on streamsS3, RedshiftReal-time aggregations, anomaly detect
Kinesis Video StreamsVideo ingestionRekognition, SageMakerComputer vision pipelines

Exam tip: Common question: "IoT sensors continuously send data, need to store in S3 for ML training with no custom code?" → Kinesis Data Firehose (managed, no code). "Need real-time processing with custom logic?" → Kinesis Data Streams.

3.1. KDS Shards & Capacity

Kinesis Data Streams Capacity:

┌─────────────────────────────────────────────┐
│  Each Shard:                                │
│  • Ingest:  1 MB/s OR 1,000 records/s       │
│  • Read:    2 MB/s                          │
│  • Retention: 24 hours (default) → 7 days  │
└─────────────────────────────────────────────┘

Stream with N shards:
• Total ingest: N × 1 MB/s
• Total read:   N × 2 MB/s

4. AWS Glue — ETL for ML

AWS Glue is a fully managed ETL service. In ML pipelines, Glue is used to transform and clean data before feeding it into training.

4.1. Glue Components

ComponentFunction
Glue Data CatalogCentral metadata repository — schemas, tables, partitions
Glue CrawlersAuto-discover schemas from S3/RDS/Redshift and populate the Data Catalog
Glue ETL JobsSpark-based transformation jobs (Python/Scala)
Glue DataBrewNo-code visual data preparation (250+ pre-built transforms)
Glue StudioVisual ETL job builder (drag-and-drop)

Exam tip: Glue Data Catalog is the shared metadata store for Athena, EMR, and Redshift Spectrum. When the question asks "centralized schema management" → Glue Data Catalog. When it asks "no-code data cleaning" → Glue DataBrew.

5. AWS Lake Formation

Lake Formation builds on S3 + Glue to provide data lake security and governance. Key feature: column-level and row-level access control.

Lake Formation Architecture:

  IAM Users ──→ Lake Formation ──→ S3 Data Lake
  IAM Roles       (Security         (Raw/Processed)
                   & Governance)
                       ↓
                  Column/Row
                  Level Access
                  Control

6. Cheat Sheet — Data Ingestion Services

ScenarioService
Streaming → S3 with no codeKinesis Data Firehose
Real-time processing with custom logicKinesis Data Streams
SQL on streaming dataKinesis Data Analytics (Flink)
Batch ETL Spark-basedAWS Glue ETL Jobs
No-code visual data prepGlue DataBrew
Schema discovery from S3Glue Crawlers + Data Catalog
SQL queries on S3Amazon Athena
Data lake governanceAWS Lake Formation
Large-scale Spark/HadoopAmazon EMR

7. Practice Questions

Q1: A company wants to ingest IoT sensor data into Amazon S3 for ML training. The data arrives continuously and no custom processing is required. Which service is the MOST cost-effective?

  • A) Amazon Kinesis Data Streams with a Lambda consumer
  • B) Amazon Kinesis Data Firehose ✓
  • C) Amazon EMR with Spark Streaming
  • D) AWS Glue ETL jobs on a schedule

Explanation: Kinesis Data Firehose is fully managed and requires no custom code — it directly delivers streaming data to S3, Redshift, or Elasticsearch. Data Streams requires custom consumers, EMR is heavy lift, and Glue is for batch ETL.

Q2: A data engineer wants to query raw CSV files in S3 using SQL without loading them into a database. Which service should be used?

  • A) Amazon RDS
  • B) Amazon DynamoDB
  • C) Amazon Athena ✓
  • D) Amazon Redshift

Explanation: Amazon Athena is serverless and allows SQL queries directly on S3 data without loading. It reads files in-place and supports formats like CSV, Parquet, ORC, JSON.

Q3: Which file format provides the BEST performance for columnar analytics queries on large ML datasets stored in Amazon S3?

  • A) CSV
  • B) JSON
  • C) XML
  • D) Apache Parquet ✓

Explanation: Parquet is a columnar format with excellent compression and predicate pushdown support. Columnar formats allow reading only the required columns, dramatically reducing I/O for analytical queries.