Chuyển đến nội dung chính

Lesson 12: Data Storage Patterns - Object Storage, Data Lake, Time-Series

Object Storage (S3) architecture and use cases. Data Lake vs Data Warehouse vs Lakehouse. Time-Series databases for metrics and IoT. Search engines (Elasticsearch). Polyglot persistence in practice.

🏗️ Architecture — Lesson 12 Lesson 12: Data Storage Patterns - Objects Storage, Data Lake, Time-Series

System Architecture: From Zero to Hero

Part 3: Database Architecture & Data Management

xdev.asia

Introduction

Not all data is suitable for a relational database. Photos, videos, logs, metrics, search indexes — each type of data needs the right storage engine. This article explores specialized storage systems.


1. Object Storage

1.1 What is Object Storage?

Traditional File System:          Object Storage:
  /home/                           ┌─────────────────────┐
    /images/                       │     Flat Namespace   │
      /2024/                       │                      │
        /01/                       │  key → object (blob) │
          photo.jpg                │  key → object (blob) │
                                   │  key → object (blob) │
  Hierarchical                     └─────────────────────┘
  Directories, inodes              Flat, HTTP API
  Mount required                   REST access

1.2 S3-Compatible Architecture

Client: PUT /bucket/images/photo.jpg
         │
         ▼
  ┌──────────────┐
  │  API Gateway  │  ← REST: GET, PUT, DELETE, LIST
  └──────┬───────┘
         ▼
  ┌──────────────┐
  │  Metadata    │  ← Key → location mapping
  │  Service     │  ← ACL, versioning, lifecycle
  └──────┬───────┘
         ▼
  ┌──────────────────────────────────┐
  │  Data Layer (Distributed)       │
  │  ┌──────┐ ┌──────┐ ┌──────┐    │
  │  │Node 1│ │Node 2│ │Node 3│    │ ← 3x replication
  │  └──────┘ └──────┘ └──────┘    │
  └──────────────────────────────────┘

1.3 Use Cases

Use CaseExample
Static assetsImages, CSS, JS → CDN origin
BackupsDatabase dumps, log archives
Data Lake storageRaw data for analytics
MediaVideos, audio files
ML artifactsModel files, training datasets

1.4 Presigned URLs Pattern

Vấn đề: User upload file 100MB qua API server → bottleneck

Giải pháp: Presigned URL - upload trực tiếp lên S3

  Client → API: "Tôi muốn upload avatar.jpg"
  API → S3: GeneratePresignedURL(PUT, bucket, key, 15min)
  API → Client: "Upload tại URL này (hết hạn 15 phút)"
  Client → S3: PUT trực tiếp (không qua API server)
  S3 → SNS/Lambda: Trigger post-processing
# Python - Generate presigned URL
import boto3

s3 = boto3.client('s3')
url = s3.generate_presigned_url(
    'put_object',
    Params={'Bucket': 'my-bucket', 'Key': 'uploads/avatar.jpg'},
    ExpiresIn=900  # 15 minutes
)

2. Data Lake vs Data Warehouse vs Lakehouse

2.1 Comparison

Data Warehouse:           Data Lake:              Lakehouse:
┌──────────────────┐   ┌──────────────────┐   ┌──────────────────┐
│ Structured only  │   │ Raw everything   │   │ Best of both     │
│ Schema-on-WRITE  │   │ Schema-on-READ   │   │ Schema evolution │
│ SQL queries      │   │ Any processing   │   │ SQL + ML + Stream│
│ Expensive        │   │ Cheap storage    │   │ Open formats     │
│                  │   │                  │   │                  │
│ Redshift, BQ,    │   │ S3 + Spark,     │   │ Delta Lake,      │
│ Snowflake        │   │ HDFS            │   │ Apache Iceberg   │
└──────────────────┘   └──────────────────┘   └──────────────────┘

2.2 Data Lake Architecture

Data Sources              Ingestion           Storage Layers
┌──────────┐             ┌─────────┐        ┌─────────────────┐
│ APIs     │────────────►│         │───────►│ Raw Zone        │
│ DBs      │────────────►│ Kafka/  │───────►│ (landing)       │
│ Logs     │────────────►│ Spark   │        ├─────────────────┤
│ IoT      │────────────►│ Airflow │───────►│ Cleaned Zone    │
│ Files    │────────────►│         │        │ (validated)     │
└──────────┘             └─────────┘        ├─────────────────┤
                                            │ Curated Zone    │
                                            │ (analytics-ready│
                                            └────────┬────────┘
                                                     │
                                            ┌────────▼────────┐
                                            │ Consumption     │
                                            │ BI, ML, Reports │
                                            └─────────────────┘

3. Time-Series Databases

3.1 Time-Series Data characteristics

Đặc điểm:
  - Append-mostly (ít update, delete)
  - Time-ordered
  - High write throughput
  - Recent data truy vấn nhiều hơn
  - Aggregations: avg, sum, percentile theo time window

Ví dụ:
  Metrics: CPU usage mỗi 10s, 100 servers → 864K points/ngày
  IoT: 10K sensors, mỗi giây 1 reading → 864M points/ngày

3.2 Optimizations

1. Time-based partitioning:
   ┌──────────┐ ┌──────────┐ ┌──────────┐
   │ Jan 2024 │ │ Feb 2024 │ │ Mar 2024 │
   └──────────┘ └──────────┘ └──────────┘
   → Drop old partitions thay vì DELETE

2. Columnar storage:
   Row:    [time, cpu, mem, disk] [time, cpu, mem, disk] ...
   Column: [time, time, time...] [cpu, cpu, cpu...] [mem, mem, mem...]
   → Compression tốt hơn (cùng type data)
   → Aggregation nhanh hơn

3. Downsampling:
   Raw: 1 point/giây (86400/ngày)
   1h:  1 point/giờ (24/ngày)
   1d:  1 point/ngày
   → Giữ raw 7 ngày, 1h 30 ngày, 1d mãi mãi

3.3 Popular TSDBs

DatabaseArchitectureAdvantages
InfluxDBStandalone/ClusterEasy setup, InfluxQL
TimescaleDBPostgreSQL extensionFull SQL, hypertables
PrometheusPull-based metricsK8s native, PromQL
ClickHouseColumnar OLAPExtremely fast aggregation
VictoriaMetricsPrometheus-compatibleStorage efficiency

3.4 PromQL example

# CPU usage trung bình 5 phút
avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) by (instance)

# Request rate per second (QPS)
rate(http_requests_total[5m])

# 99th percentile latency
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))

# Alert: CPU > 80% trong 5 phút
avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) < 0.2

4. Search Engines

4.1 Elasticsearch Architecture

Cluster
├── Node 1 (Master + Data)
│   ├── Index: products
│   │   ├── Shard 0 (Primary)
│   │   └── Shard 2 (Replica)
│   └── Index: logs
│       └── Shard 1 (Primary)
│
├── Node 2 (Data)
│   ├── Index: products
│   │   ├── Shard 1 (Primary)
│   │   └── Shard 0 (Replica)
│   └── Index: logs
│       └── Shard 0 (Primary)
│
└── Node 3 (Data)
    ├── Index: products
    │   └── Shard 2 (Primary)
    └── Index: logs
        └── Shard 1 (Replica)

4.2 Inverted Index

Documents:
  Doc 1: "Kiến trúc hệ thống phân tán"
  Doc 2: "Thiết kế hệ thống chat"
  Doc 3: "Kiến trúc microservices"

Inverted Index:
  "kiến trúc"  → [Doc 1, Doc 3]
  "hệ thống"   → [Doc 1, Doc 2]
  "phân tán"   → [Doc 1]
  "thiết kế"   → [Doc 2]
  "chat"       → [Doc 2]
  "microservices" → [Doc 3]

Query: "kiến trúc hệ thống"
  → "kiến trúc" ∩ "hệ thống" = [Doc 1]
  → Score: Doc 1 (match cả 2) > Doc 2, Doc 3

5. Polyglot Persistence

5.1 E-Commerce Example

┌─────────────────────────────────────────────────┐
│                 E-Commerce App                   │
├─────────┬──────────┬──────────┬────────┬────────┤
│ Users   │ Products │ Orders   │ Search │ Cache  │
│         │ Catalog  │          │        │        │
│PostgreSQL│ MongoDB │PostgreSQL│Elastic │ Redis  │
│         │          │          │Search  │        │
│Relational│Document │ACID txns │Full-text│Session│
│Schema   │Flexible  │Consistent│Scoring │Cart   │
│Joins    │Nested    │Foreign   │Facets  │Rate   │
│         │attrs     │keys      │        │Limit  │
└─────────┴──────────┴──────────┴────────┴────────┘
         │                                │
    ┌────▼────┐                    ┌──────▼──────┐
    │ S3      │                    │ InfluxDB    │
    │ Images  │                    │ Metrics     │
    │ Files   │                    │ Monitoring  │
    └─────────┘                    └─────────────┘

Summary

Storage TypeBest ForExample
RDBMSStructured, ACIDUsers, Orders, Billing
DocumentDBFlexible schemaProduct catalog, CMS
Key-ValueSimple, fastCache, Sessions
Object StorageFiles, blobsImages, Videos, Backups
Time SeriesMetrics, IoTMonitoring, Sensor data
Search EngineFull-text searchProduct search, Logs
GraphDBRelationshipsSocial networks, Fraud
Data LakeRaw analyticsML, BI, Reports

Exercises

  1. Storage Design: Design storage architecture for Healthcare system: patient records (HIPAA compliant), medical images (DICOM), vital signs monitoring, prescription search. Which database to choose for each use case?

  2. Time-Series: IoT system has 50K sensors, each sensor sends data every 5 seconds. Calculate storage needed for 1 year. Design retention policy.

  3. Search Architecture: E-commerce has 10M products. Search system design supports: full-text search, filters (price, brand, category), faceted navigation, auto-suggest.