Chuyển đến nội dung chính

Bài 7: Model Deployment — Endpoints & Inference

Real-time Endpoints, Batch Transform, Async Inference, Serverless Inference. Multi-Model Endpoints, Inference Pipeline. Elastic Inference, SageMaker Neo (edge deployment). A/B Testing Production Variants.

SageMaker Model Deployment Options

SageMaker Deployment: Real-time Endpoint, Serverless, Async Inference, và Batch Transform

1. SageMaker Deployment Options

SageMaker cung cấp nhiều inference patterns — mỗi loại phù hợp với workload khác nhau. Phần này thường có 5-8 câu trong đề thi MLS-C01.

Exam tip: Key decision factors: latency requirement, volume, cost, payload size. Map these to: Real-time (low latency) → Async (large payload) → Serverless (sporadic) → Batch (no latency need).

Deployment TypeLatencyThroughputCost ModelBest For
Real-time EndpointMillisecondsHighAlways-on (pay per hour)Interactive apps, APIs
Serverless InferenceSeconds (cold start)VariablePay-per-invocationSporadic, unpredictable traffic
Async InferenceMinutesHigh queuedPay per processingLarge payloads, non-urgent
Batch TransformNo real-timeVery highPay per batch jobScheduled offline predictions

2. Real-time Inference

Standard deployment — persistent endpoint chạy constantly, responds synchronously.

Real-time Endpoint Architecture:

Client ──→ HTTPS Request
              ↓
      SageMaker Endpoint
      ┌────────────────┐
      │  Model Server  │  ← Instance running 24/7
      │  (TorchServe,  │
      │  TensorFlow    │
      │  Serving, etc) │
      └────────────────┘
              ↓
         Response (ms)

2.1. Auto Scaling cho Endpoints

Endpoints có thể scale dựa trên InvocationsPerInstance metric qua Application Auto Scaling.

3. Serverless Inference

Phù hợp khi traffic không đều, khó dự đoán. AWS tự động scale, kể cả về 0 khi không có traffic.

FeatureDetail
Cold start latency~1-2 seconds (đầu tiên sau thời gian nhàn rỗi)
Memory config1 GB → 6 GB
Max payload4 MB
PricingPer inference requests + processing time

4. Async Inference

Phù hợp cho large media files, long processing time. Request được queued, response lưu vào S3.

Async Inference Flow:

Client ──→ Upload payload to S3 ──→ Invoke Endpoint
                                          ↓
                                   Queue Request
                                          ↓
                               Process when instance available
                                          ↓
                                   Save output to S3
                                          ↓
                           SNS Notification → Client
FeatureDetail
Max payload1 GB (vs 6 MB for real-time)
Auto-scale to 0Yes — scales down when queue empty
ResponseS3 output path + SNS notification

5. Batch Transform

Chạy predictions trên toàn bộ dataset theo lịch. Không có endpoint — chỉ chạy khi cần.

Batch Transform:

Input S3 ──→ Batch Transform Job ──→ Output S3
  (CSV/       (ephemeral compute)      (CSV/JSON
  JSON/                                predictions)
  Parquet)           ↑
               No persistent endpoint
               Pay only when running

6. Multi-Model Endpoints (MME)

MME cho phép host nhiều models trên một endpoint, giảm chi phí inference infrastructure.

FeatureDetail
Cost savingMột endpoint phục vụ hàng ngàn models
Dynamic loadingModels loaded into memory on-demand, cached
Use caseSaaS multi-tenant với model per customer

7. SageMaker Neo — Edge Deployment

SageMaker Neo compiles models và optimize cho specific hardware (edge devices, mobile).

Neo Workflow:

Trained Model (S3)
       ↓
  Neo Compiler
  (optimizes for target hardware)
       ↓
 Optimized Model
       ↓
    ├── Deploy to IoT Greengrass (edge)
    ├── Deploy to ARM devices
    └── Deploy to mobile (Android/iOS)

8. Cheat Sheet — Deployment Decision

ScenarioDeployment Type
Mobile app, real-time response (<100ms)Real-time Endpoint
Traffic is sporadic (few req/hour)Serverless Inference
Video/audio processing (large files)Async Inference
Nightly predictions on full datasetBatch Transform
Thousands of customer-specific modelsMulti-Model Endpoints
IoT edge device deploymentSageMaker Neo + Greengrass

9. Practice Questions

Q1: A company runs an e-commerce chatbot that requires sub-100ms response times during peak shopping hours. Which SageMaker inference type should they use?

  • A) Batch Transform
  • B) Async Inference
  • C) Serverless Inference
  • D) Real-time Endpoint ✓

Explanation: Real-time Endpoints provide persistent, always-on inference with millisecond latency. Serverless has cold start delays, Async is asynchronous (not sub-100ms), and Batch Transform is for scheduled offline processing.

Q2: A media company wants to run ML classification on 1 GB video files. Processing time is not urgent. Which SageMaker inference option is MOST appropriate?

  • A) Real-time Endpoints
  • B) Serverless Inference
  • C) Async Inference ✓
  • D) Batch Transform

Explanation: Async Inference supports payloads up to 1 GB and queues requests for processing, making it ideal for large media files. Real-time is limited to 6 MB payload, Serverless to 4 MB, and Batch Transform is for scheduled bulk predictions without real-time queue.

Q3: A SaaS company provides individual ML models for each of their 10,000 enterprise customers. Hosting each on a separate endpoint is too expensive. What is the BEST solution?

  • A) Merge all models into one large model
  • B) Use SageMaker Multi-Model Endpoints ✓
  • C) Deploy all models on a single Batch Transform job
  • D) Use Serverless Inference for each model

Explanation: Multi-Model Endpoints (MME) host multiple models on a single endpoint, dynamically loading them into memory on-demand. This is exactly designed for multi-tenant scenarios where each customer has their own model, reducing infrastructure costs by orders of magnitude.