Introduction
An AI demo running in Jupyter Notebook and an AI system running in production are two completely different worlds. This article will design an AI microservices architecture for Fashion AI Platform — where many AI models serve thousands of requests simultaneously.
1. Why is there a need for a separate architecture for AI?
Challenges of AI in Production
| Challenge | Explanation |
|---|---|
| GPU scarcity | GPU is expensive, needs to be shared between many models |
| Long inference time | Stable Diffusion takes 5–15s, cannot be blocked |
| Model size | SDXL ~6.5GB VRAM, CLIP ~2GB, SMPL ~500MB |
| Concurrency | Multiple users generate at the same time |
| Model versioning | A/B test new vs old model |
| Cold start | Loading the model takes 30–60s |
Monolith vs Microservices for AI
❌ Monolith AI Server
- 1 server load TẤT CẢ models
- VRAM overflow
- 1 model crash → toàn bộ hệ thống down
✅ AI Microservices
- Mỗi module AI = 1 service riêng
- Scale độc lập
- Isolate failures
2. Overall architecture
┌─────────────────┐
│ API Gateway │
│ (FastAPI) │
└────────┬────────┘
│
┌────────▼────────┐
│ Task Queue │
│ (Celery + Redis)│
└────────┬────────┘
│
┌────────────────┼────────────────┐
│ │ │
┌────────▼──────┐ ┌──────▼───────┐ ┌──────▼───────┐
│ Design Gen │ │ Edit │ │ Try-On │
│ Service │ │ Service │ │ Service │
│ (GPU: A100) │ │ (GPU: A100) │ │ (GPU: A10G) │
└────────┬──────┘ └──────┬───────┘ └──────┬───────┘
│ │ │
┌────────▼──────┐ ┌──────▼───────┐ ┌──────▼───────┐
│ Model Store │ │ Personalize │ │ Production │
│ (S3 + Cache) │ │ Service │ │ AI Service │
│ │ │ (CPU/T4) │ │ (CPU/T4) │
└───────────────┘ └──────────────┘ └──────────────┘
Main components
2.1 API Gateway (FastAPI)
# Nhận request từ frontend
# Validate input
# Gửi task vào queue
# Return task_id để client polling
@app.post("/api/v1/design/generate")
async def generate_design(request: GenerateRequest):
task = celery_app.send_task(
"design_gen.generate",
args=[request.dict()],
queue="gpu_high_priority"
)
return {"task_id": task.id, "status": "queued"}
2.2 Task Queue (Celery + Redis)
# Queues phân theo priority và resource
CELERY_QUEUES = {
"gpu_high_priority": {
"tasks": ["design_gen.generate", "edit.apply"],
"concurrency": 2, # Max 2 concurrent GPU tasks
},
"gpu_low_priority": {
"tasks": ["tryon.render", "production.upscale"],
"concurrency": 4,
},
"cpu_tasks": {
"tasks": ["personalize.update", "production.convert_cmyk"],
"concurrency": 16,
},
}
2.3 Model Serving Layer
Two main options:
| NVIDIA Triton | vLLM | Custom (PyTorch) | |
|---|---|---|---|
| Best for | Multi-model serving | LLM inference | Diffusion models |
| Batching | Dynamic batching | Continuous batching | Manual |
| Format | ONNX, TensorRT, PyTorch | HF models | Any |
| Use case | CLIP, classifier | Product description gene | SDXL, ControlNet |
3. GPU Scheduling Strategy
GPU Pool Architecture
GPU Pool
├── Partition A: Design Generation (A100 x 2)
│ ├── SDXL model (always loaded)
│ ├── ControlNet (loaded on demand)
│ └── IP-Adapter (loaded on demand)
│
├── Partition B: Editing & Try-On (A100 x 1 + A10G x 1)
│ ├── InstructPix2Pix
│ ├── SMPL-X
│ └── Cloth simulation
│
└── Partition C: Lightweight AI (T4 x 2)
├── CLIP (always loaded)
├── Auto-tagger
├── Real-ESRGAN upscaler
└── Size recommendation model
Model Loading Strategy
class ModelManager:
"""Quản lý load/unload models theo demand"""
def __init__(self, gpu_memory_limit: int):
self.loaded_models: dict[str, Model] = {}
self.gpu_memory_limit = gpu_memory_limit
self.lru_cache = OrderedDict()
async def get_model(self, model_name: str) -> Model:
if model_name in self.loaded_models:
self.lru_cache.move_to_end(model_name)
return self.loaded_models[model_name]
# Check if enough VRAM
required = MODEL_VRAM_MAP[model_name]
while self._used_vram() + required > self.gpu_memory_limit:
self._evict_lru()
model = await self._load_model(model_name)
self.loaded_models[model_name] = model
self.lru_cache[model_name] = True
return model
4. Model Versioning & A/B Testing
Model Registry
# model_registry.yaml
models:
design_gen_sdxl:
current: v2.1
versions:
v2.1:
path: s3://models/sdxl-fashion-v2.1/
lora: s3://models/lora-tshirt-v2.1/
metrics:
fid_score: 12.3
user_satisfaction: 4.2
v2.0:
path: s3://models/sdxl-fashion-v2.0/
metrics:
fid_score: 15.7
user_satisfaction: 3.8
style_analyzer:
current: v1.0
versions:
v1.0:
path: s3://models/clip-style-v1.0/
type: onnx
A/B Testing Pipeline
class ABTestRouter:
"""Route requests tới model versions theo experiment config"""
def __init__(self, experiments: list[Experiment]):
self.experiments = experiments
def get_model_version(
self, model_name: str, user_id: str
) -> str:
experiment = self._get_active_experiment(model_name)
if not experiment:
return self._get_default_version(model_name)
# Deterministic assignment based on user_id
bucket = hash(f"{user_id}:{experiment.id}") % 100
for variant in experiment.variants:
if bucket < variant.traffic_percentage:
return variant.model_version
bucket -= variant.traffic_percentage
return experiment.control_version
5. Async Processing Pipeline
Design Generation Flow
Client API Queue GPU Worker
│ │ │ │
│── POST /generate ──►│ │ │
│ │── send_task ──►│ │
│◄── {task_id} ──────│ │ │
│ │ │── pick task ────►│
│── GET /status ─────►│ │ │
│◄── "processing" ───│ │ │
│ │ │ (5-15s) │
│ │ │◄── result ──────│
│── GET /status ─────►│ │ │
│◄── "completed" ────│ │ │
│ + design URLs │ │ │
Webhook Alternative (for production)
# Thay vì polling, dùng webhook callback
@celery_app.task(bind=True)
def generate_design(self, request: dict):
result = model.generate(request)
# Notify client via webhook
webhook_url = request.get("callback_url")
if webhook_url:
requests.post(webhook_url, json={
"task_id": self.request.id,
"status": "completed",
"designs": result.urls,
})
return result
6. Storage Architecture
┌────────────────────────────────────────────┐
│ Storage Layer │
├────────────────────────────────────────────┤
│ │
│ S3 / MinIO │
│ ├── /models/ (AI model weights) │
│ ├── /designs/ (generated designs) │
│ ├── /uploads/ (user uploads) │
│ ├── /print-files/ (CMYK print-ready) │
│ └── /mockups/ (product mockups) │
│ │
│ Redis │
│ ├── Task queue │
│ ├── Model cache metadata │
│ ├── User session / style embeddings │
│ └── Rate limiting │
│ │
│ PostgreSQL │
│ ├── User data │
│ ├── Design metadata │
│ ├── Order history │
│ ├── A/B test results │
│ └── Behavioral logs │
│ │
│ Vector DB (Qdrant / Pinecone) │
│ ├── Design embeddings (search) │
│ ├── Style embeddings (personalization) │
│ └── User preference vectors │
│ │
└────────────────────────────────────────────┘
7. Error Handling & Resilience
Retry Strategy
@celery_app.task(
bind=True,
max_retries=3,
default_retry_delay=5,
autoretry_for=(GPUOutOfMemoryError, ModelLoadError),
)
def generate_design(self, request: dict):
try:
model = model_manager.get_model("sdxl")
return model.generate(request)
except GPUOutOfMemoryError:
# Clear GPU cache and retry
torch.cuda.empty_cache()
raise self.retry()
Fallback Strategy
# Nếu SDXL fail → fallback về lightweight model
FALLBACK_CHAIN = [
"sdxl_fashion_v2", # Primary: best quality
"sdxl_base", # Fallback 1: generic SDXL
"sd15_fashion", # Fallback 2: SD 1.5 (faster, less quality)
]
8. Monitoring & Observability
Key Metrics
# Prometheus metrics cho AI system
ai_request_duration = Histogram(
"ai_request_duration_seconds",
"Time spent processing AI request",
["model_name", "model_version", "task_type"],
)
ai_gpu_utilization = Gauge(
"ai_gpu_utilization_percent",
"GPU utilization percentage",
["gpu_id", "partition"],
)
ai_queue_length = Gauge(
"ai_queue_length",
"Number of pending tasks in queue",
["queue_name", "priority"],
)
ai_generation_quality = Histogram(
"ai_generation_quality_score",
"Quality score of generated designs",
["model_version"],
)
Summary
AI system architecture for Fashion AI Platform includes:
- API Gateway (FastAPI) — receive requests, validate, queue tasks
- Task Queue (Celery + Redis) — async processing, priority queues
- GPU Workers — partition by workload, LRU model loading
- Model Registry — versioning, A/B testing, rollback
- Storage — S3 (files), Redis (cache), PostgreSQL (metadata), Vector DB (embeddings)
- Monitoring — GPU utilization, queue length, generation quality
The next article will go into detail AI Tech Stack: comparing Stable Diffusion XL vs FLUX, ControlNet, CLIP, and setting up MLOps pipeline.