Chuyển đến nội dung chính

Lesson 14: Deployment — Serve Fine-tuned Model effectively

Deployed on Vertex AI endpoints, OpenAI API, self-hosted (vLLM, TGI). Merge LoRA adapters. Multi-adapter serving. Monitoring inference.

🧠 AI & ML — Lesson 13 Lesson 14: Deployment — Serve Fine-tuned Effective model

Fine-tuning LLM: The Art of AI Tuning

Part 6: Production & Best Practices

xdev.asia

Introduction

Fine-tuned model running in notebook ≠ running in production. This article covers deployment strategies.


1. API-based Deployment (Easiest)

Vertex AI

# Model đã deploy tự động sau tuning job
response = client.models.generate_content(
    model=tuned_model_name,
    contents="Production query here"
)

OpenAI

response = client.chat.completions.create(
    model="ft:gpt-4o-mini:org:name:id",
    messages=[{"role": "user", "content": "Production query"}]
)

2. Self-hosted Deployment

Merge LoRA + Deploy with vLLM

# Merge LoRA adapters into base model
python merge_adapters.py --base meta-llama/Llama-3-8B --adapter ./lora_output

# Serve with vLLM
python -m vllm.entrypoints.openai.api_server \
    --model ./merged_model \
    --host 0.0.0.0 --port 8000

3. Monitoring

# Track: latency, cost, quality drift
class InferenceMonitor:
    def __init__(self):
        self.metrics = []
    
    def log(self, query, response, latency, cost):
        self.metrics.append({
            "timestamp": time.time(),
            "latency_ms": latency,
            "cost_usd": cost,
            "response_length": len(response),
        })

Summary

  • API deployment: simple, Vertex AI or OpenAI
  • Self-hosted: vLLM or TGI, need to merge LoRA first
  • Multi-adapter: serve multiple fine-tuned variants from 1 base model
  • Monitoring: latency, cost, quality drift

Exercises

  1. Deploy fine-tuned model and test latency (base vs FT)
  2. Implement inference monitoring dashboard
  3. Load test: 100 concurrent requests
  4. Setup quality drift detection