Chuyển đến nội dung chính

第 14 課:部署 — 有效地服務微調模型

部署在 Vertex AI 端點、OpenAI API、自架(vLLM、TGI)上。合併 LoRA 適配器。多適配器服務。監控推斷。

🧠 人工智慧與機器學習 — 第 13 課 第 14 課:部署 — 服務微調 有效模式

微調 LLM:AI 調優的藝術

第 6 部分:生產和最佳實踐

亞洲開發網

簡介

在筆記本中運行的微調模型≠在生產中運行。本文介紹部署策略。


1.基於API的部署(最簡單)

頂點人工智慧

# Model đã deploy tự động sau tuning job
response = client.models.generate_content(
    model=tuned_model_name,
    contents="Production query here"
)

開放人工智慧

response = client.chat.completions.create(
    model="ft:gpt-4o-mini:org:name:id",
    messages=[{"role": "user", "content": "Production query"}]
)

2. 自託管部署

將 LoRA + 部署與 vLLM 合併

# Merge LoRA adapters into base model
python merge_adapters.py --base meta-llama/Llama-3-8B --adapter ./lora_output

# Serve with vLLM
python -m vllm.entrypoints.openai.api_server \
    --model ./merged_model \
    --host 0.0.0.0 --port 8000

3. 監控

# Track: latency, cost, quality drift
class InferenceMonitor:
    def __init__(self):
        self.metrics = []
    
    def log(self, query, response, latency, cost):
        self.metrics.append({
            "timestamp": time.time(),
            "latency_ms": latency,
            "cost_usd": cost,
            "response_length": len(response),
        })

總結

  • API部署:簡單、Vertex AI或OpenAI
  • 自架:vLLM或TGI,需先合併LoRA
  • 多重適配器:從 1 個基本模型提供多個微調變體
  • 監控:延遲、成本、品質漂移

練習

  1. 部署微調模型並測試延遲(基礎 vs FT)
  2. 實施推理監控儀表板 3.負載測試:100個並發請求
  3. 設定質量漂移檢測