簡介
在筆記本中運行的微調模型≠在生產中運行。本文介紹部署策略。
1.基於API的部署(最簡單)
頂點人工智慧
# Model đã deploy tự động sau tuning job
response = client.models.generate_content(
model=tuned_model_name,
contents="Production query here"
)
開放人工智慧
response = client.chat.completions.create(
model="ft:gpt-4o-mini:org:name:id",
messages=[{"role": "user", "content": "Production query"}]
)
2. 自託管部署
將 LoRA + 部署與 vLLM 合併
# Merge LoRA adapters into base model
python merge_adapters.py --base meta-llama/Llama-3-8B --adapter ./lora_output
# Serve with vLLM
python -m vllm.entrypoints.openai.api_server \
--model ./merged_model \
--host 0.0.0.0 --port 8000
3. 監控
# Track: latency, cost, quality drift
class InferenceMonitor:
def __init__(self):
self.metrics = []
def log(self, query, response, latency, cost):
self.metrics.append({
"timestamp": time.time(),
"latency_ms": latency,
"cost_usd": cost,
"response_length": len(response),
})
總結
- API部署:簡單、Vertex AI或OpenAI
- 自架:vLLM或TGI,需先合併LoRA
- 多重適配器:從 1 個基本模型提供多個微調變體
- 監控:延遲、成本、品質漂移
練習
- 部署微調模型並測試延遲(基礎 vs FT)
- 實施推理監控儀表板 3.負載測試:100個並發請求
- 設定質量漂移檢測