簡介
MLOps 很難——LLMOps 更難。 LLM 與傳統模式不同:你不需要從頭開始訓練,評估非常困難,推理成本高 100 倍。
🎯 LLMOps = 將 LLM 應用程式投入生產的有效流程。
1. MLOps 與 LLMOps — 有什麼不同?
Traditional ML:
Data → Feature Eng → Train Model → Deploy → Monitor
✅ Model nhỏ (MB)
✅ Train trên dataset nhỏ (GB)
✅ Evaluation rõ ràng (accuracy, F1)
✅ Inference rẻ (<1ms)
LLM:
Prompt → LLM API → Post-process → Deploy → Monitor
❌ Model rất lớn (100B params, 100GB+)
❌ Không train (hoặc fine-tune rất đắt)
❌ Evaluation mơ hồ (chất lượng text?)
❌ Inference đắt ($0.01-$0.10/request)
###詳細比較:
| 面向 | 傳統機器學習 | LLMOps |
|---|---|---|
| 型號 | 從頭開始訓練 | 預先訓練、提示/微調 |
| 資料 | 結構化(表格) | 非結構化(文字、程式碼) |
| 訓練 | 小時-天 | 週-月(或不適用) |
| 成本 | 火車貴,服務便宜 | 火車非常貴,服務也很貴 |
| 評估 | 清晰的指標(acc、F1) | 主觀+多重維度 |
| 版本控制 | 模型重量 | 提示+選用 |
| 部署 | 自架簡單 | API 呼叫或龐大的 GPU 伺服器 |
| 監控 | 資料漂移 | 迅速漂移、幻覺、中毒 |
| 故障模式 | 錯誤的預測 | 幻覺、有害內容 |
2.LLMOps 堆疊
┌─────────────────────────────────────────────────┐
│ LLMOps Stack │
├──────────────┬──────────────────────────────────┤
│ Layer │ Components │
├──────────────┼──────────────────────────────────┤
│ Foundation │ OpenAI, Anthropic, Google, Llama │
│ Models │ Mistral, Cohere │
├──────────────┼──────────────────────────────────┤
│ Prompt │ Prompt templates, versioning │
│ Management │ A/B testing, optimization │
├──────────────┼──────────────────────────────────┤
│ RAG │ Vector DBs, embeddings, chunking │
│ Pipeline │ retrieval, re-ranking │
├──────────────┼──────────────────────────────────┤
│ Orchestration│ LangChain, LlamaIndex, DSPy │
│ │ Agents, chains, tool use │
├──────────────┼──────────────────────────────────┤
│ Evaluation │ Human eval, LLM-as-judge │
│ │ Benchmarks, A/B testing │
├──────────────┼──────────────────────────────────┤
│ Observability│ LangSmith, Langfuse, Arize │
│ │ Tracing, logging, debugging │
├──────────────┼──────────────────────────────────┤
│ Guardrails │ Content filtering, PII detection │
│ │ Output validation, rate limiting │
├──────────────┼──────────────────────────────────┤
│ Cost Mgmt │ Caching, routing, token counting │
│ │ Budget alerts, model selection │
└──────────────┴──────────────────────────────────┘
3. LLM申請模式
3.1 直接API調用
"""Pattern 1: Direct API call — đơn giản nhất"""
from openai import OpenAI
client = OpenAI()
def classify_sentiment(text):
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "Classify sentiment as positive/negative/neutral. Return ONLY the label."},
{"role": "user", "content": text},
],
temperature=0,
max_tokens=10,
)
return response.choices[0].message.content.strip()
# Simple, nhưng:
# ❌ Không cache
# ❌ Không track prompts
# ❌ Không monitor
# ❌ Không fallback
3.2 生產就緒模式
"""Pattern 2: Production-ready LLM call"""
from openai import OpenAI
import hashlib
import json
import time
import logging
from functools import lru_cache
logger = logging.getLogger(__name__)
class LLMService:
def __init__(self):
self.client = OpenAI()
self.cache = {} # Redis in production
self.total_tokens = 0
self.total_cost = 0
def call(self, messages, model="gpt-4o-mini", **kwargs):
"""Production LLM call with caching, logging, and fallback"""
# 1. Check cache
cache_key = self._cache_key(messages, model)
if cache_key in self.cache:
logger.info(f"Cache hit: {cache_key[:8]}")
return self.cache[cache_key]
# 2. Call with retry
for attempt in range(3):
try:
start = time.time()
response = self.client.chat.completions.create(
model=model,
messages=messages,
**kwargs,
)
latency = time.time() - start
# 3. Track usage
usage = response.usage
self.total_tokens += usage.total_tokens
cost = self._calculate_cost(model, usage)
self.total_cost += cost
# 4. Log
logger.info(
f"LLM call: model={model}, tokens={usage.total_tokens}, "
f"latency={latency:.2f}s, cost=${cost:.4f}"
)
result = response.choices[0].message.content
self.cache[cache_key] = result
return result
except Exception as e:
logger.warning(f"Attempt {attempt+1} failed: {e}")
if attempt == 2:
# Fallback to cheaper model
if model != "gpt-4o-mini":
logger.info("Falling back to gpt-4o-mini")
return self.call(messages, model="gpt-4o-mini", **kwargs)
raise
time.sleep(2 ** attempt)
def _cache_key(self, messages, model):
content = json.dumps({"messages": messages, "model": model})
return hashlib.md5(content.encode()).hexdigest()
def _calculate_cost(self, model, usage):
prices = {
"gpt-4o": {"input": 2.50/1e6, "output": 10.0/1e6},
"gpt-4o-mini": {"input": 0.15/1e6, "output": 0.60/1e6},
}
p = prices.get(model, prices["gpt-4o-mini"])
return usage.prompt_tokens * p["input"] + \
usage.completion_tokens * p["output"]
3.3 RAG 模式
"""Pattern 3: RAG (Retrieval-Augmented Generation)"""
from openai import OpenAI
import chromadb
class RAGService:
def __init__(self):
self.llm = OpenAI()
self.db = chromadb.HttpClient(host="localhost", port=8000)
self.collection = self.db.get_collection("knowledge_base")
def query(self, question, top_k=5):
# 1. Retrieve relevant documents
results = self.collection.query(
query_texts=[question],
n_results=top_k,
)
# 2. Build context
context = "\n\n".join(results['documents'][0])
# 3. Generate answer
response = self.llm.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": f"""
Answer based ONLY on the provided context.
If the answer is not in the context, say "I don't know."
Context:
{context}
"""},
{"role": "user", "content": question},
],
)
return {
"answer": response.choices[0].message.content,
"sources": results['metadatas'][0],
}
4. LLM評估-巨大挑戰
4.1 問題
Traditional ML: accuracy = 0.92 → Good!
LLM: "Summarize this article" → ???
- Chính xác?
- Đầy đủ?
- Ngắn gọn?
- Đúng tone?
- Không hallucinate?
→ Rất khó đo tự động
4.2 評估框架
"""LLM Evaluation pipeline"""
from openai import OpenAI
import json
client = OpenAI()
def evaluate_with_llm_judge(question, answer, reference_answer=None):
"""LLM-as-a-Judge evaluation"""
eval_prompt = f"""
Evaluate the following answer on these criteria (1-5 scale):
1. **Correctness**: Is the answer factually correct?
2. **Completeness**: Does it cover all key points?
3. **Conciseness**: Is it appropriately concise?
4. **Relevance**: Does it actually answer the question?
5. **Harmlessness**: Is it safe and appropriate?
Question: {question}
Answer: {answer}
{"Reference: " + reference_answer if reference_answer else ""}
Respond in JSON: {{"correctness": X, "completeness": X, "conciseness": X, "relevance": X, "harmlessness": X, "explanation": "..."}}
"""
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": eval_prompt}],
response_format={"type": "json_object"},
)
return json.loads(response.choices[0].message.content)
# Chạy evaluation trên test set
test_cases = [
{
"question": "Python list comprehension là gì?",
"reference": "List comprehension là cú pháp ngắn gọn để tạo list mới từ iterable.",
},
# ... more test cases
]
results = []
for tc in test_cases:
answer = my_llm_service.query(tc["question"])
eval_result = evaluate_with_llm_judge(
tc["question"], answer, tc.get("reference")
)
results.append(eval_result)
print(f"Q: {tc['question'][:50]}... → Score: {eval_result}")
# Aggregate
avg_scores = {
key: sum(r[key] for r in results) / len(results)
for key in ["correctness", "completeness", "relevance"]
}
print(f"\n📊 Average scores: {avg_scores}")
5. 決策架構:提示、微調、RAG
Cần domain knowledge?
/ \
Yes No
| |
Data có sẵn? Task phức tạp?
/ \ / \
Yes No Yes No
| | | |
Fine-tune RAG + Prompt Few-shot Zero-shot
+ RAG Engineering Prompting Prompting
Quy tắc ngón cái:
1. Thử Prompt Engineering trước (rẻ, nhanh)
2. Thêm RAG nếu cần knowledge (medium effort)
3. Fine-tune cuối cùng (đắt, chậm, nhưng mạnh)
| 方法 | 成本 | 努力 | 何時使用 |
|---|---|---|---|
| 零射擊 | 💰 | 🔨 | 簡單的任務,強大的模型 |
| 少射 | 💰 | 🔨🔨 | 需要例子,具體格式 |
| 抹布 | 💰💰 | 🔨🔨🔨 | 需要領域知識、資料變更 |
| 微調 | 💰💰💰 | 🔨🔨🔨🔨 | 具體任務,需要獨特風格 |
6.LLMOps 生命週期
"""LLMOps Lifecycle trong practice"""
# Phase 1: Prototype (1-2 tuần)
# - Thử nhiều prompts trong playground
# - So sánh models (GPT-4o vs Claude vs Gemini)
# - Build basic RAG nếu cần
# Phase 2: Evaluation (1-2 tuần)
# - Tạo eval dataset (50-200 test cases)
# - LLM-as-judge evaluation
# - Human evaluation (sample)
# - Benchmark: latency, cost, accuracy
# Phase 3: Production (1-2 tuần)
# - Prompt versioning
# - Caching (semantic cache)
# - Rate limiting
# - Error handling & fallbacks
# - Guardrails (content filter)
# Phase 4: Monitoring (ongoing)
# - Track: latency, cost, token usage
# - Track: user feedback, thumbs up/down
# - Track: hallucination rate
# - Data drift detection
# - A/B testing prompts
總結
| 概念 | 記住 |
|---|---|
| LLMOps | 適合 LLM 申請的 MLOps |
| 主要區別 | 無需訓練、推理昂貴、評估困難 |
| LLMOps 堆疊 | 模型 → 提示 → RAG → 編排 → 評估 → 可觀察性 |
| 評估 | 法學碩士作為法官+人工評估+自動化指標 |
| 決定 | 首先提示 → RAG → 微調(依需求升級) |
| 生產 | 快取+重試+回退+護欄 |
練習
- 堆疊分析: 您的團隊使用哪種法學碩士?列出生產所需的組件。
- 評估: 建立包含 20 個測試案例的評估資料集。以法官身分運行法學碩士。報告成績。
- 生產包裝器: 使用快取、重試、回退、成本追蹤來包裝 OpenAI API 呼叫。
- 決策矩陣: 對於 3 個用例,決定:提示、RAG 與微調。解釋。
下一篇文章: 及時管理和 A/B 測試。