Giới thiệu
Gateway là lớp quan trọng nhất để local AI stack chạy ổn định. Nó giúp team kiểm soát chất lượng và bảo mật, thay vì để client gọi trực tiếp model runtime.
1. Mẫu API tối thiểu
from fastapi import FastAPI
from pydantic import BaseModel
import requests
app = FastAPI()
class ChatReq(BaseModel):
prompt: str
model: str = "gemma4"
@app.post("/chat")
def chat(req: ChatReq):
r = requests.post(
"http://127.0.0.1:11434/api/generate",
json={"model": req.model, "prompt": req.prompt, "stream": False},
timeout=90,
)
r.raise_for_status()
data = r.json()
return {"answer": data.get("response", ""), "model": req.model}
2. Policy bắt buộc ở gateway
- Timeout cứng theo endpoint
- Retry giới hạn cho lỗi tạm thời
- Whitelist model được phép dùng
- Limit kích thước prompt để tránh abuse
3. Structured output
Khi app cần JSON, ép contract ngay từ gateway:
- Prompt contract rõ ràng
- Validate schema trước khi trả client
- Nếu fail schema, trả lỗi có mã cụ thể
4. Authentication nội bộ
Tối thiểu triển khai:
- API key theo service
- Rate limit theo key
- Logging theo tenant/team
Nếu dùng trong doanh nghiệp, kết nối SSO ở gateway thay vì tại model layer.
5. Logging và tracing
Mỗi request nên log:
request_idendpointmodellatency_msprompt_tokens_eststatus
Không log dữ liệu nhạy cảm nguyên văn nếu có PII.
6. Fallback model strategy
Thiết kế fallback để hệ thống không chết cứng:
- Model chính timeout
- Tự chuyển model nhẹ hơn
- Trả response có cờ
degraded_mode=true
Demo code
Kết quả test Chat API qua gateway:

Model policy enforcement — chặn model không được phép:

Source code: 02-api-gateway
Tóm tắt
Gateway biến local LLM thành một dịch vụ đúng nghĩa. Bài tiếp theo sẽ đi sâu vào prompt contract, JSON schema và regression test để giữ hành vi model ổn định theo thời gian.