Chuyển đến nội dung chính

課程 3:建構帶應用層策略的 Gemma 4 API 閘道

建構具備逾時、重試、結構化輸出、日誌中繼資料與 團隊別模型存取控制的 FastAPI 閘道。

🧠 AI & ML — L0 課程 3:建構帶應用層策略的 Gemma 4 API 閘道 Gemma 4 本地 AI 工程實戰 on Mac 第二部分:Integration — API、Prompting 與應用整合 xdev.asia

前言

閘道是穩定運作本地 AI 技術棧最重要的層。取代客戶端直接呼叫模型執行環境,為團隊提供品質與安全控制。

1. 最小 API 範本

from fastapi import FastAPI
from pydantic import BaseModel
import requests

app = FastAPI()

class ChatReq(BaseModel):
    prompt: str
    model: str = "gemma4"

@app.post("/chat")
def chat(req: ChatReq):
    r = requests.post(
        "http://127.0.0.1:11434/api/generate",
        json={"model": req.model, "prompt": req.prompt, "stream": False},
        timeout=90,
    )
    r.raise_for_status()
    data = r.json()
    return {"answer": data.get("response", ""), "model": req.model}

2. 必要的閘道策略

  • 每個端點的硬性逾時
  • 暫時性錯誤的有限重試
  • 允許模型的白名單
  • 防止濫用的 prompt 大小限制

3. 結構化輸出

當應用程式需要 JSON 時,在閘道強制 contract:

  • 明確的 prompt contract
  • 回傳給客戶端前驗證 schema
  • Schema 失敗時以特定代碼返回錯誤

4. 內部認證

最低限度實作:

  • 每個服務的 API key
  • 每個 key 的速率限制
  • 按租戶/團隊的日誌記錄

企業級使用時,在閘道層而非模型層接入 SSO。

5. 日誌與追蹤

每個請求應記錄:

  • request_id
  • endpoint
  • model
  • latency_ms
  • prompt_tokens_est
  • status

含有 PII 時,不要記錄原始機密資料。

6. 備援模型策略

設計備援以避免系統硬性崩潰:

  1. 主要模型逾時
  2. 自動切換至輕量模型
  3. 以 degraded_mode=true 標記回應

Demo 程式碼

透過閘道的聊天 API 回應:

聊天回應

模型策略強制——阻擋未授權的模型:

策略強制

原始碼:02-api-gateway

總結