Chuyển đến nội dung chính

Lesson 3: Building an API gateway for Gemma 4 with application-layer policies

Build a FastAPI gateway with timeout, retry, structured output, logging metadata, and model access control by team.

🧠 AI & ML — L0 Lesson 3: Building an API gateway for Gemma 4 with app-layer policies Gemma 4 Local AI Engineering on Mac Part 2: Integration - API, Prompting & App Embedding xdev.asia

Introduction

The gateway is the most important layer for keeping a local AI stack running reliably. It gives the team control over quality and security, instead of letting clients call the model runtime directly.

1. Minimal API Template

from fastapi import FastAPI
from pydantic import BaseModel
import requests

app = FastAPI()

class ChatReq(BaseModel):
    prompt: str
    model: str = "gemma4"

@app.post("/chat")
def chat(req: ChatReq):
    r = requests.post(
        "http://127.0.0.1:11434/api/generate",
        json={"model": req.model, "prompt": req.prompt, "stream": False},
        timeout=90,
    )
    r.raise_for_status()
    data = r.json()
    return {"answer": data.get("response", ""), "model": req.model}

2. Required Gateway Policies

  • Hard timeout per endpoint
  • Limited retries for transient errors
  • Whitelist of allowed models
  • Prompt size limit to prevent abuse

3. Structured Output

When the app needs JSON, enforce the contract at the gateway:

  • Clear prompt contract
  • Validate schema before returning to client
  • If schema fails, return an error with a specific code

4. Internal Authentication

At minimum, implement:

  • API key per service
  • Rate limit per key
  • Logging per tenant/team

For enterprise use, connect SSO at the gateway rather than at the model layer.

5. Logging and Tracing

Each request should log:

  • request_id
  • endpoint
  • model
  • latency_ms
  • prompt_tokens_est
  • status

Don't log raw sensitive data if PII is present.

6. Fallback Model Strategy

Design fallback so the system doesn't hard crash:

  1. Primary model times out
  2. Automatically switch to a lighter model
  3. Return response with degraded_mode=true flag

Demo Code

Chat API response through the gateway:

Chat Response

Model policy enforcement — blocking unauthorized models:

Policy Enforcement

Source code: 02-api-gateway

Summary