Introduction
The gateway is the most important layer for keeping a local AI stack running reliably. It gives the team control over quality and security, instead of letting clients call the model runtime directly.
1. Minimal API Template
from fastapi import FastAPI
from pydantic import BaseModel
import requests
app = FastAPI()
class ChatReq(BaseModel):
prompt: str
model: str = "gemma4"
@app.post("/chat")
def chat(req: ChatReq):
r = requests.post(
"http://127.0.0.1:11434/api/generate",
json={"model": req.model, "prompt": req.prompt, "stream": False},
timeout=90,
)
r.raise_for_status()
data = r.json()
return {"answer": data.get("response", ""), "model": req.model}
2. Required Gateway Policies
- Hard timeout per endpoint
- Limited retries for transient errors
- Whitelist of allowed models
- Prompt size limit to prevent abuse
3. Structured Output
When the app needs JSON, enforce the contract at the gateway:
- Clear prompt contract
- Validate schema before returning to client
- If schema fails, return an error with a specific code
4. Internal Authentication
At minimum, implement:
- API key per service
- Rate limit per key
- Logging per tenant/team
For enterprise use, connect SSO at the gateway rather than at the model layer.
5. Logging and Tracing
Each request should log:
request_idendpointmodellatency_msprompt_tokens_eststatus
Don't log raw sensitive data if PII is present.
6. Fallback Model Strategy
Design fallback so the system doesn't hard crash:
- Primary model times out
- Automatically switch to a lighter model
- Return response with
degraded_mode=trueflag
Demo Code
Chat API response through the gateway:

Model policy enforcement — blocking unauthorized models:

Source code: 02-api-gateway