Chuyển đến nội dung chính

Lesson 7: Eval framework, observability & SLOs for GenAI

Design a golden set, online feedback loop, latency and groundedness metrics, and define practical SLOs for GenAI features in an internal environment.

🧠 AI & ML — L0 Lesson 7: Eval framework, observability & SLOs for GenAI Gemma 4 Local AI Engineering on Mac Part 4: Reliability, Cost & Production Hardening xdev.asia

Introduction

If you don't measure, you can't manage. A local AI stack needs the same operational discipline as other backend services: with SLOs, dashboards, and postmortems.

1. Core Metrics

  • Latency: p50, p95, p99
  • Quality: groundedness, citation accuracy
  • Retrieval: recall@k, hit rate
  • Internal cost: CPU/GPU time, memory pressure

2. Designing a Golden Set

A golden set includes these groups:

  • Short and clear FAQ questions
  • Multi-step questions
  • Trick questions with missing context
  • Vietnamese questions with technical terminology

Each case has expected behavior and pass/fail criteria.

3. Online Feedback Loop

In the UI, add buttons:

  • Helpful
  • Not helpful
  • Wrong citation

Log feedback by request_id to map back to the prompt, model, and retriever version.

4. Defining Practical SLOs

Example SLOs for an internal Q&A feature:

  • 95% of requests respond within 3 seconds
  • 90% of answers have valid citations
  • Fallback rate due to missing context below 12%

SLOs should include an error budget and a remediation plan for violations.

5. Minimum Dashboard

The dashboard should include:

  • Latency by model
  • Quality by use case
  • Schema error frequency
  • Fallback rate

Track by day and by release to detect drift quickly.

Demo Code

Eval framework tests and golden set runner:

Eval Framework

Source code: 06-eval-observability

Summary