Introduction
If you don't measure, you can't manage. A local AI stack needs the same operational discipline as other backend services: with SLOs, dashboards, and postmortems.
1. Core Metrics
- Latency: p50, p95, p99
- Quality: groundedness, citation accuracy
- Retrieval: recall@k, hit rate
- Internal cost: CPU/GPU time, memory pressure
2. Designing a Golden Set
A golden set includes these groups:
- Short and clear FAQ questions
- Multi-step questions
- Trick questions with missing context
- Vietnamese questions with technical terminology
Each case has expected behavior and pass/fail criteria.
3. Online Feedback Loop
In the UI, add buttons:
- Helpful
- Not helpful
- Wrong citation
Log feedback by request_id to map back to the prompt, model, and retriever version.
4. Defining Practical SLOs
Example SLOs for an internal Q&A feature:
- 95% of requests respond within 3 seconds
- 90% of answers have valid citations
- Fallback rate due to missing context below 12%
SLOs should include an error budget and a remediation plan for violations.
5. Minimum Dashboard
The dashboard should include:
- Latency by model
- Quality by use case
- Schema error frequency
- Fallback rate
Track by day and by release to detect drift quickly.
Demo Code
Eval framework tests and golden set runner:

Source code: 06-eval-observability