Introduction
"It works on my laptop" is not enough for a production agent. You need to see what the agent is thinking, why it chose that tool, and measure the output quality in a systematic way.
1. Observability Stack
1.1 Tracing
from langsmith import traceable
@traceable(name="research_agent")
def run_agent(query):
# Mọi LLM call, tool call đều được trace
...
1.2 Key Metrics
- Latency: Time to first token, total response time
- Cost: Per-request cost, daily budget usage
- Success rate: % tasks completed successfully
- Tool usage: Which tools called most, failure rates
2. Evaluation
2.1 LLM-as-a-Judge
def evaluate_output(task, agent_output, reference_output):
judge_prompt = f"""
Evaluate agent output on a scale of 1-5:
Task: {task}
Agent output: {agent_output}
Reference: {reference_output}
Score: [1-5]
Reasoning: [why]
"""
return call_llm(judge_prompt)
Summary
- Observability = tracing + logging + metrics + cost tracking
- LangSmith / Langfuse for agent tracing
- Evaluation: LLM-as-Judge, golden sets, human eval
- A/B test prompts/tools to optimize
Exercises
- Setup LangSmith tracing for agent
- Build custom dashboard: cost, latency, success rate
- Create golden test set (20 cases) and run evaluation
- A/B test 2 different system prompts