Chuyển đến nội dung chính

Lesson 16: Observability & Evaluation — Monitor what Agents "think"

Tracing agent decisions with LangSmith, Langfuse. Logging, metrics, cost tracking. Evaluation: LLM-as-a-Judge, golden test sets, human evaluation. A/B testing agent prompts.

🧠 AI & ML — Lesson 15 Lesson 16: Observability & Evaluation — According Watch what Agent "thinks"

Build AI Agents: From Zero to Production

Part 6: Production & Actual Deployment

xdev.asia

Introduction

"It works on my laptop" is not enough for a production agent. You need to see what the agent is thinking, why it chose that tool, and measure the output quality in a systematic way.


1. Observability Stack

1.1 Tracing

from langsmith import traceable

@traceable(name="research_agent")
def run_agent(query):
    # Mọi LLM call, tool call đều được trace
    ...

1.2 Key Metrics

  • Latency: Time to first token, total response time
  • Cost: Per-request cost, daily budget usage
  • Success rate: % tasks completed successfully
  • Tool usage: Which tools called most, failure rates

2. Evaluation

2.1 LLM-as-a-Judge

def evaluate_output(task, agent_output, reference_output):
    judge_prompt = f"""
    Evaluate agent output on a scale of 1-5:
    Task: {task}
    Agent output: {agent_output}
    Reference: {reference_output}
    
    Score: [1-5]
    Reasoning: [why]
    """
    return call_llm(judge_prompt)

Summary

  • Observability = tracing + logging + metrics + cost tracking
  • LangSmith / Langfuse for agent tracing
  • Evaluation: LLM-as-Judge, golden sets, human eval
  • A/B test prompts/tools to optimize

Exercises

  1. Setup LangSmith tracing for agent
  2. Build custom dashboard: cost, latency, success rate
  3. Create golden test set (20 cases) and run evaluation
  4. A/B test 2 different system prompts