Introduction
Before building an agent, you need to master the core tool — LLM APIs. This article focuses on the three most popular providers: OpenAI (GPT-4o), Anthropic (Claude 3.5 Sonnet), and Google (Gemini). You will learn how to call APIs, handle streaming, structured output, and most importantly — optimize costs.
1. Overview of 3 LLM Providers
Quick comparison
| OpenAI GPT-4o | Anthropic Claude 3.5 | Google Gemini 1.5 | |
|---|---|---|---|
| Input price | $2.50/1M tokens | $3.00/1M tokens | $1.25/1M tokens |
| Output price | $10.00/1M tokens | $15.00/1M tokens | $5.00/1M tokens |
| Context window | 128K | 200K | 1M |
| Strengths | Tool use, coding | Reasoning long, safety | Huge Context, search |
| Vision | ✅ | ✅ | ✅ |
| Streaming | ✅ | ✅ | ✅ |
When to use what?
- OpenAI: Default choice, largest ecosystem, stable function calling
- Claude: When needing complex reasoning, handling long documents, or needing high safety
- Gemini: When you need a very large context window or grounding with search
2. OpenAI API
2.1 Setup & Authentication
pip install openai
from openai import OpenAI
# Cách 1: Environment variable (khuyến nghị)
# export OPENAI_API_KEY=sk-...
client = OpenAI()
# Cách 2: Truyền trực tiếp
client = OpenAI(api_key="sk-...")
2.2 Chat Completions — Basic
response = client.chat.completions.create(
model="gpt-4o-mini", # Rẻ và nhanh, đủ cho hầu hết use case
messages=[
{"role": "system", "content": "Bạn là trợ lý AI nói tiếng Việt."},
{"role": "user", "content": "Giải thích AI Agent trong 3 câu."}
],
temperature=0.7, # Creativity level (0 = deterministic, 2 = creative)
max_tokens=500, # Giới hạn output
)
print(response.choices[0].message.content)
print(f"Tokens used: {response.usage.total_tokens}")
print(f"Cost: ~${response.usage.total_tokens * 0.00000015:.6f}")
2.3 Streaming
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Viết một bài thơ về AI"}],
stream=True,
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
2.4 Structured Output (JSON Mode)
from pydantic import BaseModel
class AgentStep(BaseModel):
thought: str
action: str
tool_name: str | None
tool_args: dict | None
response = client.beta.chat.completions.parse(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "Analyze the user request and plan the next agent step."},
{"role": "user", "content": "Find the weather in Hanoi and compare with Saigon"},
],
response_format=AgentStep,
)
step = response.choices[0].message.parsed
print(f"Thought: {step.thought}")
print(f"Action: {step.action}")
print(f"Tool: {step.tool_name}({step.tool_args})")
3. Anthropic Claude API
3.1 Setup
pip install anthropic
from anthropic import Anthropic
client = Anthropic() # Dùng ANTHROPIC_API_KEY env var
3.2 Messages API
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
system="Bạn là trợ lý AI cho developer Việt Nam.",
messages=[
{"role": "user", "content": "So sánh LangChain vs LangGraph"}
]
)
print(message.content[0].text)
print(f"Input tokens: {message.usage.input_tokens}")
print(f"Output tokens: {message.usage.output_tokens}")
3.3 Streaming
with client.messages.stream(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[{"role": "user", "content": "Explain ReAct pattern"}]
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
4. Google Gemini API
4.1 Setup
pip install google-genai
from google import genai
client = genai.Client() # Dùng GOOGLE_API_KEY env var
4.2 Generate Content
response = client.models.generate_content(
model="gemini-2.0-flash",
contents="Giải thích MCP protocol cho AI Agents"
)
print(response.text)
5. Best Practices for Agent Development
5.1 Choose the appropriate model
# Quy tắc ngón tay cái:
MODEL_SELECTION = {
"simple_tasks": "gpt-4o-mini", # Rẻ, nhanh
"complex_reasoning": "claude-sonnet-4-20250514", # Chính xác
"long_context": "gemini-1.5-pro", # 1M tokens
"tool_calling": "gpt-4o", # Ổn định nhất
"cost_sensitive": "gpt-4o-mini", # Rẻ nhất
}
5.2 Error Handling
from openai import RateLimitError, APIError
import time
def call_llm_with_retry(messages, max_retries=3):
for attempt in range(max_retries):
try:
return client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
)
except RateLimitError:
wait = 2 ** attempt
print(f"Rate limited. Waiting {wait}s...")
time.sleep(wait)
except APIError as e:
print(f"API error: {e}")
raise
raise Exception("Max retries exceeded")
5.3 Cost Tracking
class CostTracker:
PRICING = {
"gpt-4o-mini": {"input": 0.15/1e6, "output": 0.60/1e6},
"gpt-4o": {"input": 2.50/1e6, "output": 10.00/1e6},
}
def __init__(self):
self.total_cost = 0
def track(self, model, usage):
pricing = self.PRICING.get(model, {"input": 0, "output": 0})
cost = (usage.prompt_tokens * pricing["input"] +
usage.completion_tokens * pricing["output"])
self.total_cost += cost
return cost
tracker = CostTracker()
Summary
- Proficient in 3 LLM APIs: OpenAI, Anthropic, Google Gemini
- Know when to use which model (cost vs capability trade-off)
- Streaming, structured output, error handling — ready for agent development
- Cost tracking is required when building an agent (agent calls LLM multiple times!)
Exercises
- Write a wrapper function that calls all 3 providers with the same interface
- Compare the response quality of 3 models for the same complex prompt
- Implement cost tracker and run 10 requests, calculating total cost
- Try Structured Output: force LLM to return JSON with a specific schema