1. Agent 抽象層 — LLM + 記憶 + 工具 + 規劃
1.1. 什麼是 Agent?
第 8 課介紹了簡單的 ReAct agent — 一個選擇工具然後回答的 LLM。第 9 課擴展到 Agentic AI:LLM 作為中央大腦的系統 — 自主規劃、選擇工具、回應結果,並與其他代理協調。
一個 Agent 由 4 個核心元件組成:
- LLM(大腦) — 推理、決策、文字生成
- 記憶 — 短期(對話緩衝區)與長期(向量儲存、資料庫)
- 工具 — 代理可呼叫的函式:搜尋、計算器、API、程式碼執行
- 規劃 — 建立計畫、分解任務、發生錯誤時反思
Agent Abstraction — Core Components
══════════════════════════════════════════════════════════════
┌──────────────────────┐
│ USER │
│ (Task / Query) │
└──────────┬───────────┘
│
▼
┌────────────────────────────────────────────────────────┐
│ AGENT │
│ ┌──────────┐ ┌───────────┐ ┌──────────────────┐ │
│ │ Planning │ │ LLM Core │ │ Memory │ │
│ │ │◄─┤ (Brain) ├─►│ Short-term: chat │ │
│ │ Decompose│ │ Reasoning │ │ Long-term: VDB │ │
│ │ Reflect │ │ Decisions │ │ Episodic: logs │ │
│ └──────────┘ └─────┬─────┘ └──────────────────┘ │
│ │ │
│ ┌────────┼────────┐ │
│ ▼ ▼ ▼ │
│ ┌────────┐┌───────┐┌────────┐ │
│ │Search ││ Code ││ API │ ◄── Tools │
│ │Engine ││ Exec ││ Calls │ │
│ └────────┘└───────┘└────────┘ │
└────────────────────────────────────────────────────────┘
1.2. 感知 → 推理 → 行動 → 觀察迴圈
每個代理都在一個基本迴圈上運行:
- 感知 — 接收輸入(使用者查詢、工具輸出、環境回饋)
- 推理 — LLM 推理:「接下來該做什麼?用哪個工具?資訊夠了嗎?」
- 行動 — 執行動作:呼叫工具、生成文字、回答使用者
- 觀察 — 接收行動的結果,回饋至感知 → 重複
Agent Loop — Perception → Reasoning → Action → Observation
═══════════════════════════════════════════════════════════
┌──────────────┐ ┌───────────────┐ ┌──────────────┐
│ PERCEPTION │────►│ REASONING │────►│ ACTION │
│ │ │ │ │ │
│ User query │ │ "Which tool?" │ │ Call tool │
│ Tool output │ │ "Enough info?"│ │ Generate text │
│ Error msg │ │ "Need retry?" │ │ Return answer │
└──────┬───────┘ └───────────────┘ └───────┬───────┘
▲ │
│ ┌───────────────┐ │
└───────────│ OBSERVATION │◄──────────────┘
│ │
│ Tool result │
│ Error / OK │
└───────────────┘
Loop continues until Final Answer
1.3. 代理能力等級
並非每個 LLM 應用都需要完整的代理。NVIDIA DLI 區分了不同的代理能力等級:
| 等級 | 模式 | LLM 角色 | 範例 |
|---|---|---|---|
| L0 — 無代理能力 | 簡單提示 → 回應 | 文字生成器 | FAQ 聊天機器人 |
| L1 — 工具使用 | LLM 選擇 1 個工具 | 路由器 | Function calling API |
| L2 — 單一代理 | ReAct 迴圈、多步驟 | 規劃者 + 執行者 | RAG Agent(第 8 課) |
| L3 — 多代理 | 多個代理協調 | 協調者 | 監督者 + 工作者 |
| L4 — 自主式 | 自我改進、長時間運行 | 自主系統 | AI Scientist、Devin |
考試提示:「LLM 自主分解任務、呼叫多個工具、根據需要重試」→ Agent(L2+)。「多個 LLM 協調,每個專精一項任務」→ Multi-Agent(L3)。DLI 考試常問:「Agent 與 Chain 有何不同?」→ Agent 具有動態控制流(LLM 決定下一步),Chain 具有固定控制流。

2. LLM Agent 的認知架構
2.1. ReAct — 推理 + 行動
ReAct(第 8 課介紹)交替進行思考(推理)與行動(執行)。優點:簡單、透明。缺點:沒有長期規劃 — 代理只思考下一步,而非全局。
2.2. Plan-and-Execute
Plan-and-Execute 明確分離兩個階段:(1) Planner LLM 預先建立完整計畫,(2) Executor LLM 執行每個步驟。每個步驟之後,Planner 可以重新規劃(調整計畫)。
Plan-and-Execute Architecture
══════════════════════════════════════════════════════════
User: "Analyze Q3 revenue, compare with Q2, write a report"
│
▼
┌─────────────────────────────────────────────┐
│ PLANNER LLM │
│ Plan: │
│ Step 1: Retrieve Q3 revenue data │
│ Step 2: Retrieve Q2 revenue data │
│ Step 3: Calculate Q2→Q3 change │
│ Step 4: Write comparison report │
└─────────────────────┬───────────────────────┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Execute S1 Execute S2 Execute S3 ...
(retriever) (retriever) (calculator)
│ │ │
└─────────────┼─────────────┘
│
▼
┌───────────────┐
│ REPLAN? │──► If step fails → adjust plan
│ All done? │──► If done → Step 4: report
└───────────────┘
2.3. LATS — Language Agent Tree Search
LATS 結合了蒙特卡羅樹搜尋(MCTS)與 LLM 推理。不同於 ReAct 只走一條路徑,LATS 探索多個解決方案分支,使用 LLM 評估每個分支,然後選擇最佳路徑。就像 LLM 下棋一樣 — 提前思考好幾步。
2.4. Reflexion — 從錯誤中學習
Reflexion 增加了自我反思步驟:完成任務後,代理自我評估結果 → 如果錯誤,將「經驗教訓」寫入記憶 → 帶著前次嘗試的經驗重試。這是一種透過自我回饋的上下文學習形式。
2.5. 認知架構比較
| 架構 | 規劃 | 執行 | 優勢 | 劣勢 |
|---|---|---|---|---|
| ReAct | 逐步(近視的) | 交替思考+行動 | 簡單、透明 | 無全局規劃,可能迴圈 |
| Plan-and-Execute | 預先完整規劃 | 循序執行 | 全局視角,LLM 呼叫較少 | 執行步驟後計畫可能過時 |
| LATS | 樹搜尋(探索) | 最佳優先搜尋 | 探索替代方案,穩健 | 非常昂貴(大量 LLM 呼叫) |
| Reflexion | 試錯 + 記憶 | 執行 → 反思 → 重試 | 從錯誤中學習 | 收斂速度慢,需要評估器 |
考試提示:「代理需要在執行前先想好完整計畫」→ Plan-and-Execute。「代理嘗試多條解決路徑,選擇最佳」→ LATS。「代理自我評估結果並改進」→ Reflexion。「代理交替思考與行動」→ ReAct。DLI 考試通常聚焦於 ReAct 和 Plan-and-Execute 作為兩種最廣泛使用的架構。
3. LangGraph — 有狀態的圖形化代理編排
3.1. 為什麼選擇 LangGraph?
LangChain 中的 AgentExecutor(第 8 課)是「黑盒子」— 難以自訂控制流。LangGraph 是一個 LangChain 函式庫,讓你將代理建構為有向圖:每個節點是一個處理步驟,邊定義流程,條件邊允許根據狀態進行分支。
| 特性 | AgentExecutor | LangGraph |
|---|---|---|
| 控制流 | 固定 ReAct 迴圈 | 自訂圖形 — 你設計流程 |
| 狀態管理 | 隱藏的內部狀態 | 明確的 TypedDict 狀態 |
| 多代理 | 無原生支援 | 一等支援:每個代理 = 子圖 |
| Human-in-the-loop | 有限 | 內建:中斷、批准、編輯 |
| 持久化 | 無內建 | Checkpointer:儲存/恢復狀態 |
| 串流 | 基本 | 逐事件串流 |
| 除錯 | 透過 LangSmith 追蹤 | 圖形視覺化 + LangSmith |
3.2. 核心概念 — StateGraph、節點、邊
LangGraph 建立在 3 個概念之上:
- State — 一個 TypedDict,保存所有在節點間傳遞的資料。每個節點讀取/寫入狀態。
- Nodes — Python 函式。輸入:state → 輸出:部分狀態更新(僅需要更新的欄位)。
- Edges — 節點間的連接。
add_edge(A, B)= 永遠走 A→B。add_conditional_edges(A, func)= func 決定走向。
LangGraph Concepts
══════════════════════════════════════════════════════════
State = TypedDict(messages, plan, results, ...)
────────────────────────────────────────────────
START ──► [Node: agent] ──conditional──► [Node: tools]
│ │
│ (if done) │ (tool result)
▼ │
END ◄───────────────────────────┘
Nodes: Python functions that read/write State
Edges: Static (always) or Conditional (function decides)
3.3. 程式碼:基本 LangGraph Agent
from typing import TypedDict, Annotated, Sequence
from langchain_core.messages import BaseMessage, HumanMessage, AIMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
import operator
# === 1. Define State ===
class AgentState(TypedDict):
messages: Annotated[Sequence[BaseMessage], operator.add]
# === 2. Define Tools ===
@tool
def search_docs(query: str) -> str:
"""Search internal documents for company information."""
# Simulate retrieval
docs = {
"leave": "Employees get 12 days annual leave per year.",
"refund": "Refund within 30 days with original receipt.",
}
for key, val in docs.items():
if key in query.lower():
return val
return "No relevant documents found."
@tool
def calculator(expression: str) -> str:
"""Calculate mathematical expressions."""
try:
return str(eval(expression)) # production: use safe eval
except Exception as e:
return f"Error: {e}"
tools = [search_docs, calculator]
# === 3. Define LLM with tools ===
llm = ChatNVIDIA(
model="meta/llama-3.1-70b-instruct",
temperature=0.1
).bind_tools(tools)
# === 4. Define Nodes ===
def agent_node(state: AgentState) -> dict:
"""LLM decides: call tool or respond."""
response = llm.invoke(state["messages"])
return {"messages": [response]}
tool_node = ToolNode(tools)
# === 5. Define Routing ===
def should_continue(state: AgentState) -> str:
last_message = state["messages"][-1]
if last_message.tool_calls:
return "tools" # LLM wants to call a tool
return "end" # LLM is done, return answer
# === 6. Build Graph ===
graph = StateGraph(AgentState)
graph.add_node("agent", agent_node)
graph.add_node("tools", tool_node)
graph.set_entry_point("agent")
graph.add_conditional_edges("agent", should_continue, {
"tools": "tools",
"end": END,
})
graph.add_edge("tools", "agent") # after tool → back to agent
app = graph.compile()
# === 7. Run ===
result = app.invoke({
"messages": [HumanMessage(content="What is the leave policy?")]
})
print(result["messages"][-1].content)
3.4. Human-in-the-Loop
LangGraph 支援在執行危險操作前中斷 — 例如發送電子郵件、刪除資料、執行程式碼。代理暫停,等待使用者批准,然後繼續。
from langgraph.checkpoint.memory import MemorySaver
# Compile with checkpointer for interruption + resume
checkpointer = MemorySaver()
app = graph.compile(
checkpointer=checkpointer,
interrupt_before=["tools"] # pause BEFORE executing tools
)
# Run — will pause before the "tools" node
config = {"configurable": {"thread_id": "user-123"}}
result = app.invoke(
{"messages": [HumanMessage(content="Delete file report.pdf")]},
config=config,
)
# Inspect pending tool call
pending = result["messages"][-1].tool_calls
print(f"Agent wants to: {pending}")
# → Agent wants to: [{'name': 'delete_file', 'args': {'path': 'report.pdf'}}]
# Human approves → continue
final = app.invoke(None, config=config) # resume from checkpoint
3.5. Checkpointing — 儲存與恢復狀態
Checkpointer 在每個節點後儲存狀態,提供以下功能:
- 恢復 — 代理執行中途崩潰 → 載入檢查點 → 繼續
- 時間旅行 — 回到任何檢查點 → 以不同輸入重試
- Human-in-the-loop — 暫停、等待使用者、恢復(如上所示)
- 多輪對話 — 跨多輪維護對話歷史
考試提示:「建構具有自訂控制流、條件分支的代理」→ LangGraph(而非 AgentExecutor)。「暫停代理執行以等待人工批准」→ interrupt_before + checkpointer。「儲存代理狀態,稍後恢復」→ LangGraph checkpointing。DLI C-FX-25 常問:「為什麼用 LangGraph 而不是 AgentExecutor?」→ 自訂流程、多代理、持久化、Human-in-the-loop。
4. 多代理模式
4.1. 為什麼需要多代理?
單一代理配備 20+ 工具會面臨問題:工具選擇混淆(工具太多,LLM 選錯),提示過長(必須包含所有指令),難以除錯(不清楚代理在哪裡失敗)。多代理透過分解來解決:每個代理專精一項任務,配備較少的工具。
4.2. 監督者模式
一個監督者代理(LLM)從使用者接收任務,委派給工作者代理,收集結果,並綜合最終答案。
Supervisor Pattern
══════════════════════════════════════════════════════════
┌──────────────┐
│ USER │
└──────┬───────┘
│
▼
┌────────────────────────┐
│ SUPERVISOR AGENT │
│ (Orchestrator LLM) │
│ │
│ Decides: │
│ • Which worker next? │
│ • All done? │
│ • Need to re-route? │
└────┬──────┬──────┬─────┘
│ │ │
┌────────┘ │ └────────┐
▼ ▼ ▼
┌─────────────┐┌─────────────┐┌─────────────┐
│ Researcher ││ Coder ││ Reporter │
│ Agent ││ Agent ││ Agent │
│ ││ ││ │
│ Tools: ││ Tools: ││ Tools: │
│ • web_search││ • python ││ • write_doc │
│ • doc_search││ • shell ││ • format │
└─────────────┘└─────────────┘└─────────────┘
4.3. 階層式模式
階層式擴展了監督者模式:每個工作者本身可以是子工作者的監督者。適合複雜的組織結構 — 例如 CEO 代理 → 經理代理 → 專家代理。
Hierarchical Multi-Agent
══════════════════════════════════════════════════════════
┌──────────────────────┐
│ TOP SUPERVISOR │
│ (Project Manager) │
└───┬─────────────┬────┘
│ │
┌────────┘ └────────┐
▼ ▼
┌──────────────┐ ┌──────────────┐
│ RESEARCH │ │ ENGINEERING │
│ SUPERVISOR │ │ SUPERVISOR │
└──┬───────┬───┘ └──┬───────┬───┘
│ │ │ │
▼ ▼ ▼ ▼
[Web [Paper [Backend [Frontend
Searcher] Analyzer] Dev] Dev]
4.4. 群集模式
Swarm(OpenAI Swarm 概念)— 無監督者。代理根據上下文互相移交。Agent A 意識到「這個任務屬於 Agent B 的專長」→ 自動移交。
4.5. 辯論模式
Debate — 兩個或多個代理就一個問題進行辯論。每個代理提出觀點並反駁對方。最後由 Judge 代理選擇最佳結論。此模式提升了複雜問題的推理品質。
4.6. 多代理模式比較
| 模式 | 控制流 | 通訊方式 | 最適用場景 |
|---|---|---|---|
| 監督者 | 集中式 — 監督者路由 | 星形拓撲 | 明確的任務委派,中等複雜度 |
| 階層式 | 多層級監督 | 樹狀結構 | 複雜組織,多個專業化子團隊 |
| 群集 | 去中心化 — 代理互相移交 | 對等網路 | 客戶服務、路由、彈性流程 |
| 辯論 | 輪流論證 | 廣播 + 裁判 | 複雜推理、事實驗證 |
考試提示:「一個 LLM 將任務路由給專業化代理」→ 監督者。「代理在無中央控制的情況下互相移交」→ Swarm。「多個代理辯論,裁判做決定」→ 辯論。「巢狀監督者管理子團隊」→ 階層式。DLI 考試通常聚焦於監督者模式,因為它在生產環境中最常見。
4.7. 程式碼:使用 LangGraph 的監督者多代理系統
from typing import TypedDict, Annotated, Literal, Sequence
from langchain_core.messages import BaseMessage, HumanMessage, SystemMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
import operator
# === State ===
class MultiAgentState(TypedDict):
messages: Annotated[Sequence[BaseMessage], operator.add]
next_agent: str
# === Worker Tools ===
@tool
def web_search(query: str) -> str:
"""Search the web for current information."""
return f"[Web Result] Top findings for '{query}': ..."
@tool
def run_python(code: str) -> str:
"""Execute Python code and return output."""
try:
exec_globals = {}
exec(code, exec_globals)
return str(exec_globals.get("result", "Code executed successfully."))
except Exception as e:
return f"Error: {e}"
@tool
def write_report(content: str) -> str:
"""Format content into a professional report."""
return f"=== REPORT ===\n{content}\n=== END ==="
# === Worker Agents ===
researcher_llm = ChatNVIDIA(
model="meta/llama-3.1-70b-instruct", temperature=0.1
).bind_tools([web_search])
coder_llm = ChatNVIDIA(
model="meta/llama-3.1-70b-instruct", temperature=0.0
).bind_tools([run_python])
reporter_llm = ChatNVIDIA(
model="meta/llama-3.1-70b-instruct", temperature=0.3
).bind_tools([write_report])
# === Supervisor ===
supervisor_llm = ChatNVIDIA(
model="meta/llama-3.1-70b-instruct", temperature=0.0
)
WORKERS = ["researcher", "coder", "reporter"]
def supervisor_node(state: MultiAgentState) -> dict:
"""Supervisor decides which worker to route to next."""
system_prompt = f"""You are a supervisor managing these workers: {WORKERS}.
Given the conversation, decide which worker should act next,
or if the task is complete respond with FINISH.
Respond with ONLY the worker name or FINISH."""
messages = [SystemMessage(content=system_prompt)] + state["messages"]
response = supervisor_llm.invoke(messages)
next_agent = response.content.strip().lower()
if next_agent not in WORKERS:
next_agent = "FINISH"
return {"next_agent": next_agent}
def researcher_node(state: MultiAgentState) -> dict:
system = SystemMessage(content="You are a research specialist. "
"Use web_search to find information. Be thorough.")
response = researcher_llm.invoke([system] + state["messages"])
return {"messages": [response]}
def coder_node(state: MultiAgentState) -> dict:
system = SystemMessage(content="You are a Python coding specialist. "
"Use run_python to execute code for analysis and calculations.")
response = coder_llm.invoke([system] + state["messages"])
return {"messages": [response]}
def reporter_node(state: MultiAgentState) -> dict:
system = SystemMessage(content="You are a report writer. "
"Use write_report to create formatted reports from gathered info.")
response = reporter_llm.invoke([system] + state["messages"])
return {"messages": [response]}
# === Routing ===
def route_supervisor(state: MultiAgentState) -> str:
next_agent = state.get("next_agent", "FINISH")
if next_agent == "FINISH":
return "end"
return next_agent
# === Build Graph ===
graph = StateGraph(MultiAgentState)
graph.add_node("supervisor", supervisor_node)
graph.add_node("researcher", researcher_node)
graph.add_node("coder", coder_node)
graph.add_node("reporter", reporter_node)
graph.add_node("researcher_tools", ToolNode([web_search]))
graph.add_node("coder_tools", ToolNode([run_python]))
graph.add_node("reporter_tools", ToolNode([write_report]))
graph.set_entry_point("supervisor")
# Supervisor routes to workers
graph.add_conditional_edges("supervisor", route_supervisor, {
"researcher": "researcher",
"coder": "coder",
"reporter": "reporter",
"end": END,
})
# Workers → tool nodes → back to supervisor
for worker in WORKERS:
def make_router(w):
def router(state):
last = state["messages"][-1]
if hasattr(last, "tool_calls") and last.tool_calls:
return f"{w}_tools"
return "supervisor"
return router
graph.add_conditional_edges(worker, make_router(worker), {
f"{worker}_tools": f"{worker}_tools",
"supervisor": "supervisor",
})
graph.add_edge(f"{worker}_tools", worker)
app = graph.compile()
# === Run ===
result = app.invoke({
"messages": [HumanMessage(
content="Research NVIDIA H100 GPU specs, calculate price-performance "
"ratio vs A100, and write a comparison report."
)],
"next_agent": "",
})
for msg in result["messages"]:
print(f"[{msg.type}] {msg.content[:200]}...")
5. 建構生產環境的多代理應用程式
5.1. 研究助手 — 完整範例
建構一個完整的研究助手,包含 3 個代理:研究者(尋找資訊)、程式設計師(分析資料)、報告撰寫者(撰寫報告)。包含錯誤處理、重試邏輯和結構化輸出。
from typing import TypedDict, Annotated, Sequence, Optional
from langchain_core.messages import BaseMessage, HumanMessage, SystemMessage, AIMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.checkpoint.memory import MemorySaver
import operator
import json
# === Enhanced State ===
class ResearchState(TypedDict):
messages: Annotated[Sequence[BaseMessage], operator.add]
research_data: Optional[str] # collected research
analysis_result: Optional[str] # code analysis output
final_report: Optional[str] # formatted report
current_agent: str
iteration: int # track iterations to prevent loops
MAX_ITERATIONS = 10
# === Tools ===
@tool
def search_arxiv(query: str) -> str:
"""Search academic papers on arxiv for research topics."""
return json.dumps({
"papers": [
{"title": f"Paper on {query}", "abstract": f"Study of {query}...",
"year": 2025, "citations": 42},
]
})
@tool
def search_web(query: str) -> str:
"""Search the web for current news, blog posts, documentation."""
return json.dumps({
"results": [
{"title": f"Latest news: {query}", "snippet": f"Updated info on {query}..."},
]
})
@tool
def execute_analysis(code: str) -> str:
"""Run Python code for data analysis. Variable 'result' will be returned."""
exec_globals = {}
try:
exec(code, exec_globals)
return str(exec_globals.get("result", "Executed OK, no 'result' variable."))
except Exception as e:
return f"Error: {e}"
@tool
def generate_report(title: str, sections: str) -> str:
"""Generate a formatted markdown report from title and section content."""
return f"# {title}\n\n{sections}\n\n---\nGenerated by Research Assistant"
# === Agent Nodes ===
base_llm = ChatNVIDIA(model="meta/llama-3.1-70b-instruct")
def researcher_node(state: ResearchState) -> dict:
llm = base_llm.bind_tools([search_arxiv, search_web])
system = SystemMessage(content=(
"You are a research specialist. Search for papers and web results "
"to gather comprehensive information. Summarize findings clearly."
))
response = llm.invoke([system] + list(state["messages"]))
# If no tool calls, research is done — extract data
if not response.tool_calls:
return {
"messages": [response],
"research_data": response.content,
"current_agent": "supervisor",
}
return {"messages": [response], "current_agent": "researcher_tools"}
def coder_node(state: ResearchState) -> dict:
llm = base_llm.bind_tools([execute_analysis])
context = state.get("research_data", "No research data yet.")
system = SystemMessage(content=(
f"You are a data analyst. Use the research data below to perform "
f"analysis with Python code.\n\nResearch Data:\n{context}"
))
response = llm.invoke([system] + list(state["messages"]))
if not response.tool_calls:
return {
"messages": [response],
"analysis_result": response.content,
"current_agent": "supervisor",
}
return {"messages": [response], "current_agent": "coder_tools"}
def reporter_node(state: ResearchState) -> dict:
llm = base_llm.bind_tools([generate_report])
research = state.get("research_data", "N/A")
analysis = state.get("analysis_result", "N/A")
system = SystemMessage(content=(
f"You are a report writer. Create a professional report.\n"
f"Research:\n{research}\n\nAnalysis:\n{analysis}"
))
response = llm.invoke([system] + list(state["messages"]))
if not response.tool_calls:
return {
"messages": [response],
"final_report": response.content,
"current_agent": "supervisor",
}
return {"messages": [response], "current_agent": "reporter_tools"}
def supervisor_node(state: ResearchState) -> dict:
iteration = state.get("iteration", 0) + 1
if iteration > MAX_ITERATIONS:
return {
"messages": [AIMessage(content="Max iterations reached. Returning results.")],
"current_agent": "FINISH",
"iteration": iteration,
}
system = SystemMessage(content="""You are a project supervisor. Based on the current state:
- If no research data → route to "researcher"
- If research done but no analysis → route to "coder"
- If analysis done but no report → route to "reporter"
- If report is ready → respond "FINISH"
Respond with ONLY one of: researcher, coder, reporter, FINISH""")
response = base_llm.invoke([system] + list(state["messages"]))
next_agent = response.content.strip().lower()
valid = ["researcher", "coder", "reporter", "finish"]
if next_agent not in valid:
next_agent = "researcher" # default fallback
return {"current_agent": next_agent, "iteration": iteration}
# === Routing ===
def route_from_supervisor(state: ResearchState) -> str:
agent = state.get("current_agent", "FINISH")
if agent in ["researcher", "coder", "reporter"]:
return agent
return "end"
def route_from_worker(worker_name: str):
def router(state: ResearchState) -> str:
current = state.get("current_agent", "supervisor")
if current == f"{worker_name}_tools":
return f"{worker_name}_tools"
return "supervisor"
return router
# === Build Graph ===
from langgraph.prebuilt import ToolNode
graph = StateGraph(ResearchState)
graph.add_node("supervisor", supervisor_node)
graph.add_node("researcher", researcher_node)
graph.add_node("coder", coder_node)
graph.add_node("reporter", reporter_node)
graph.add_node("researcher_tools", ToolNode([search_arxiv, search_web]))
graph.add_node("coder_tools", ToolNode([execute_analysis]))
graph.add_node("reporter_tools", ToolNode([generate_report]))
graph.set_entry_point("supervisor")
graph.add_conditional_edges("supervisor", route_from_supervisor, {
"researcher": "researcher",
"coder": "coder",
"reporter": "reporter",
"end": END,
})
for worker in ["researcher", "coder", "reporter"]:
graph.add_conditional_edges(worker, route_from_worker(worker), {
f"{worker}_tools": f"{worker}_tools",
"supervisor": "supervisor",
})
graph.add_edge(f"{worker}_tools", worker)
# Compile with checkpointing
checkpointer = MemorySaver()
app = graph.compile(checkpointer=checkpointer)
# === Execute ===
config = {"configurable": {"thread_id": "research-001"}}
result = app.invoke(
{
"messages": [HumanMessage(
content="Research the latest advances in mixture-of-experts (MoE) "
"models, analyze their parameter efficiency compared to "
"dense models, and write a summary report."
)],
"current_agent": "",
"iteration": 0,
},
config=config,
)
# Print final report
print(result.get("final_report", result["messages"][-1].content))
5.2. 生產環境最佳實踐
| 實踐 | 原因 | 實作方式 |
|---|---|---|
| 最大迭代次數 | 防止無限迴圈 | 在狀態中設置 iteration 計數器,在監督者處檢查 |
| 錯誤處理 | 工具失敗不應導致代理崩潰 | 工具中使用 try/except,回傳錯誤訊息 |
| 檢查點 | 崩潰後恢復 | MemorySaver(開發)/ SqliteSaver(生產) |
| 結構化輸出 | 可靠的路由決策 | 限制監督者輸出為有效選項 |
| 可觀測性 | 多代理除錯困難 | LangSmith 追蹤,記錄每個節點進出 |
| 每節點超時 | 單一節點不應阻塞 | 對 LLM 呼叫和工具執行設定超時 |
| Human-in-the-loop | 關鍵操作需要批准 | 在危險的工具節點使用 interrupt_before |
考試提示:「如何防止代理無限迴圈?」→ max_iterations + 迭代計數器。「如何除錯多代理系統?」→ LangSmith 追蹤 + 日誌記錄。「代理崩潰恢復?」→ 使用持久化儲存的 Checkpointing。生產部署 → LangGraph Platform(託管)或 LangServe(自行部署)。
6. DLI C-FX-25 — Agentic AI 課程總覽
6.1. 課程結構
課程 C-FX-25:「Building Agentic AI Applications」是 DLI 的進階模組,專注於建構生產環境的 Agentic AI 系統。它補充了 S-FX-15,深入探討代理架構。
| 模組 | 主題 | 實作練習 |
|---|---|---|
| 模組 1 | Agent 基礎、ReAct、工具呼叫 | 使用 NVIDIA NIM 建構單一代理 |
| 模組 2 | LangGraph 入門、StateGraph | 實作自訂代理圖 |
| 模組 3 | 多代理架構 | 建構監督者多代理系統 |
| 模組 4 | 進階:記憶、規劃、評估 | 生產部署練習 |
6.2. 評量重點領域
C-FX-25 評量聚焦於實作:
- LangGraph StateGraph — 定義狀態、節點、條件邊
- 工具整合 — 將工具綁定到 LLM、處理工具呼叫
- 監督者路由 — 實作監督者邏輯、路由至工作者
- 檢查點 — 儲存/恢復代理狀態
- Human-in-the-loop — interrupt_before、批准、恢復
6.3. 需要記住的關鍵 API
| API / 概念 | 用途 |
|---|---|
StateGraph(State) | 建立帶有型別狀態的圖 |
graph.add_node(name, func) | 新增處理節點 |
graph.add_edge(A, B) | 永遠路由 A → B |
graph.add_conditional_edges(A, func, map) | 根據函式輸出路由 |
graph.set_entry_point(name) | 設定起始節點 |
graph.compile(checkpointer=...) | 編譯圖,可選檢查點 |
ToolNode(tools) | LangGraph 預建節點,用於執行工具呼叫 |
MemorySaver() | 記憶體內檢查點(僅限開發) |
interrupt_before=[node] | 在執行節點前暫停 |
llm.bind_tools(tools) | 將工具附加到 LLM 以進行 function calling |
考試提示:C-FX-25 評量要求從零開始撰寫 LangGraph 程式碼。記住這個模式:(1) 定義 State TypedDict,(2) 定義節點為函式,(3) 新增節點 + 邊,(4) 編譯 + 執行。你不需要背誦 API,但必須理解流程:狀態在節點間傳遞,條件邊動態路由。
7. 速查表
| 概念 | 重點 |
|---|---|
| Agent 元件 | LLM + 記憶 + 工具 + 規劃 |
| Agent 迴圈 | 感知 → 推理 → 行動 → 觀察 |
| Agent vs Chain | Agent = 動態流程(LLM 決定);Chain = 固定流程 |
| ReAct | 交替思考 + 行動 + 觀察。簡單,近視的 |
| Plan-and-Execute | 預先規劃 → 執行步驟 → 需要時重新規劃 |
| LATS | 推理路徑的樹搜尋。昂貴但穩健 |
| Reflexion | 執行 → 自我反思 → 帶著經驗重試 |
| LangGraph | StateGraph:節點 + 邊 + 條件路由 |
| LangGraph State | 所有節點共享的 TypedDict |
| 條件邊 | 路由函式決定下一個節點 |
| 檢查點 | MemorySaver(開發)、SqliteSaver(生產)。啟用恢復 |
| Human-in-the-loop | interrupt_before=[node] + 使用 checkpointer 編譯 |
| 監督者模式 | 中央 LLM 路由至專業化工作者代理 |
| 階層式 | 巢狀監督者 — 代理樹 |
| 群集 | 去中心化移交,無中央監督者 |
| 辯論 | 代理辯論,裁判決定。更好的推理 |
| 最大迭代次數 | 始終設定以防止代理無限迴圈 |
| ToolNode | LangGraph 預建節點,用於執行工具呼叫 |
| C-FX-25 重點 | LangGraph 程式碼、多代理、檢查點、HITL |
8. 練習題 — 程式碼
Q1:建構基本的 LangGraph ReAct 代理
建構一個簡單的 LangGraph 代理,包含 2 個工具:search_docs(搜尋文件)和 calculator(執行計算)。實作完整流程:State、agent 節點、tool 節點、條件邊路由、編譯並執行。
顯示答案 Q1
from typing import TypedDict, Annotated, Sequence
from langchain_core.messages import BaseMessage, HumanMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
import operator
# 1. State
class AgentState(TypedDict):
messages: Annotated[Sequence[BaseMessage], operator.add]
# 2. Tools
@tool
def search_docs(query: str) -> str:
"""Search internal knowledge base for relevant documents."""
return f"Found: Documentation about {query} — key facts here."
@tool
def calculator(expression: str) -> str:
"""Calculate a mathematical expression."""
return str(eval(expression))
tools = [search_docs, calculator]
# 3. LLM with tools
llm = ChatNVIDIA(model="meta/llama-3.1-70b-instruct", temperature=0.0)
llm_with_tools = llm.bind_tools(tools)
# 4. Nodes
def agent_node(state: AgentState) -> dict:
response = llm_with_tools.invoke(state["messages"])
return {"messages": [response]}
tool_node = ToolNode(tools)
# 5. Router
def should_continue(state: AgentState) -> str:
last = state["messages"][-1]
if hasattr(last, "tool_calls") and last.tool_calls:
return "tools"
return "end"
# 6. Build graph
graph = StateGraph(AgentState)
graph.add_node("agent", agent_node)
graph.add_node("tools", tool_node)
graph.set_entry_point("agent")
graph.add_conditional_edges("agent", should_continue, {
"tools": "tools",
"end": END,
})
graph.add_edge("tools", "agent")
app = graph.compile()
# 7. Run
result = app.invoke({
"messages": [HumanMessage(content="What is 25 * 4 + 100?")]
})
print(result["messages"][-1].content)
Q2:使用 LangGraph 檢查點實作 Human-in-the-loop
修改 Q1 的代理以新增 Human-in-the-loop:代理在執行工具前暫停,使用者可以批准或拒絕。示範:(1) 使用 checkpointer + interrupt_before 編譯,(2) 執行並看到代理暫停,(3) 恢復執行。
顯示答案 Q2
from langgraph.checkpoint.memory import MemorySaver
# Reuse graph from Q1, compile with HITL
checkpointer = MemorySaver()
app_hitl = graph.compile(
checkpointer=checkpointer,
interrupt_before=["tools"] # Pause BEFORE tool execution
)
# Run — agent will pause before calling tools
config = {"configurable": {"thread_id": "hitl-demo-001"}}
result = app_hitl.invoke(
{"messages": [HumanMessage(content="Calculate 1000 / 4")]},
config=config,
)
# Agent paused — inspect what it wants to do
last_msg = result["messages"][-1]
print("Agent wants to call:")
for tc in last_msg.tool_calls:
print(f" Tool: {tc['name']}, Args: {tc['args']}")
# User approves → resume (pass None to continue from checkpoint)
final_result = app_hitl.invoke(None, config=config)
print("\nFinal answer:", final_result["messages"][-1].content)
# If user REJECTS → could modify state or stop here
# To reject: simply don't call invoke(None, config)
Q3:建構監督者多代理系統
實作一個監督者模式,包含 2 個工作者:researcher(使用 web_search 工具)和 writer(使用 write_report 工具)。監督者從使用者接收任務,路由至適當的工作者,並收集結果。使用條件邊實作路由邏輯。
顯示答案 Q3
from typing import TypedDict, Annotated, Sequence
from langchain_core.messages import BaseMessage, HumanMessage, SystemMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
import operator
class SupervisorState(TypedDict):
messages: Annotated[Sequence[BaseMessage], operator.add]
next: str
@tool
def web_search(query: str) -> str:
"""Search the internet for information."""
return f"Search results for '{query}': ..."
@tool
def write_report(content: str) -> str:
"""Write and format a professional report."""
return f"=== Report ===\n{content}\n=== End ==="
llm = ChatNVIDIA(model="meta/llama-3.1-70b-instruct", temperature=0.0)
def supervisor(state: SupervisorState) -> dict:
sys = SystemMessage(content=(
"You are a supervisor. Workers: researcher, writer. "
"Route to appropriate worker or say FINISH if task complete. "
"Respond with ONLY: researcher, writer, or FINISH."
))
resp = llm.invoke([sys] + list(state["messages"]))
next_val = resp.content.strip().lower()
if next_val not in ["researcher", "writer"]:
next_val = "FINISH"
return {"next": next_val}
def researcher(state: SupervisorState) -> dict:
r_llm = llm.bind_tools([web_search])
sys = SystemMessage(content="You are a researcher. Use web_search.")
resp = r_llm.invoke([sys] + list(state["messages"]))
return {"messages": [resp]}
def writer(state: SupervisorState) -> dict:
w_llm = llm.bind_tools([write_report])
sys = SystemMessage(content="You are a report writer. Use write_report.")
resp = w_llm.invoke([sys] + list(state["messages"]))
return {"messages": [resp]}
def route(state: SupervisorState) -> str:
n = state.get("next", "FINISH")
return n if n in ["researcher", "writer"] else "end"
# Build
g = StateGraph(SupervisorState)
g.add_node("supervisor", supervisor)
g.add_node("researcher", researcher)
g.add_node("writer", writer)
g.add_node("research_tools", ToolNode([web_search]))
g.add_node("writer_tools", ToolNode([write_report]))
g.set_entry_point("supervisor")
g.add_conditional_edges("supervisor", route, {
"researcher": "researcher",
"writer": "writer",
"end": END,
})
# Researcher flow
def route_researcher(state):
last = state["messages"][-1]
if hasattr(last, "tool_calls") and last.tool_calls:
return "research_tools"
return "supervisor"
g.add_conditional_edges("researcher", route_researcher, {
"research_tools": "research_tools",
"supervisor": "supervisor",
})
g.add_edge("research_tools", "researcher")
# Writer flow
def route_writer(state):
last = state["messages"][-1]
if hasattr(last, "tool_calls") and last.tool_calls:
return "writer_tools"
return "supervisor"
g.add_conditional_edges("writer", route_writer, {
"writer_tools": "writer_tools",
"supervisor": "supervisor",
})
g.add_edge("writer_tools", "writer")
app = g.compile()
result = app.invoke({
"messages": [HumanMessage(content="Research AI trends 2025 and write a report")],
"next": "",
})
print(result["messages"][-1].content)
Q4:為 LangGraph 代理新增 Plan-and-Execute
實作 Plan-and-Execute 模式:(1) Planner 節點從使用者查詢建立步驟列表,(2) Executor 節點執行每個步驟,(3) Replanner 節點檢查進度並在需要時調整計畫。將計畫儲存在狀態中。
顯示答案 Q4
from typing import TypedDict, Annotated, Sequence, List, Optional
from langchain_core.messages import BaseMessage, HumanMessage, SystemMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langgraph.graph import StateGraph, END
import operator, json
class PlanExecState(TypedDict):
messages: Annotated[Sequence[BaseMessage], operator.add]
plan: List[str] # list of steps
current_step: int # index of current step
step_results: List[str] # results of each step
done: bool
llm = ChatNVIDIA(model="meta/llama-3.1-70b-instruct", temperature=0.0)
def planner_node(state: PlanExecState) -> dict:
"""Create a plan from user request."""
sys = SystemMessage(content=(
"You are a planner. Break the user's request into 3-5 concrete steps. "
"Return ONLY a JSON array of strings, e.g. [\"step1\", \"step2\"]."
))
resp = llm.invoke([sys] + list(state["messages"]))
try:
plan = json.loads(resp.content)
except json.JSONDecodeError:
plan = [resp.content]
return {"plan": plan, "current_step": 0, "step_results": []}
def executor_node(state: PlanExecState) -> dict:
"""Execute the current step of the plan."""
step_idx = state["current_step"]
plan = state["plan"]
if step_idx >= len(plan):
return {"done": True}
current = plan[step_idx]
sys = SystemMessage(content=(
f"Execute this step: {current}\n"
f"Previous results: {state['step_results']}\n"
"Provide a concise result."
))
resp = llm.invoke([sys] + list(state["messages"]))
new_results = list(state["step_results"]) + [resp.content]
return {
"step_results": new_results,
"current_step": step_idx + 1,
"messages": [resp],
}
def replanner_node(state: PlanExecState) -> dict:
"""Check progress, adjust plan if needed."""
if state["current_step"] >= len(state["plan"]):
return {"done": True}
sys = SystemMessage(content=(
f"Plan: {state['plan']}\n"
f"Completed: {state['current_step']}/{len(state['plan'])}\n"
f"Results so far: {state['step_results']}\n"
"Should the remaining plan continue as-is? "
"Reply 'CONTINUE' or provide updated remaining steps as JSON array."
))
resp = llm.invoke([sys])
if "CONTINUE" in resp.content.upper():
return {"done": False}
try:
remaining = json.loads(resp.content)
new_plan = state["plan"][:state["current_step"]] + remaining
return {"plan": new_plan, "done": False}
except json.JSONDecodeError:
return {"done": False}
def route_after_exec(state: PlanExecState) -> str:
if state.get("done", False):
return "end"
return "replanner"
def route_after_replan(state: PlanExecState) -> str:
if state.get("done", False):
return "end"
return "executor"
# Build graph
g = StateGraph(PlanExecState)
g.add_node("planner", planner_node)
g.add_node("executor", executor_node)
g.add_node("replanner", replanner_node)
g.set_entry_point("planner")
g.add_edge("planner", "executor")
g.add_conditional_edges("executor", route_after_exec, {
"replanner": "replanner",
"end": END,
})
g.add_conditional_edges("replanner", route_after_replan, {
"executor": "executor",
"end": END,
})
app = g.compile()
result = app.invoke({
"messages": [HumanMessage(
content="Analyze the pros and cons of microservices architecture "
"and recommend when to use it vs monolith."
)],
"plan": [],
"current_step": 0,
"step_results": [],
"done": False,
})
for i, res in enumerate(result["step_results"]):
print(f"Step {i+1}: {res[:150]}...")
Q5:在多代理系統中實作錯誤處理和重試邏輯
為多代理系統新增錯誤處理:(1) 工具失敗回傳錯誤訊息而非崩潰,(2) 代理收到錯誤 → 以不同策略重試(最多 2 次重試),(3) 代理狀態追蹤重試次數。實作節點包裝器模式。
顯示答案 Q5
from typing import TypedDict, Annotated, Sequence, Dict
from langchain_core.messages import BaseMessage, HumanMessage, AIMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
import operator
class RobustState(TypedDict):
messages: Annotated[Sequence[BaseMessage], operator.add]
error_count: int
max_retries: int
# Tools with error handling built-in
@tool
def risky_api_call(endpoint: str) -> str:
"""Call an external API that might fail."""
import random
if random.random() < 0.5:
raise ConnectionError(f"API {endpoint} unreachable")
return f"API response from {endpoint}: success data"
@tool
def safe_search(query: str) -> str:
"""Search with built-in error handling."""
return f"Results for {query}: ..."
# Wrap tools with error handling
def safe_tool_node(tools):
"""ToolNode wrapper that catches errors and returns error messages."""
base_node = ToolNode(tools)
def wrapper(state: RobustState) -> dict:
try:
return base_node.invoke(state)
except Exception as e:
error_msg = AIMessage(content=f"Tool error: {str(e)}. Try different approach.")
return {
"messages": [error_msg],
"error_count": state.get("error_count", 0) + 1,
}
return wrapper
tools = [risky_api_call, safe_search]
llm = ChatNVIDIA(model="meta/llama-3.1-70b-instruct", temperature=0.0)
llm_with_tools = llm.bind_tools(tools)
def agent_node(state: RobustState) -> dict:
error_count = state.get("error_count", 0)
max_retries = state.get("max_retries", 2)
# If too many errors, give up gracefully
if error_count >= max_retries:
return {"messages": [AIMessage(
content="I encountered multiple errors. Here's what I could gather "
"from successful attempts: " +
" | ".join(m.content for m in state["messages"][-3:])
)]}
# Add retry context if there were errors
msgs = list(state["messages"])
if error_count > 0:
msgs.append(HumanMessage(
content=f"Previous attempt failed ({error_count}/{max_retries} retries). "
"Try a different tool or approach."
))
response = llm_with_tools.invoke(msgs)
return {"messages": [response]}
def should_continue(state: RobustState) -> str:
last = state["messages"][-1]
error_count = state.get("error_count", 0)
max_retries = state.get("max_retries", 2)
# Stop if max retries exceeded
if error_count >= max_retries and not (
hasattr(last, "tool_calls") and last.tool_calls
):
return "end"
if hasattr(last, "tool_calls") and last.tool_calls:
return "tools"
return "end"
# Build
g = StateGraph(RobustState)
g.add_node("agent", agent_node)
g.add_node("tools", safe_tool_node(tools))
g.set_entry_point("agent")
g.add_conditional_edges("agent", should_continue, {
"tools": "tools",
"end": END,
})
g.add_edge("tools", "agent") # tool result → back to agent
app = g.compile()
result = app.invoke({
"messages": [HumanMessage(content="Call the user-data API endpoint")],
"error_count": 0,
"max_retries": 2,
})
print(result["messages"][-1].content)