Chuyển đến nội dung chính

Bài 9: Agentic AI & Multi-Agent Systems

Agent abstraction: perception → reasoning → action loop. Cognitive architectures: ReAct, Plan-and-Execute, LATS. LangGraph: stateful graph-based agent orchestration. Multi-agent systems: supervisor, hierarchical, swarm patterns. Build production-ready multi-agent application.

1. Agent Abstraction — LLM + Memory + Tools + Planning

1.1. Agent là gì?

Bài 8 đã giới thiệu ReAct agent đơn giản — LLM chọn tool rồi trả lời. Bài 9 mở rộng sang Agentic AI: hệ thống nơi LLM đóng vai trò bộ não trung tâm — tự lên kế hoạch, chọn công cụ, phản hồi từ kết quả, và phối hợp với nhiều agent khác.

Một Agent gồm 4 thành phần cốt lõi:

  • LLM (Brain) — suy luận, ra quyết định, sinh text
  • Memory — ngắn hạn (conversation buffer) và dài hạn (vector store, database)
  • Tools — hàm mà agent gọi được: search, calculator, API, code execution
  • Planning — lên kế hoạch (plan), chia nhỏ task, phản hồi (reflect) khi có lỗi

Agent Abstraction — Core Components
══════════════════════════════════════════════════════════════

                    ┌──────────────────────┐
                    │       USER           │
                    │   (Task / Query)     │
                    └──────────┬───────────┘
                               │
                               ▼
  ┌────────────────────────────────────────────────────────┐
  │                      AGENT                             │
  │  ┌──────────┐  ┌───────────┐  ┌──────────────────┐    │
  │  │ Planning │  │  LLM Core │  │     Memory       │    │
  │  │          │◄─┤  (Brain)  ├─►│ Short-term: chat │    │
  │  │ Decompose│  │ Reasoning │  │ Long-term: VDB   │    │
  │  │ Reflect  │  │ Decisions │  │ Episodic: logs   │    │
  │  └──────────┘  └─────┬─────┘  └──────────────────┘    │
  │                       │                                │
  │              ┌────────┼────────┐                       │
  │              ▼        ▼        ▼                       │
  │         ┌────────┐┌───────┐┌────────┐                  │
  │         │Search  ││ Code  ││  API   │  ◄── Tools      │
  │         │Engine  ││ Exec  ││ Calls  │                  │
  │         └────────┘└───────┘└────────┘                  │
  └────────────────────────────────────────────────────────┘

1.2. Perception → Reasoning → Action → Observation Loop

Mọi agent đều chạy theo một vòng lặp cơ bản:

  1. Perception — nhận input (user query, tool output, environment feedback)
  2. Reasoning — LLM suy luận: "Tôi cần làm gì tiếp? Dùng tool nào? Đã đủ info chưa?"
  3. Action — thực hiện hành động: gọi tool, sinh text, trả lời user
  4. Observation — nhận kết quả từ action, đưa lại vào bước Perception → lặp lại

Agent Loop — Perception → Reasoning → Action → Observation
═══════════════════════════════════════════════════════════

  ┌──────────────┐     ┌───────────────┐     ┌──────────────┐
  │  PERCEPTION  │────►│   REASONING   │────►│    ACTION     │
  │              │     │               │     │               │
  │ User query   │     │ "Which tool?" │     │ Call tool     │
  │ Tool output  │     │ "Enough info?"│     │ Generate text │
  │ Error msg    │     │ "Need retry?" │     │ Return answer │
  └──────┬───────┘     └───────────────┘     └───────┬───────┘
         ▲                                           │
         │           ┌───────────────┐               │
         └───────────│  OBSERVATION  │◄──────────────┘
                     │               │
                     │ Tool result   │
                     │ Error / OK    │
                     └───────────────┘
         Loop tiếp tục cho đến khi có Final Answer

1.3. Levels of Agency

Không phải mọi ứng dụng LLM đều cần full agent. NVIDIA DLI phân biệt rõ các mức độ agency:

LevelPatternLLM RoleExample
L0 — No agencySimple prompt → responseText generatorChatbot trả lời FAQ
L1 — Tool useLLM chọn 1 toolRouterFunction calling API
L2 — Single agentReAct loop, multi-stepPlanner + executorRAG Agent (Bài 8)
L3 — Multi-agentNhiều agents phối hợpCoordinatorSupervisor + workers
L4 — AutonomousSelf-improving, long-runningAutonomous systemAI Scientist, Devin

Exam tip: "LLM tự chia nhỏ task, gọi nhiều tools, lặp lại khi cần" → Agent (L2+). "Nhiều LLM phối hợp, mỗi cái chuyên một việc" → Multi-Agent (L3). DLI exam thường hỏi: "What differentiates an agent from a chain?" → Agent có dynamic control flow (LLM quyết định bước tiếp theo), chain có fixed control flow.

Multi-Agent System — Orchestrator, Specialized Agents, LangGraph State Machine
Multi-Agent System — Orchestrator, Specialized Agents, LangGraph State Machine

2. Cognitive Architectures cho LLM Agents

2.1. ReAct — Reasoning + Acting

ReAct (đã giới thiệu Bài 8) xen kẽ Thought (suy nghĩ) với Action (hành động). Ưu điểm: đơn giản, transparent. Nhược điểm: không có planning dài hạn — agent chỉ nghĩ bước tiếp theo, không nhìn toàn cảnh.

2.2. Plan-and-Execute

Plan-and-Execute tách rõ hai giai đoạn: (1) Planner LLM lên kế hoạch toàn bộ trước, (2) Executor LLM thực hiện từng bước. Sau mỗi bước, Planner có thể replan (điều chỉnh kế hoạch).


Plan-and-Execute Architecture
══════════════════════════════════════════════════════════

  User: "Phân tích doanh thu Q3, so sánh với Q2, viết báo cáo"
       │
       ▼
  ┌─────────────────────────────────────────────┐
  │  PLANNER LLM                                │
  │  Plan:                                      │
  │    Step 1: Retrieve Q3 revenue data         │
  │    Step 2: Retrieve Q2 revenue data         │
  │    Step 3: Calculate Q2→Q3 change           │
  │    Step 4: Write comparison report          │
  └─────────────────────┬───────────────────────┘
                        │
          ┌─────────────┼─────────────┐
          ▼             ▼             ▼
     Execute S1    Execute S2    Execute S3 ...
     (retriever)   (retriever)   (calculator)
          │             │             │
          └─────────────┼─────────────┘
                        │
                        ▼
                ┌───────────────┐
                │ REPLAN?       │──► Nếu step fail → adjust plan
                │ All done?     │──► Nếu done → Step 4: report
                └───────────────┘

2.3. LATS — Language Agent Tree Search

LATS kết hợp Monte Carlo Tree Search (MCTS) với LLM reasoning. Thay vì đi theo 1 path (ReAct), LATS explore nhiều nhánh giải pháp, đánh giá từng nhánh bằng LLM, rồi chọn nhánh tốt nhất. Giống như LLM chơi cờ — suy nghĩ trước nhiều bước.

2.4. Reflexion — Learn from Mistakes

Reflexion thêm bước self-reflection: sau khi hoàn thành task, agent tự đánh giá kết quả → nếu sai, viết "bài học" vào memory → thử lại với kinh nghiệm từ lần trước. Đây là dạng in-context learning qua self-feedback.

2.5. So sánh Cognitive Architectures

ArchitecturePlanningExecutionStrengthWeakness
ReActStep-by-step (myopic)Interleaved think+actSimple, transparentNo global planning, can loop
Plan-and-ExecuteUpfront full planSequential executionGlobal view, fewer LLM callsPlan có thể outdated after steps
LATSTree search (explore)Best-first searchExplores alternatives, robustVery expensive (nhiều LLM calls)
ReflexionTrial-and-error + memoryExecute → reflect → retryLearns from mistakesSlow convergence, needs evaluator

Exam tip: "Agent cần suy nghĩ toàn bộ kế hoạch trước khi thực hiện" → Plan-and-Execute. "Agent thử nhiều hướng giải quyết, chọn tốt nhất" → LATS. "Agent tự đánh giá kết quả và cải thiện" → Reflexion. "Agent xen kẽ suy nghĩ và hành động" → ReAct. Đề DLI thường ưu tiên hỏi ReAct và Plan-and-Execute vì hai kiến trúc này phổ biến nhất.

3. LangGraph — Stateful Graph-Based Agent Orchestration

3.1. Tại sao cần LangGraph?

AgentExecutor trong LangChain (Bài 8) là "black box" — khó customize control flow. LangGraph là thư viện của LangChain cho phép xây dựng agent dạng directed graph: mỗi node là một bước xử lý, edges định nghĩa flow, conditional edges cho phép rẽ nhánh dựa trên state.

FeatureAgentExecutorLangGraph
Control flowFixed ReAct loopCustom graph — bạn thiết kế flow
State managementHidden internal stateExplicit TypedDict state
Multi-agentKhông hỗ trợ nativeFirst-class: mỗi agent = sub-graph
Human-in-the-loopHạn chếBuilt-in: interrupt, approve, edit
PersistenceNo built-inCheckpointer: save/resume state
StreamingBasicEvent-by-event streaming
DebugTrace via LangSmithGraph visualization + LangSmith

3.2. Core Concepts — StateGraph, Nodes, Edges

LangGraph xây dựng trên 3 khái niệm:

  • State — TypedDict giữ toàn bộ data chuyền giữa các nodes. Mỗi node đọc/ghi state.
  • Nodes — Python functions. Input: state → Output: partial state update (chỉ fields cần update).
  • Edges — kết nối giữa nodes. add_edge(A, B) = always go A→B. add_conditional_edges(A, func) = func quyết định đi đâu.

LangGraph Concepts
══════════════════════════════════════════════════════════

  State = TypedDict(messages, plan, results, ...)
  ────────────────────────────────────────────────

  START ──► [Node: agent]  ──conditional──► [Node: tools]
                 │                              │
                 │  (if done)                   │ (tool result)
                 ▼                              │
               END ◄───────────────────────────┘

  Nodes: Python functions that read/write State
  Edges: Static (always) or Conditional (function decides)

3.3. Code: Basic LangGraph Agent


from typing import TypedDict, Annotated, Sequence
from langchain_core.messages import BaseMessage, HumanMessage, AIMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
import operator

# === 1. Define State ===
class AgentState(TypedDict):
    messages: Annotated[Sequence[BaseMessage], operator.add]

# === 2. Define Tools ===
@tool
def search_docs(query: str) -> str:
    """Search internal documents for company information."""
    # Simulate retrieval
    docs = {
        "leave": "Employees get 12 days annual leave per year.",
        "refund": "Refund within 30 days with original receipt.",
    }
    for key, val in docs.items():
        if key in query.lower():
            return val
    return "No relevant documents found."

@tool
def calculator(expression: str) -> str:
    """Calculate mathematical expressions."""
    try:
        return str(eval(expression))  # production: use safe eval
    except Exception as e:
        return f"Error: {e}"

tools = [search_docs, calculator]

# === 3. Define LLM with tools ===
llm = ChatNVIDIA(
    model="meta/llama-3.1-70b-instruct",
    temperature=0.1
).bind_tools(tools)

# === 4. Define Nodes ===
def agent_node(state: AgentState) -> dict:
    """LLM decides: call tool or respond."""
    response = llm.invoke(state["messages"])
    return {"messages": [response]}

tool_node = ToolNode(tools)

# === 5. Define Routing ===
def should_continue(state: AgentState) -> str:
    last_message = state["messages"][-1]
    if last_message.tool_calls:
        return "tools"    # LLM wants to call a tool
    return "end"          # LLM is done, return answer

# === 6. Build Graph ===
graph = StateGraph(AgentState)

graph.add_node("agent", agent_node)
graph.add_node("tools", tool_node)

graph.set_entry_point("agent")
graph.add_conditional_edges("agent", should_continue, {
    "tools": "tools",
    "end": END,
})
graph.add_edge("tools", "agent")  # after tool → back to agent

app = graph.compile()

# === 7. Run ===
result = app.invoke({
    "messages": [HumanMessage(content="Chính sách nghỉ phép là gì?")]
})
print(result["messages"][-1].content)

3.4. Human-in-the-Loop

LangGraph hỗ trợ interrupt trước khi thực hiện action nguy hiểm — ví dụ: gửi email, xóa dữ liệu, thực thi code. Agent tạm dừng, chờ user approve, rồi tiếp tục.


from langgraph.checkpoint.memory import MemorySaver

# Compile with checkpointer for interruption + resume
checkpointer = MemorySaver()
app = graph.compile(
    checkpointer=checkpointer,
    interrupt_before=["tools"]   # pause BEFORE executing tools
)

# Run — sẽ dừng trước node "tools"
config = {"configurable": {"thread_id": "user-123"}}
result = app.invoke(
    {"messages": [HumanMessage(content="Delete file report.pdf")]},
    config=config,
)

# Inspect pending tool call
pending = result["messages"][-1].tool_calls
print(f"Agent wants to: {pending}")
# → Agent wants to: [{'name': 'delete_file', 'args': {'path': 'report.pdf'}}]

# Human approves → continue
final = app.invoke(None, config=config)  # resume from checkpoint

3.5. Checkpointing — Save & Resume State

Checkpointer lưu state sau mỗi node, cho phép:

  • Resume — agent crash giữa chừng → load checkpoint → chạy tiếp
  • Time travel — quay lại bất kỳ checkpoint nào → thử lại với input khác
  • Human-in-the-loop — pause, đợi user, resume (như code trên)
  • Multi-turn — giữ conversation history qua nhiều turns

Exam tip: "Build agent with custom control flow, conditional branching" → LangGraph (không phải AgentExecutor). "Pause agent execution for human approval" → interrupt_before + checkpointer. "Save agent state, resume later" → LangGraph checkpointing. DLI C-FX-25 thường hỏi: "Why use LangGraph over AgentExecutor?" → Custom flow, multi-agent, persistence, human-in-the-loop.

4. Multi-Agent Patterns

4.1. Tại sao cần Multi-Agent?

Một single agent với 20+ tools sẽ gặp vấn đề: tool selection confusion (quá nhiều tool, LLM chọn sai), prompt quá dài (phải nhét hết instructions), khó debug (không rõ agent fail ở bước nào). Multi-agent giải quyết bằng cách chia nhỏ: mỗi agent chuyên một nhiệm vụ với ít tools hơn.

4.2. Supervisor Pattern

Một Supervisor agent (LLM) nhận task từ user, phân công cho các Worker agents, thu thập kết quả, và tổng hợp câu trả lời.


Supervisor Pattern
══════════════════════════════════════════════════════════

                    ┌──────────────┐
                    │     USER     │
                    └──────┬───────┘
                           │
                           ▼
              ┌────────────────────────┐
              │   SUPERVISOR AGENT     │
              │   (Orchestrator LLM)   │
              │                        │
              │   Decides:             │
              │   • Which worker next? │
              │   • All done?          │
              │   • Need to re-route?  │
              └────┬──────┬──────┬─────┘
                   │      │      │
          ┌────────┘      │      └────────┐
          ▼               ▼               ▼
   ┌─────────────┐┌─────────────┐┌─────────────┐
   │ Researcher  ││   Coder     ││  Reporter   │
   │ Agent       ││   Agent     ││  Agent      │
   │             ││             ││             │
   │ Tools:      ││ Tools:      ││ Tools:      │
   │ • web_search││ • python    ││ • write_doc │
   │ • doc_search││ • shell     ││ • format    │
   └─────────────┘└─────────────┘└─────────────┘

4.3. Hierarchical Pattern

Hierarchical mở rộng Supervisor: mỗi worker có thể là supervisor của sub-workers. Phù hợp cho organization phức tạp — ví dụ: CEO agent → Manager agents → Specialist agents.


Hierarchical Multi-Agent
══════════════════════════════════════════════════════════

              ┌──────────────────────┐
              │   TOP SUPERVISOR     │
              │   (Project Manager)  │
              └───┬─────────────┬────┘
                  │             │
         ┌────────┘             └────────┐
         ▼                               ▼
  ┌──────────────┐                ┌──────────────┐
  │ RESEARCH     │                │ ENGINEERING  │
  │ SUPERVISOR   │                │ SUPERVISOR   │
  └──┬───────┬───┘                └──┬───────┬───┘
     │       │                       │       │
     ▼       ▼                       ▼       ▼
  [Web    [Paper                 [Backend [Frontend
  Searcher] Analyzer]             Dev]     Dev]

4.4. Swarm Pattern

Swarm (OpenAI Swarm concept) — không có supervisor. Agents tự chuyển tiếp (handoff) cho nhau dựa trên context. Agent A nhận ra "task này thuộc chuyên môn Agent B" → tự handoff.

4.5. Debate Pattern

Debate — hai hoặc nhiều agents tranh luận về một câu hỏi. Mỗi agent đưa ra quan điểm, phản bác quan điểm kia. Cuối cùng, một Judge agent chọn kết luận tốt nhất. Pattern này improve reasoning quality cho câu hỏi phức tạp.

4.6. So sánh Multi-Agent Patterns

PatternControl FlowCommunicationBest For
SupervisorCentralized — supervisor routesHub-and-spokeClear task delegation, moderate complexity
HierarchicalMulti-level supervisionTree structureComplex orgs, many specialized sub-teams
SwarmDecentralized — agents handoffPeer-to-peerCustomer service, routing, flexible flow
DebateRound-robin argumentationBroadcast + judgeComplex reasoning, fact verification

Exam tip: "One LLM routes tasks to specialized agents" → Supervisor. "Agents hand off to each other without central control" → Swarm. "Multiple agents argue, a judge decides" → Debate. "Nested supervisors managing sub-teams" → Hierarchical. DLI exam thường tập trung: Supervisor pattern vì là phổ biến nhất trong production.

4.7. Code: Supervisor Multi-Agent với LangGraph


from typing import TypedDict, Annotated, Literal, Sequence
from langchain_core.messages import BaseMessage, HumanMessage, SystemMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
import operator

# === State ===
class MultiAgentState(TypedDict):
    messages: Annotated[Sequence[BaseMessage], operator.add]
    next_agent: str

# === Worker Tools ===
@tool
def web_search(query: str) -> str:
    """Search the web for current information."""
    return f"[Web Result] Top findings for '{query}': ..."

@tool
def run_python(code: str) -> str:
    """Execute Python code and return output."""
    try:
        exec_globals = {}
        exec(code, exec_globals)
        return str(exec_globals.get("result", "Code executed successfully."))
    except Exception as e:
        return f"Error: {e}"

@tool
def write_report(content: str) -> str:
    """Format content into a professional report."""
    return f"=== REPORT ===\n{content}\n=== END ==="

# === Worker Agents ===
researcher_llm = ChatNVIDIA(
    model="meta/llama-3.1-70b-instruct", temperature=0.1
).bind_tools([web_search])

coder_llm = ChatNVIDIA(
    model="meta/llama-3.1-70b-instruct", temperature=0.0
).bind_tools([run_python])

reporter_llm = ChatNVIDIA(
    model="meta/llama-3.1-70b-instruct", temperature=0.3
).bind_tools([write_report])

# === Supervisor ===
supervisor_llm = ChatNVIDIA(
    model="meta/llama-3.1-70b-instruct", temperature=0.0
)

WORKERS = ["researcher", "coder", "reporter"]

def supervisor_node(state: MultiAgentState) -> dict:
    """Supervisor decides which worker to route to next."""
    system_prompt = f"""You are a supervisor managing these workers: {WORKERS}.
Given the conversation, decide which worker should act next,
or if the task is complete respond with FINISH.
Respond with ONLY the worker name or FINISH."""

    messages = [SystemMessage(content=system_prompt)] + state["messages"]
    response = supervisor_llm.invoke(messages)
    next_agent = response.content.strip().lower()

    if next_agent not in WORKERS:
        next_agent = "FINISH"
    return {"next_agent": next_agent}

def researcher_node(state: MultiAgentState) -> dict:
    system = SystemMessage(content="You are a research specialist. "
        "Use web_search to find information. Be thorough.")
    response = researcher_llm.invoke([system] + state["messages"])
    return {"messages": [response]}

def coder_node(state: MultiAgentState) -> dict:
    system = SystemMessage(content="You are a Python coding specialist. "
        "Use run_python to execute code for analysis and calculations.")
    response = coder_llm.invoke([system] + state["messages"])
    return {"messages": [response]}

def reporter_node(state: MultiAgentState) -> dict:
    system = SystemMessage(content="You are a report writer. "
        "Use write_report to create formatted reports from gathered info.")
    response = reporter_llm.invoke([system] + state["messages"])
    return {"messages": [response]}

# === Routing ===
def route_supervisor(state: MultiAgentState) -> str:
    next_agent = state.get("next_agent", "FINISH")
    if next_agent == "FINISH":
        return "end"
    return next_agent

# === Build Graph ===
graph = StateGraph(MultiAgentState)

graph.add_node("supervisor", supervisor_node)
graph.add_node("researcher", researcher_node)
graph.add_node("coder", coder_node)
graph.add_node("reporter", reporter_node)
graph.add_node("researcher_tools", ToolNode([web_search]))
graph.add_node("coder_tools", ToolNode([run_python]))
graph.add_node("reporter_tools", ToolNode([write_report]))

graph.set_entry_point("supervisor")

# Supervisor routes to workers
graph.add_conditional_edges("supervisor", route_supervisor, {
    "researcher": "researcher",
    "coder": "coder",
    "reporter": "reporter",
    "end": END,
})

# Workers → tool nodes → back to supervisor
for worker in WORKERS:
    def make_router(w):
        def router(state):
            last = state["messages"][-1]
            if hasattr(last, "tool_calls") and last.tool_calls:
                return f"{w}_tools"
            return "supervisor"
        return router
    graph.add_conditional_edges(worker, make_router(worker), {
        f"{worker}_tools": f"{worker}_tools",
        "supervisor": "supervisor",
    })
    graph.add_edge(f"{worker}_tools", worker)

app = graph.compile()

# === Run ===
result = app.invoke({
    "messages": [HumanMessage(
        content="Research NVIDIA H100 GPU specs, calculate price-performance "
                "ratio vs A100, and write a comparison report."
    )],
    "next_agent": "",
})

for msg in result["messages"]:
    print(f"[{msg.type}] {msg.content[:200]}...")

5. Build Production-Ready Multi-Agent App

5.1. Research Assistant — Complete Example

Xây dựng Research Assistant hoàn chỉnh với 3 agents: Researcher (tìm thông tin), Coder (phân tích data), Reporter (viết báo cáo). Có error handling, retry logic, và structured output.


from typing import TypedDict, Annotated, Sequence, Optional
from langchain_core.messages import BaseMessage, HumanMessage, SystemMessage, AIMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.checkpoint.memory import MemorySaver
import operator
import json

# === Enhanced State ===
class ResearchState(TypedDict):
    messages: Annotated[Sequence[BaseMessage], operator.add]
    research_data: Optional[str]       # collected research
    analysis_result: Optional[str]     # code analysis output
    final_report: Optional[str]        # formatted report
    current_agent: str
    iteration: int                     # track iterations to prevent loops

MAX_ITERATIONS = 10

# === Tools ===
@tool
def search_arxiv(query: str) -> str:
    """Search academic papers on arxiv for research topics."""
    return json.dumps({
        "papers": [
            {"title": f"Paper on {query}", "abstract": f"Study of {query}...",
             "year": 2025, "citations": 42},
        ]
    })

@tool
def search_web(query: str) -> str:
    """Search the web for current news, blog posts, documentation."""
    return json.dumps({
        "results": [
            {"title": f"Latest news: {query}", "snippet": f"Updated info on {query}..."},
        ]
    })

@tool
def execute_analysis(code: str) -> str:
    """Run Python code for data analysis. Variable 'result' will be returned."""
    exec_globals = {}
    try:
        exec(code, exec_globals)
        return str(exec_globals.get("result", "Executed OK, no 'result' variable."))
    except Exception as e:
        return f"Error: {e}"

@tool
def generate_report(title: str, sections: str) -> str:
    """Generate a formatted markdown report from title and section content."""
    return f"# {title}\n\n{sections}\n\n---\nGenerated by Research Assistant"

# === Agent Nodes ===
base_llm = ChatNVIDIA(model="meta/llama-3.1-70b-instruct")

def researcher_node(state: ResearchState) -> dict:
    llm = base_llm.bind_tools([search_arxiv, search_web])
    system = SystemMessage(content=(
        "You are a research specialist. Search for papers and web results "
        "to gather comprehensive information. Summarize findings clearly."
    ))
    response = llm.invoke([system] + list(state["messages"]))

    # If no tool calls, research is done — extract data
    if not response.tool_calls:
        return {
            "messages": [response],
            "research_data": response.content,
            "current_agent": "supervisor",
        }
    return {"messages": [response], "current_agent": "researcher_tools"}

def coder_node(state: ResearchState) -> dict:
    llm = base_llm.bind_tools([execute_analysis])
    context = state.get("research_data", "No research data yet.")
    system = SystemMessage(content=(
        f"You are a data analyst. Use the research data below to perform "
        f"analysis with Python code.\n\nResearch Data:\n{context}"
    ))
    response = llm.invoke([system] + list(state["messages"]))

    if not response.tool_calls:
        return {
            "messages": [response],
            "analysis_result": response.content,
            "current_agent": "supervisor",
        }
    return {"messages": [response], "current_agent": "coder_tools"}

def reporter_node(state: ResearchState) -> dict:
    llm = base_llm.bind_tools([generate_report])
    research = state.get("research_data", "N/A")
    analysis = state.get("analysis_result", "N/A")
    system = SystemMessage(content=(
        f"You are a report writer. Create a professional report.\n"
        f"Research:\n{research}\n\nAnalysis:\n{analysis}"
    ))
    response = llm.invoke([system] + list(state["messages"]))

    if not response.tool_calls:
        return {
            "messages": [response],
            "final_report": response.content,
            "current_agent": "supervisor",
        }
    return {"messages": [response], "current_agent": "reporter_tools"}

def supervisor_node(state: ResearchState) -> dict:
    iteration = state.get("iteration", 0) + 1
    if iteration > MAX_ITERATIONS:
        return {
            "messages": [AIMessage(content="Max iterations reached. Returning results.")],
            "current_agent": "FINISH",
            "iteration": iteration,
        }

    system = SystemMessage(content="""You are a project supervisor. Based on the current state:
- If no research data → route to "researcher"
- If research done but no analysis → route to "coder"
- If analysis done but no report → route to "reporter"
- If report is ready → respond "FINISH"
Respond with ONLY one of: researcher, coder, reporter, FINISH""")

    response = base_llm.invoke([system] + list(state["messages"]))
    next_agent = response.content.strip().lower()

    valid = ["researcher", "coder", "reporter", "finish"]
    if next_agent not in valid:
        next_agent = "researcher"  # default fallback

    return {"current_agent": next_agent, "iteration": iteration}

# === Routing ===
def route_from_supervisor(state: ResearchState) -> str:
    agent = state.get("current_agent", "FINISH")
    if agent in ["researcher", "coder", "reporter"]:
        return agent
    return "end"

def route_from_worker(worker_name: str):
    def router(state: ResearchState) -> str:
        current = state.get("current_agent", "supervisor")
        if current == f"{worker_name}_tools":
            return f"{worker_name}_tools"
        return "supervisor"
    return router

# === Build Graph ===
from langgraph.prebuilt import ToolNode

graph = StateGraph(ResearchState)

graph.add_node("supervisor", supervisor_node)
graph.add_node("researcher", researcher_node)
graph.add_node("coder", coder_node)
graph.add_node("reporter", reporter_node)
graph.add_node("researcher_tools", ToolNode([search_arxiv, search_web]))
graph.add_node("coder_tools", ToolNode([execute_analysis]))
graph.add_node("reporter_tools", ToolNode([generate_report]))

graph.set_entry_point("supervisor")

graph.add_conditional_edges("supervisor", route_from_supervisor, {
    "researcher": "researcher",
    "coder": "coder",
    "reporter": "reporter",
    "end": END,
})

for worker in ["researcher", "coder", "reporter"]:
    graph.add_conditional_edges(worker, route_from_worker(worker), {
        f"{worker}_tools": f"{worker}_tools",
        "supervisor": "supervisor",
    })
    graph.add_edge(f"{worker}_tools", worker)

# Compile with checkpointing
checkpointer = MemorySaver()
app = graph.compile(checkpointer=checkpointer)

# === Execute ===
config = {"configurable": {"thread_id": "research-001"}}
result = app.invoke(
    {
        "messages": [HumanMessage(
            content="Research the latest advances in mixture-of-experts (MoE) "
                    "models, analyze their parameter efficiency compared to "
                    "dense models, and write a summary report."
        )],
        "current_agent": "",
        "iteration": 0,
    },
    config=config,
)

# Print final report
print(result.get("final_report", result["messages"][-1].content))

5.2. Production Best Practices

PracticeWhyImplementation
Max iterationsPrevent infinite loopsiteration counter in state, check at supervisor
Error handlingTool failures shouldn't crash agenttry/except in tools, return error message
CheckpointingResume after crashMemorySaver (dev) / SqliteSaver (prod)
Structured outputReliable routing decisionsConstrain supervisor output to valid choices
ObservabilityDebug multi-agent is hardLangSmith tracing, log each node entry/exit
Timeout per nodeSingle node shouldn't blockSet timeout on LLM calls and tool executions
Human-in-the-loopCritical actions need approvalinterrupt_before on dangerous tool nodes

Exam tip: "How to prevent agent infinite loops?" → max_iterations + iteration counter. "How to debug multi-agent systems?" → LangSmith tracing + logging. "Agent crash recovery?" → Checkpointing with persistent storage. Production deployment → LangGraph Platform (managed) hoặc LangServe (self-hosted).

6. DLI C-FX-25 — Agentic AI Course Overview

6.1. Course Structure

Course C-FX-25: "Building Agentic AI Applications" là module nâng cao trong DLI, tập trung vào xây dựng hệ thống Agentic AI production-ready. Course bổ sung cho S-FX-15 bằng cách đi sâu vào agent architectures.

ModuleTopicsHands-On
Module 1Agent fundamentals, ReAct, tool callingBuild single agent with NVIDIA NIM
Module 2LangGraph introduction, StateGraphImplement custom agent graph
Module 3Multi-agent architecturesBuild supervisor multi-agent system
Module 4Advanced: memory, planning, evalProduction deployment exercise

6.2. Assessment Focus Areas

C-FX-25 assessment tập trung vào hands-on implementation:

  • LangGraph StateGraph — define state, nodes, conditional edges
  • Tool integration — bind tools to LLM, handle tool calls
  • Supervisor routing — implement supervisor logic, route to workers
  • Checkpointing — save/restore agent state
  • Human-in-the-loop — interrupt_before, approve, resume

6.3. Key APIs to Memorize

API / ConceptUsage
StateGraph(State)Create graph with typed state
graph.add_node(name, func)Add processing node
graph.add_edge(A, B)Always route A → B
graph.add_conditional_edges(A, func, map)Route based on function output
graph.set_entry_point(name)Set starting node
graph.compile(checkpointer=...)Compile graph, optional checkpointer
ToolNode(tools)Pre-built node that executes tool calls
MemorySaver()In-memory checkpointer (dev only)
interrupt_before=[node]Pause before executing node
llm.bind_tools(tools)Attach tools to LLM for function calling

Exam tip: C-FX-25 assessment yêu cầu viết code LangGraph từ đầu. Nhớ rõ pattern: (1) Define State TypedDict, (2) Define nodes as functions, (3) Add nodes + edges, (4) Compile + run. Không cần thuộc lòng API nhưng phải hiểu flow: state travels through nodes, conditional edges route dynamically.

7. Cheat Sheet

ConceptKey Point
Agent componentsLLM + Memory + Tools + Planning
Agent loopPerception → Reasoning → Action → Observation
Agent vs ChainAgent = dynamic flow (LLM decides); Chain = fixed flow
ReActInterleave Thought + Action + Observation. Simple, myopic
Plan-and-ExecutePlan upfront → execute steps → replan if needed
LATSTree search over reasoning paths. Expensive but robust
ReflexionExecute → self-reflect → retry with lessons
LangGraphStateGraph: nodes + edges + conditional routing
LangGraph StateTypedDict shared across all nodes
Conditional edgesRouter function decides next node
CheckpointingMemorySaver (dev), SqliteSaver (prod). Enable resume
Human-in-the-loopinterrupt_before=[node] + compile with checkpointer
Supervisor patternCentral LLM routes to specialized worker agents
HierarchicalNested supervisors — tree of agents
SwarmDecentralized handoffs, no central supervisor
DebateAgents argue, judge decides. Better reasoning
Max iterationsAlways set to prevent infinite agent loops
ToolNodeLangGraph pre-built node to execute tool calls
C-FX-25 focusLangGraph coding, multi-agent, checkpointing, HITL

8. Practice Questions — Coding

Q1: Build a basic LangGraph ReAct agent

Xây dựng một LangGraph agent đơn giản với 2 tools: search_docs (tìm tài liệu) và calculator (tính toán). Implement đầy đủ: State, agent node, tool node, conditional edge routing, compile và run.

Xem đáp án Q1

from typing import TypedDict, Annotated, Sequence
from langchain_core.messages import BaseMessage, HumanMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
import operator

# 1. State
class AgentState(TypedDict):
    messages: Annotated[Sequence[BaseMessage], operator.add]

# 2. Tools
@tool
def search_docs(query: str) -> str:
    """Search internal knowledge base for relevant documents."""
    return f"Found: Documentation about {query} — key facts here."

@tool
def calculator(expression: str) -> str:
    """Calculate a mathematical expression."""
    return str(eval(expression))

tools = [search_docs, calculator]

# 3. LLM with tools
llm = ChatNVIDIA(model="meta/llama-3.1-70b-instruct", temperature=0.0)
llm_with_tools = llm.bind_tools(tools)

# 4. Nodes
def agent_node(state: AgentState) -> dict:
    response = llm_with_tools.invoke(state["messages"])
    return {"messages": [response]}

tool_node = ToolNode(tools)

# 5. Router
def should_continue(state: AgentState) -> str:
    last = state["messages"][-1]
    if hasattr(last, "tool_calls") and last.tool_calls:
        return "tools"
    return "end"

# 6. Build graph
graph = StateGraph(AgentState)
graph.add_node("agent", agent_node)
graph.add_node("tools", tool_node)
graph.set_entry_point("agent")
graph.add_conditional_edges("agent", should_continue, {
    "tools": "tools",
    "end": END,
})
graph.add_edge("tools", "agent")

app = graph.compile()

# 7. Run
result = app.invoke({
    "messages": [HumanMessage(content="What is 25 * 4 + 100?")]
})
print(result["messages"][-1].content)

Q2: Implement human-in-the-loop with LangGraph checkpointing

Modify agent từ Q1 để thêm human-in-the-loop: agent pause trước khi thực thi tools, user có thể approve hoặc reject. Demonstrate: (1) compile với checkpointer + interrupt_before, (2) run và thấy agent pause, (3) resume execution.

Xem đáp án Q2

from langgraph.checkpoint.memory import MemorySaver

# Reuse graph from Q1, compile with HITL
checkpointer = MemorySaver()
app_hitl = graph.compile(
    checkpointer=checkpointer,
    interrupt_before=["tools"]  # Pause BEFORE tool execution
)

# Run — agent will pause before calling tools
config = {"configurable": {"thread_id": "hitl-demo-001"}}
result = app_hitl.invoke(
    {"messages": [HumanMessage(content="Calculate 1000 / 4")]},
    config=config,
)

# Agent paused — inspect what it wants to do
last_msg = result["messages"][-1]
print("Agent wants to call:")
for tc in last_msg.tool_calls:
    print(f"  Tool: {tc['name']}, Args: {tc['args']}")

# User approves → resume (pass None to continue from checkpoint)
final_result = app_hitl.invoke(None, config=config)
print("\nFinal answer:", final_result["messages"][-1].content)

# If user REJECTS → could modify state or stop here
# To reject: simply don't call invoke(None, config)

Q3: Build a Supervisor multi-agent system

Implement Supervisor pattern với 2 workers: researcher (sử dụng web_search tool) và writer (sử dụng write_report tool). Supervisor nhận task từ user, route đến worker phù hợp, thu thập kết quả. Implement routing logic với conditional edges.

Xem đáp án Q3

from typing import TypedDict, Annotated, Sequence
from langchain_core.messages import BaseMessage, HumanMessage, SystemMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
import operator

class SupervisorState(TypedDict):
    messages: Annotated[Sequence[BaseMessage], operator.add]
    next: str

@tool
def web_search(query: str) -> str:
    """Search the internet for information."""
    return f"Search results for '{query}': ..."

@tool
def write_report(content: str) -> str:
    """Write and format a professional report."""
    return f"=== Report ===\n{content}\n=== End ==="

llm = ChatNVIDIA(model="meta/llama-3.1-70b-instruct", temperature=0.0)

def supervisor(state: SupervisorState) -> dict:
    sys = SystemMessage(content=(
        "You are a supervisor. Workers: researcher, writer. "
        "Route to appropriate worker or say FINISH if task complete. "
        "Respond with ONLY: researcher, writer, or FINISH."
    ))
    resp = llm.invoke([sys] + list(state["messages"]))
    next_val = resp.content.strip().lower()
    if next_val not in ["researcher", "writer"]:
        next_val = "FINISH"
    return {"next": next_val}

def researcher(state: SupervisorState) -> dict:
    r_llm = llm.bind_tools([web_search])
    sys = SystemMessage(content="You are a researcher. Use web_search.")
    resp = r_llm.invoke([sys] + list(state["messages"]))
    return {"messages": [resp]}

def writer(state: SupervisorState) -> dict:
    w_llm = llm.bind_tools([write_report])
    sys = SystemMessage(content="You are a report writer. Use write_report.")
    resp = w_llm.invoke([sys] + list(state["messages"]))
    return {"messages": [resp]}

def route(state: SupervisorState) -> str:
    n = state.get("next", "FINISH")
    return n if n in ["researcher", "writer"] else "end"

# Build
g = StateGraph(SupervisorState)
g.add_node("supervisor", supervisor)
g.add_node("researcher", researcher)
g.add_node("writer", writer)
g.add_node("research_tools", ToolNode([web_search]))
g.add_node("writer_tools", ToolNode([write_report]))

g.set_entry_point("supervisor")
g.add_conditional_edges("supervisor", route, {
    "researcher": "researcher",
    "writer": "writer",
    "end": END,
})

# Researcher flow
def route_researcher(state):
    last = state["messages"][-1]
    if hasattr(last, "tool_calls") and last.tool_calls:
        return "research_tools"
    return "supervisor"

g.add_conditional_edges("researcher", route_researcher, {
    "research_tools": "research_tools",
    "supervisor": "supervisor",
})
g.add_edge("research_tools", "researcher")

# Writer flow
def route_writer(state):
    last = state["messages"][-1]
    if hasattr(last, "tool_calls") and last.tool_calls:
        return "writer_tools"
    return "supervisor"

g.add_conditional_edges("writer", route_writer, {
    "writer_tools": "writer_tools",
    "supervisor": "supervisor",
})
g.add_edge("writer_tools", "writer")

app = g.compile()

result = app.invoke({
    "messages": [HumanMessage(content="Research AI trends 2025 and write a report")],
    "next": "",
})
print(result["messages"][-1].content)

Q4: Add Plan-and-Execute to a LangGraph agent

Implement Plan-and-Execute pattern: (1) Planner node tạo danh sách steps từ user query, (2) Executor node thực hiện từng step, (3) Replanner node kiểm tra progress và điều chỉnh plan nếu cần. Lưu plan trong state.

Xem đáp án Q4

from typing import TypedDict, Annotated, Sequence, List, Optional
from langchain_core.messages import BaseMessage, HumanMessage, SystemMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langgraph.graph import StateGraph, END
import operator, json

class PlanExecState(TypedDict):
    messages: Annotated[Sequence[BaseMessage], operator.add]
    plan: List[str]          # list of steps
    current_step: int        # index of current step
    step_results: List[str]  # results of each step
    done: bool

llm = ChatNVIDIA(model="meta/llama-3.1-70b-instruct", temperature=0.0)

def planner_node(state: PlanExecState) -> dict:
    """Create a plan from user request."""
    sys = SystemMessage(content=(
        "You are a planner. Break the user's request into 3-5 concrete steps. "
        "Return ONLY a JSON array of strings, e.g. [\"step1\", \"step2\"]."
    ))
    resp = llm.invoke([sys] + list(state["messages"]))
    try:
        plan = json.loads(resp.content)
    except json.JSONDecodeError:
        plan = [resp.content]
    return {"plan": plan, "current_step": 0, "step_results": []}

def executor_node(state: PlanExecState) -> dict:
    """Execute the current step of the plan."""
    step_idx = state["current_step"]
    plan = state["plan"]
    if step_idx >= len(plan):
        return {"done": True}

    current = plan[step_idx]
    sys = SystemMessage(content=(
        f"Execute this step: {current}\n"
        f"Previous results: {state['step_results']}\n"
        "Provide a concise result."
    ))
    resp = llm.invoke([sys] + list(state["messages"]))
    new_results = list(state["step_results"]) + [resp.content]
    return {
        "step_results": new_results,
        "current_step": step_idx + 1,
        "messages": [resp],
    }

def replanner_node(state: PlanExecState) -> dict:
    """Check progress, adjust plan if needed."""
    if state["current_step"] >= len(state["plan"]):
        return {"done": True}

    sys = SystemMessage(content=(
        f"Plan: {state['plan']}\n"
        f"Completed: {state['current_step']}/{len(state['plan'])}\n"
        f"Results so far: {state['step_results']}\n"
        "Should the remaining plan continue as-is? "
        "Reply 'CONTINUE' or provide updated remaining steps as JSON array."
    ))
    resp = llm.invoke([sys])
    if "CONTINUE" in resp.content.upper():
        return {"done": False}
    try:
        remaining = json.loads(resp.content)
        new_plan = state["plan"][:state["current_step"]] + remaining
        return {"plan": new_plan, "done": False}
    except json.JSONDecodeError:
        return {"done": False}

def route_after_exec(state: PlanExecState) -> str:
    if state.get("done", False):
        return "end"
    return "replanner"

def route_after_replan(state: PlanExecState) -> str:
    if state.get("done", False):
        return "end"
    return "executor"

# Build graph
g = StateGraph(PlanExecState)
g.add_node("planner", planner_node)
g.add_node("executor", executor_node)
g.add_node("replanner", replanner_node)

g.set_entry_point("planner")
g.add_edge("planner", "executor")
g.add_conditional_edges("executor", route_after_exec, {
    "replanner": "replanner",
    "end": END,
})
g.add_conditional_edges("replanner", route_after_replan, {
    "executor": "executor",
    "end": END,
})

app = g.compile()

result = app.invoke({
    "messages": [HumanMessage(
        content="Analyze the pros and cons of microservices architecture "
                "and recommend when to use it vs monolith."
    )],
    "plan": [],
    "current_step": 0,
    "step_results": [],
    "done": False,
})

for i, res in enumerate(result["step_results"]):
    print(f"Step {i+1}: {res[:150]}...")

Q5: Implement error handling and retry logic in multi-agent system

Thêm error handling vào multi-agent system: (1) Tool failures return error message thay vì crash, (2) Agent nhận error → retry với strategy khác (max 2 retries), (3) Agent state track số lần retry. Implement node wrapper pattern.

Xem đáp án Q5

from typing import TypedDict, Annotated, Sequence, Dict
from langchain_core.messages import BaseMessage, HumanMessage, AIMessage
from langchain_nvidia_ai_endpoints import ChatNVIDIA
from langchain_core.tools import tool
from langgraph.graph import StateGraph, END
from langgraph.prebuilt import ToolNode
import operator

class RobustState(TypedDict):
    messages: Annotated[Sequence[BaseMessage], operator.add]
    error_count: int
    max_retries: int

# Tools with error handling built-in
@tool
def risky_api_call(endpoint: str) -> str:
    """Call an external API that might fail."""
    import random
    if random.random() < 0.5:
        raise ConnectionError(f"API {endpoint} unreachable")
    return f"API response from {endpoint}: success data"

@tool
def safe_search(query: str) -> str:
    """Search with built-in error handling."""
    return f"Results for {query}: ..."

# Wrap tools with error handling
def safe_tool_node(tools):
    """ToolNode wrapper that catches errors and returns error messages."""
    base_node = ToolNode(tools)
    def wrapper(state: RobustState) -> dict:
        try:
            return base_node.invoke(state)
        except Exception as e:
            error_msg = AIMessage(content=f"Tool error: {str(e)}. Try different approach.")
            return {
                "messages": [error_msg],
                "error_count": state.get("error_count", 0) + 1,
            }
    return wrapper

tools = [risky_api_call, safe_search]
llm = ChatNVIDIA(model="meta/llama-3.1-70b-instruct", temperature=0.0)
llm_with_tools = llm.bind_tools(tools)

def agent_node(state: RobustState) -> dict:
    error_count = state.get("error_count", 0)
    max_retries = state.get("max_retries", 2)

    # If too many errors, give up gracefully
    if error_count >= max_retries:
        return {"messages": [AIMessage(
            content="I encountered multiple errors. Here's what I could gather "
                    "from successful attempts: " +
                    " | ".join(m.content for m in state["messages"][-3:])
        )]}

    # Add retry context if there were errors
    msgs = list(state["messages"])
    if error_count > 0:
        msgs.append(HumanMessage(
            content=f"Previous attempt failed ({error_count}/{max_retries} retries). "
                    "Try a different tool or approach."
        ))

    response = llm_with_tools.invoke(msgs)
    return {"messages": [response]}

def should_continue(state: RobustState) -> str:
    last = state["messages"][-1]
    error_count = state.get("error_count", 0)
    max_retries = state.get("max_retries", 2)

    # Stop if max retries exceeded
    if error_count >= max_retries and not (
        hasattr(last, "tool_calls") and last.tool_calls
    ):
        return "end"

    if hasattr(last, "tool_calls") and last.tool_calls:
        return "tools"
    return "end"

# Build
g = StateGraph(RobustState)
g.add_node("agent", agent_node)
g.add_node("tools", safe_tool_node(tools))
g.set_entry_point("agent")
g.add_conditional_edges("agent", should_continue, {
    "tools": "tools",
    "end": END,
})
g.add_edge("tools", "agent")  # tool result → back to agent

app = g.compile()

result = app.invoke({
    "messages": [HumanMessage(content="Call the user-data API endpoint")],
    "error_count": 0,
    "max_retries": 2,
})
print(result["messages"][-1].content)