Chuyển đến nội dung chính

第6課:LLM推論パイプライン設計

LLM推論パラメータ:temperature、top-k、top-p。 NVIDIA NIMマイクロサービスによるモデルデプロイ。 LangChain LCELパイプライン。 GradioとLangServe:UI + APIの構築。 対話管理とマルチターン会話。

1. Diffusion ModelsからLLMアプリケーションへ

パート2では、Diffusion Modelsを習得しました — forward/reverseプロセスからCLIPガイド生成まで。パート3では、大規模言語モデル(LLM)と実世界のアプリケーション構築に焦点を移します:推論パイプライン、RAG、チャットボット。

本レッスンでは、LLM推論パイプライン設計に焦点を当てます — サンプリングパラメータによるLLM出力の制御方法、NVIDIA NIMによるモデルデプロイ、LangChain LCELによるパイプライン構築、そしてGradio + LangServeによるUI/APIの作成です。

試験のヒント: NVIDIA DLI試験では、推論パラメータ(temperature、top-k、top-p)やNIMと他のフレームワークの使い分けについて頻繁に出題されます。本レッスン末尾の比較表を必ず押さえてください。

LLM Inference Pipeline — Prompt Template, NIM, LCEL Chain, Gradio UI
LLM推論パイプライン — Prompt Template、NIM、LCEL Chain、Gradio UI

2. LLM推論の基礎

2.1. 自己回帰生成

LLMは自己回帰メカニズムでテキストを生成します:各ステップで、モデルはそれまでのすべてのトークンに基づいて次のトークンを予測します。このプロセスは停止トークンが出現するか、max_tokensに達するまで繰り返されます。


自己回帰生成の流れ
═══════════════════════════════

Input: "Hanoi is"
         │
         ▼
┌─────────────────────┐
│   LLM Forward Pass   │
│   P(token | context) │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│   Sampling Strategy  │──► temperature, top-k, top-p
│   次のトークンを選択  │
└──────────┬──────────┘
           │
           ▼
    token = "the"
           │
           ▼
Input: "Hanoi is the"
         │
         ▼
┌─────────────────────┐
│   LLM Forward Pass   │
└──────────┬──────────┘
           │
           ▼
    token = "capital"
           │
           ▼
   ... <EOS> または max_tokens まで繰り返す

2.2. サンプリングパラメータ

出力の創造性を制御する最も重要な3つのパラメータ:

パラメータ範囲効果低い値高い値
temperature0.0 – 2.0確率分布のエントロピーを調整決定論的、繰り返しが多い創造的、よりランダム
top_k1 – vocab_size確率上位K個のトークンのみに制限より選択的、多様性が低い選択肢が多い
top_p0.0 – 1.0Nucleusサンプリング:累積確率≤pのトークンのみ考慮最も確実なトークンのみより多くのトークンを考慮

トークンサンプリングプロセス(temperature + top-p)
═════════════════════════════════════════════

Raw logits:  [2.1, 1.8, 0.5, 0.3, -1.0, -2.5, ...]
                │
                ▼
         ┌──────────────┐
         │  ÷ temperature │  (temp=0.7 → より鋭い分布)
         └──────┬───────┘
                │
                ▼
Scaled probs: [0.35, 0.28, 0.12, 0.09, 0.08, 0.05, 0.03]
                │
                ▼
         ┌──────────────┐
         │   top-p=0.8   │  cumsum: 0.35→0.63→0.75→0.84 ✓
         │   上位4つを保持 │  → トークン5,6,7...を除外
         └──────┬───────┘
                │
                ▼
Filtered:   [0.41, 0.33, 0.14, 0.12]  (再正規化)
                │
                ▼
         ランダムサンプリング → token "the"

2.3. その他のパラメータ

パラメータ説明ユースケース
max_tokens出力トークンの最大数を制限コスト・レイテンシの制御
stopこの文字列が出現したら生成を停止構造化出力、function calling
repetition_penalty既出トークンにペナルティを付与(>1.0 = 強いペナルティ)単語・文の繰り返しを回避
frequency_penalty出現頻度に基づいて確率を低下より多様な出力
presence_penalty1回でも出現したトークンにペナルティ新しいトピックを促進

試験のヒント: よくある問題:「常に同じ出力(決定論的)を得るにはどのパラメータを設定すべきか?」→ temperature = 0.0。「単語の繰り返しを減らす」場合 → repetition_penalty > 1.0 または frequency_penalty > 0 を使用します。

3. NVIDIA NIM(NVIDIA Inference Microservices)

3.1. NIMとは?

NVIDIA NIMは、NVIDIA GPU上でLLM/マルチモーダルモデルを最高のパフォーマンスでデプロイするための最適化済み推論コンテナのセットです。NIMにはTensorRT-LLM、量子化、メモリ最適化が組み込まれています。

主な特徴:

  • OpenAI互換API — ドロップイン置換可能、openaiクライアントで直接呼び出し
  • TensorRT-LLMバックエンド — NVIDIA GPU向け最適化カーネル
  • Continuous batching — 複数リクエストを効率的に同時処理
  • gRPC + REST API — 柔軟な統合
  • マルチGPUサポート — 自動テンソル並列化

3.2. NIMアーキテクチャ


NVIDIA NIMアーキテクチャ
════════════════════════

┌─────────────────────────────────────────────┐
│              NIMコンテナ                      │
│                                              │
│  ┌──────────┐   ┌──────────────────────┐    │
│  │  REST API │   │   gRPC Endpoint      │    │
│  │ :8000     │   │   :8001              │    │
│  └─────┬────┘   └──────────┬───────────┘    │
│        │                    │                │
│        └────────┬───────────┘                │
│                 ▼                             │
│  ┌──────────────────────────────────┐       │
│  │     Request Router & Batcher     │       │
│  │     (Continuous Batching)        │       │
│  └──────────────┬───────────────────┘       │
│                 ▼                             │
│  ┌──────────────────────────────────┐       │
│  │     TensorRT-LLM Engine          │       │
│  │  ┌────────┐ ┌────────────────┐   │       │
│  │  │ KV Cache│ │ Paged Attention│   │       │
│  │  └────────┘ └────────────────┘   │       │
│  └──────────────┬───────────────────┘       │
│                 ▼                             │
│  ┌──────────────────────────────────┐       │
│  │       NVIDIA GPU(s)              │       │
│  │   A100 / H100 / L40S            │       │
│  └──────────────────────────────────┘       │
└─────────────────────────────────────────────┘

3.3. NIMコンテナのPull & Run


# Llama-3用NIMコンテナのPullと実行
# 要件:NVIDIA GPU、Docker + NVIDIA Container Toolkit

# ターミナルコマンド:
# docker run -it --rm --gpus all \
#   -p 8000:8000 \
#   -e NGC_API_KEY=$NGC_API_KEY \
#   nvcr.io/nim/meta/llama-3.1-8b-instruct:latest

3.4. NIM APIの呼び出し


from openai import OpenAI

# NIMはOpenAI API互換 — base_urlを変更するだけ
client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-used"  # ローカルNIMではキー不要
)

response = client.chat.completions.create(
    model="meta/llama-3.1-8b-instruct",
    messages=[
        {"role": "system", "content": "You are a helpful AI assistant."},
        {"role": "user", "content": "Explain the Transformer architecture"}
    ],
    temperature=0.7,
    top_p=0.9,
    max_tokens=512
)

print(response.choices[0].message.content)

3.5. NIMとHuggingFace推論の比較

基準NVIDIA NIMHuggingFace Transformers
バックエンドTensorRT-LLMPyTorch
スループット(tokens/s)~2500-4000~300-800
レイテンシ(TTFT)~50-100ms~200-500ms
バッチ処理Continuous batching手動 / 静的
APIOpenAI互換RESTPython API
セットアップdocker runコマンド1つライブラリインストール + コード
量子化組み込み(FP8、INT4)別途GPTQ/AWQが必要
本番対応対応済み(モニタリング、スケーリング)追加のサービングレイヤーが必要

試験のヒント: 試験で「NVIDIA GPUでLLMを最速でデプロイする方法」や「TensorRT-LLM最適化による本番対応推論」を問われた場合、NIMが常に正解です。NIM ≠ 学習フレームワーク — 推論専用です。

4. LangChain LCELパイプライン設計

4.1. LCELとは?

LangChain Expression Language(LCEL)は、LLM処理パイプラインを構築するための宣言的構文です。|(パイプ)演算子でコンポーネントをチェーン接続します — Unixパイプに似ています。

LCELの利点:

  • ストリーミング — トークン単位の出力ストリーミングに対応
  • 非同期 — ネイティブasyncサポート
  • バッチ処理 — 複数入力の同時処理
  • リトライ/フォールバック — エラー時の自動リトライ
  • トレーシング — LangSmithとの統合でデバッグ

4.2. コアプリミティブ

コンポーネント役割入力 → 出力
PromptTemplate変数を使ってプロンプトをフォーマットdict → PromptValue
ChatPromptTemplateチャットメッセージをフォーマットdict → ChatPromptValue
ChatModelLLMを呼び出す(ChatOpenAI、ChatNVIDIA...)PromptValue → AIMessage
StrOutputParserAIMessageから文字列を抽出AIMessage → str
JsonOutputParser出力からJSONをパースAIMessage → dict
RunnablePassthrough入力をそのまま通過any → any
RunnableLambda関数をRunnableとしてラップany → any
RunnableParallel複数チェーンを並列実行dict → dict

4.3. LCELパイプラインの流れ


LCELパイプラインアーキテクチャ
════════════════════════════

シンプルチェーン:
─────────────
  {"topic": "AI"}
        │
        ▼
┌───────────────┐    ┌─────────────┐    ┌────────────────┐
│ PromptTemplate │──►│  ChatModel   │──►│ StrOutputParser │──► "AI is..."
│ "Explain {topic}"│  │ (ChatNVIDIA) │    │                │
└───────────────┘    └─────────────┘    └────────────────┘

       prompt      |      llm       |      parser
                   LCEL: prompt | llm | parser


並列チェーン(RunnableParallel):
───────────────────────────────────
                 {"topic": "AI"}
                       │
              ┌────────┴────────┐
              ▼                 ▼
     ┌──────────────┐  ┌──────────────┐
     │  chain_summary│  │  chain_quiz  │
     │  prompt | llm │  │  prompt | llm│
     └──────┬───────┘  └──────┬───────┘
              │                 │
              └────────┬────────┘
                       ▼
            {"summary": "...", "quiz": "..."}

4.4. コード:LCELチェーン


from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_nvidia_ai_endpoints import ChatNVIDIA

# 1. コンポーネントの初期化
prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a {domain} expert. Answer concisely."),
    ("human", "{question}")
])

llm = ChatNVIDIA(
    model="meta/llama-3.1-8b-instruct",
    temperature=0.3,
    top_p=0.9,
    max_tokens=512
)

parser = StrOutputParser()

# 2. LCELパイプ構文でチェーンを作成
chain = prompt | llm | parser

# 3. 同期呼び出し
result = chain.invoke({
    "domain": "deep learning",
    "question": "How does Transformer self-attention work?"
})
print(result)

# 4. ストリーミング(トークン単位)
for chunk in chain.stream({
    "domain": "deep learning",
    "question": "Compare RNN and Transformer"
}):
    print(chunk, end="", flush=True)

4.5. 応用:RunnableParallelとRunnableLambda


from langchain_core.runnables import (
    RunnablePassthrough,
    RunnableParallel,
    RunnableLambda
)

# カスタム関数をRunnableとしてラップ
def word_count(text: str) -> dict:
    return {"text": text, "word_count": len(text.split())}

# 並列チェーン:要約と単語数カウントを同時実行
parallel_chain = RunnableParallel(
    summary=prompt | llm | parser,
    metadata=RunnableLambda(
        lambda x: f"Query: {x['question']}"
    )
)

# パススルー付きチェーン — パイプライン全体で元の入力を保持
chain_with_context = (
    RunnablePassthrough.assign(
        answer=prompt | llm | parser
    )
)

# 並列呼び出し
result = parallel_chain.invoke({
    "domain": "AI",
    "question": "What is Generative AI?"
})
# result = {"summary": "...", "metadata": "Query: What is Generative AI?"}

試験のヒント: 試験でLCELコードが提示され「出力の型は何か?」と問われた場合、各ステップをトレースしてください:PromptTemplate → PromptValue、ChatModel → AIMessage、StrOutputParser → str。パーサーを忘れると、出力は文字列ではなくAIMessageオブジェクトになります。

5. GradioでUIを構築 & LangServeでAPIを構築

5.1. Gradio:高速チャットボットUI

Gradioを使えば、わずか数行のコードでMLモデルのWeb UIを作成できます。gr.ChatInterfaceコンポーネントはチャットボットに特に適しています。


import gradio as gr
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_nvidia_ai_endpoints import ChatNVIDIA

# チェーンのセットアップ
prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a friendly AI assistant."),
    ("human", "{message}")
])
llm = ChatNVIDIA(model="meta/llama-3.1-8b-instruct")
chain = prompt | llm | StrOutputParser()

# Gradioハンドラー
def respond(message, history):
    """チャットメッセージを処理 — historyは[user, bot]ペアのリスト。"""
    response = chain.invoke({"message": message})
    return response

# UIを起動
demo = gr.ChatInterface(
    fn=respond,
    title="NVIDIA NIM Chatbot",
    description="Chatbot powered by Llama 3.1 via NIM",
    examples=["What is Generative AI?", "Compare GAN and Diffusion"],
    theme="soft"
)
demo.launch(server_port=7860)

5.2. LangServe:チェーンをREST APIとして公開

LangServeは任意のLCELチェーンを自動ドキュメント(Swagger)付きのREST APIに変換します。本番デプロイに適しています。


# === サーバー (server.py) ===
from fastapi import FastAPI
from langserve import add_routes
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_nvidia_ai_endpoints import ChatNVIDIA

app = FastAPI(title="LLM API")

# チェーンの作成
chain = (
    ChatPromptTemplate.from_messages([
        ("system", "AI assistant specializing in {domain}."),
        ("human", "{question}")
    ])
    | ChatNVIDIA(model="meta/llama-3.1-8b-instruct")
    | StrOutputParser()
)

# /chatエンドポイントでチェーンを公開
add_routes(app, chain, path="/chat")

# 実行: uvicorn server:app --port 8080

# === クライアント (client.py) ===
from langserve import RemoteRunnable

# LangServeエンドポイントに接続
chain = RemoteRunnable("http://localhost:8080/chat")

# ローカルチェーンと同様に呼び出し
result = chain.invoke({
    "domain": "machine learning",
    "question": "What is overfitting?"
})
print(result)

# ストリーミングも対応
for chunk in chain.stream({
    "domain": "NLP",
    "question": "How does tokenization work?"
}):
    print(chunk, end="")

Gradio + LangServeデプロイパターン
═══════════════════════════════════════

   ブラウザ(ユーザー)         モバイルアプリ / サービス
        │                          │
        ▼                          ▼
┌──────────────┐          ┌──────────────┐
│ Gradio UI     │          │ RESTクライアント│
│ :7860         │          │               │
└──────┬───────┘          └──────┬───────┘
       │                         │
       └────────┬────────────────┘
                ▼
      ┌──────────────────┐
      │  LangServe API    │
      │  FastAPI :8080    │
      │  /chat/invoke     │
      │  /chat/stream     │
      └────────┬─────────┘
               ▼
      ┌──────────────────┐
      │  LCELチェーン      │
      │  prompt|llm|parser│
      └────────┬─────────┘
               ▼
      ┌──────────────────┐
      │  NVIDIA NIM       │
      │  :8000            │
      └──────────────────┘

試験のヒント: Gradio = プロトタイピング/デモUI、LangServe = 本番REST API。試験で「チャットボットを最速でデモする方法」→ Gradio。「複数クライアントにチェーンを公開」→ LangServe。両方を組み合わせて使うことも可能です。

6. 対話管理とマルチターン会話

6.1. メモリの種類

チャットボットは前の会話ターンのコンテキストを記憶する必要があります。LangChainは複数のメモリタイプを提供しています:

メモリタイプ仕組み利点欠点
ConversationBufferMemory履歴全体を保存情報の損失なしトークン数が急速に増加
ConversationBufferWindowMemory直近N回のターンを保持トークン使用量を制御古いコンテキストを失う
ConversationSummaryMemoryLLMで履歴を要約効率的な圧縮追加のLLM呼び出しコスト
ConversationSummaryBufferMemory古いものを要約 + 直近はそのまま保持詳細と圧縮のバランスより複雑

6.2. メッセージタイプ

LangChainは型付きメッセージでロールを区別します:


from langchain_core.messages import (
    SystemMessage,
    HumanMessage,
    AIMessage
)

messages = [
    SystemMessage(content="You are an AI assistant."),
    HumanMessage(content="Hello!"),
    AIMessage(content="Hi there! How can I help?"),
    HumanMessage(content="Explain the attention mechanism"),
]

6.3. コード:メモリ付きマルチターンチャットボット


from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain_core.output_parsers import StrOutputParser
from langchain_core.chat_history import InMemoryChatMessageHistory
from langchain_core.runnables.history import RunnableWithMessageHistory
from langchain_nvidia_ai_endpoints import ChatNVIDIA

# 1. メッセージ履歴用スロット付きプロンプト
prompt = ChatPromptTemplate.from_messages([
    ("system", "You are an AI assistant. Answer concisely."),
    MessagesPlaceholder(variable_name="history"),
    ("human", "{input}")
])

llm = ChatNVIDIA(model="meta/llama-3.1-8b-instruct")
chain = prompt | llm | StrOutputParser()

# 2. セッションストア — 各ユーザーが独自の履歴を持つ
session_store = {}

def get_session_history(session_id: str):
    if session_id not in session_store:
        session_store[session_id] = InMemoryChatMessageHistory()
    return session_store[session_id]

# 3. メッセージ履歴でチェーンをラップ
chain_with_history = RunnableWithMessageHistory(
    chain,
    get_session_history,
    input_messages_key="input",
    history_messages_key="history"
)

# 4. チャット — 同じsession_idでコンテキストを保持
config = {"configurable": {"session_id": "user-123"}}

r1 = chain_with_history.invoke(
    {"input": "My name is Minh"},
    config=config
)
print(r1)  # "Hello Minh!..."

r2 = chain_with_history.invoke(
    {"input": "What is my name?"},
    config=config
)
print(r2)  # "Your name is Minh."  ← コンテキストを記憶!

6.4. ウィンドウメモリパターン


ウィンドウメモリ(k=3):直近3ターンのみ保持
═══════════════════════════════════════════════════════

Turn 1: User: "Hello"              ─┐
Turn 2: AI: "Hi there!"             │ ← ターン数 > 3+k で削除
Turn 3: User: "My name is Minh"     │
Turn 4: AI: "Hello Minh!"          ─┘

Turn 5: User: "Explain CNN"         ─┐
Turn 6: AI: "CNN is..."              │ ← 保持
Turn 7: User: "Compare with RNN?"   ─┘

送信されるプロンプト:[System] + [Turn 5,6,7] + [Turn 8 input]
→ トークンを節約するが、「name is Minh」のコンテキストは失われる

試験のヒント: 「チャットボットが数ターン後にコンテキストを忘れる」→ BufferWindowMemoryが小さすぎるか、メモリが一切ない状態です。「トークン上限超過」→ ConversationSummaryMemoryに切り替えて履歴を圧縮します。

7. 推論フレームワーク比較

機能NVIDIA NIMvLLMTGI(HuggingFace)Ollama
バックエンドTensorRT-LLMPagedAttentionPyTorch + Flashllama.cpp
GPU必須NVIDIA(A100/H100)NVIDIANVIDIA不要(CPU可)
スループット最高非常に高い高い低い
量子化FP8、INT4組み込みAWQ、GPTQGPTQ、bitsandbytesGGUF
APIOpenAI互換OpenAI互換カスタム + MessagesOpenAI互換
セットアップDocker(NGC)pip installDockerバイナリ1つ
最適な用途エンタープライズ、本番環境研究、高スループットHFエコシステムローカル開発、ノートPC
NVIDIA最適化✅ 最も深い✅ 良好部分的❌

試験のヒント: NVIDIA DLI試験では、本番デプロイの問題はすべてNIMが有利です。「NVIDIA GPUで最高性能」→ NIM。「ノートPCで素早くローカルテスト」→ Ollama。「オープンソースの高スループット」→ vLLM。

8. チートシート

概念ポイント
temperature = 0.0決定論的出力(再現可能)
temperature = 1.0+創造的、よりランダム
top_p = 0.1最も確実なトークンのみ選択
top_k = 50候補を50トークンに制限
NIM最適化済みコンテナ、TensorRT-LLM、OpenAI API
LCELパイプprompt | llm | parser
RunnableParallel複数チェーンを同時実行
GradioデモUI、gr.ChatInterface
LangServeLCELチェーンからREST API、FastAPI
BufferMemory履歴全体を保存 → トークン数が急速に増加
SummaryMemoryLLMで履歴を圧縮 → トークンを節約
WindowMemory(k=N)直近Nターンを保持
MessagesPlaceholderプロンプト内のチャット履歴スロット
RunnableWithMessageHistoryチェーン + セッションベースメモリのラップ

9. 練習問題

Q1:ストリーミング付きLCELチェーンの構築

PromptTemplate → ChatNVIDIA → StrOutputParserを使ったLCELチェーンを作成してください。プロンプトはtopicを受け取り、LLMにその説明を求めます。ストリーミング出力を追加してください。

回答 Q1を表示

from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_nvidia_ai_endpoints import ChatNVIDIA

# プロンプトテンプレートの作成
prompt = ChatPromptTemplate.from_messages([
    ("system", "You are an AI teacher. Explain clearly and simply."),
    ("human", "Explain in detail: {topic}")
])

# LLMの作成
llm = ChatNVIDIA(
    model="meta/llama-3.1-8b-instruct",
    temperature=0.5,
    max_tokens=1024
)

# パーサーの作成
parser = StrOutputParser()

# LCELチェーン
chain = prompt | llm | parser

# 同期呼び出し(結果を一括で返す)
result = chain.invoke({"topic": "Diffusion Models"})
print(result)

# ストリーミング(トークン単位) — .invoke()の代わりに.stream()を使用
for chunk in chain.stream({"topic": "Diffusion Models"}):
    print(chunk, end="", flush=True)

# 解説:
# - .invoke()はチェーンを呼び出し、完全な出力を待つ
# - .stream()はイテレータを返し、各チャンクが出力の一部
# - StrOutputParserは文字列チャンクをそのまま通すのでストリーミング可能
# - JsonOutputParserを使用すると、ストリームは部分的なJSONを返す

Q2:NIMの設定とTemperatureの比較

OpenAIクライアントを使ってNIMエンドポイントを呼び出してください。同じプロンプトでtemperature=0.0とtemperature=1.0の出力を比較してください。各設定で3回実行し、違いを観察してください。

回答 Q2を表示

from openai import OpenAI

# NIMエンドポイントに接続
client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-used"
)

prompt_msg = [
    {"role": "system", "content": "Answer concisely in 1-2 sentences."},
    {"role": "user", "content": "Why is the sky blue?"}
]

print("=== Temperature = 0.0(決定論的)===")
for i in range(3):
    resp = client.chat.completions.create(
        model="meta/llama-3.1-8b-instruct",
        messages=prompt_msg,
        temperature=0.0,  # 常に最高確率のトークンを選択
        max_tokens=100
    )
    print(f"Run {i+1}: {resp.choices[0].message.content}")
# → 3回とも同一の出力

print("\n=== Temperature = 1.0(創造的)===")
for i in range(3):
    resp = client.chat.completions.create(
        model="meta/llama-3.1-8b-instruct",
        messages=prompt_msg,
        temperature=1.0,  # より広い分布、よりランダム
        max_tokens=100
    )
    print(f"Run {i+1}: {resp.choices[0].message.content}")
# → 3回とも異なる出力

# 重要なポイント:
# - temp=0.0:貪欲デコーディング、再現可能、事実に基づくタスクに使用
# - temp=1.0:より広いサンプリング、創造的、ブレインストーミングに使用
# - NIMはOpenAI互換APIのため、クライアントコードは同一

Q3:メモリ付きマルチターンチャットボット

ConversationBufferMemoryをRunnableWithMessageHistory経由でLCELチェーンに統合したチャットボットを作成してください。ボットは前のターンでユーザーの名前を記憶できる必要があります。

回答 Q3を表示

from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain_core.output_parsers import StrOutputParser
from langchain_core.chat_history import InMemoryChatMessageHistory
from langchain_core.runnables.history import RunnableWithMessageHistory
from langchain_nvidia_ai_endpoints import ChatNVIDIA

# 1. 履歴プレースホルダー付きプロンプト
prompt = ChatPromptTemplate.from_messages([
    ("system", "You are a friendly assistant. Remember information the user shares."),
    MessagesPlaceholder(variable_name="history"),
    ("human", "{input}")
])

# 2. チェーン
llm = ChatNVIDIA(model="meta/llama-3.1-8b-instruct", temperature=0.3)
chain = prompt | llm | StrOutputParser()

# 3. セッションストア
store = {}
def get_history(session_id: str):
    if session_id not in store:
        store[session_id] = InMemoryChatMessageHistory()
    return store[session_id]

# 4. メッセージ履歴でラップ
chatbot = RunnableWithMessageHistory(
    chain,
    get_history,
    input_messages_key="input",
    history_messages_key="history"
)

# 5. マルチターンのテスト
cfg = {"configurable": {"session_id": "demo-001"}}

print(chatbot.invoke({"input": "My name is Lan"}, config=cfg))
# → "Hello Lan! Nice to meet you..."

print(chatbot.invoke({"input": "What is my name?"}, config=cfg))
# → "Your name is Lan." ← ボットがコンテキストを記憶!

print(chatbot.invoke({"input": "I like machine learning"}, config=cfg))
# → "That's great, Lan! Machine learning is..."

# 保存された履歴を確認
history = store["demo-001"]
for msg in history.messages:
    print(f"{msg.type}: {msg.content[:50]}...")

Q4:Gradio ChatInterface + LangChain

gr.ChatInterfaceを使ったGradio UIチャットボットを作成し、バックエンドとしてNIMを呼び出すLCELチェーンを使用してください。ストリーミングレスポンスに対応してください。

回答 Q4を表示

import gradio as gr
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_nvidia_ai_endpoints import ChatNVIDIA

# 1. LCELチェーンのセットアップ
prompt = ChatPromptTemplate.from_messages([
    ("system", "You are an AI assistant specializing in deep learning."),
    ("human", "{message}")
])
llm = ChatNVIDIA(
    model="meta/llama-3.1-8b-instruct",
    temperature=0.7
)
chain = prompt | llm | StrOutputParser()

# 2. Gradio用ストリーミングハンドラー
def respond_stream(message, history):
    """
    Gradio ChatInterfaceがこの関数を呼び出します。
    - message:ユーザーの新しいメッセージ
    - history:[user_msg, bot_msg]ペアのリスト
    各チャンクをyieldしてGradioがストリーミング表示。
    """
    partial = ""
    for chunk in chain.stream({"message": message}):
        partial += chunk
        yield partial  # yieldのたびにGradioがUIを更新

# 3. Gradioアプリを起動
demo = gr.ChatInterface(
    fn=respond_stream,
    title="🤖 DL Assistant (NIM-powered)",
    description="Ask anything about Deep Learning",
    examples=[
        "How does Transformer work?",
        "Compare CNN and ViT",
        "What is Batch Normalization used for?"
    ],
    theme="soft"
)

demo.launch(server_port=7860, share=False)

# アクセス:http://localhost:7860
# Gradioがストリーミングレスポンスをリアルタイムで表示

Q5:デバッグ — チェーンが空の出力を返す

以下のコードは実行されますが、出力が常に空または予期しないオブジェクトです。バグを見つけて修正してください。


# バグ:チェーンが文字列ではなくAIMessageオブジェクトを返す
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import JsonOutputParser  # ← うーん...
from langchain_nvidia_ai_endpoints import ChatNVIDIA

prompt = ChatPromptTemplate.from_messages([
    ("system", "Answer concisely in plain text."),
    ("human", "{question}")
])

llm = ChatNVIDIA(model="meta/llama-3.1-8b-instruct")
chain = prompt | llm | JsonOutputParser()  # ← バグはここ

result = chain.invoke({"question": "What is AI?"})
print(result)  # → エラーまたは空/おかしな出力
回答 Q5を表示

# バグ分析:
# - プロンプトはLLMにプレーンテキストで回答するよう指示している
# - しかしパーサーはJsonOutputParser → JSON形式を期待
# - LLMは "AI is artificial intelligence..."(JSONではない)を返す
# - JsonOutputParserがパースを試みる → 失敗または空の出力

# 修正:JsonOutputParserをStrOutputParserに置き換え

from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser  # ← 修正!
from langchain_nvidia_ai_endpoints import ChatNVIDIA

prompt = ChatPromptTemplate.from_messages([
    ("system", "Answer concisely in plain text."),
    ("human", "{question}")
])

llm = ChatNVIDIA(model="meta/llama-3.1-8b-instruct")
chain = prompt | llm | StrOutputParser()  # ← StrOutputParser

result = chain.invoke({"question": "What is AI?"})
print(result)  # → "AI (Artificial Intelligence) is..."

# ルール:OutputParserの型は出力形式と一致させる必要があります:
# - プレーンテキスト → StrOutputParser
# - JSON出力(プロンプトでJSONを要求する必要あり) → JsonOutputParser
# - 構造化出力 → PydanticOutputParser
# 不一致の場合 → チェーンがサイレントに失敗するかエラーが発生