Chuyển đến nội dung chính

レッスン 15: 質問応答 — スマートな質疑応答システム

QA の種類: 抽出的、抽象的、オープンドメイン。 SQuAD データセットと形式。抽出的な QA のために BERT を微調整します。検索拡張 QA。検索のためのクロスエンコーダーとバイエンコーダー。ハンズオンはベトナム人向けの QA システムを構築します。

🧠 AI と ML — レッスン 14 レッスン 15: 質問への回答 — 質問システム 賢く答える

NLP の基礎から上級まで: 自然言語処理をマスターする

パート 5: NLP の応用問題 — 実践プロジェクト

xdev.asia

はじめに

質問応答 (QA) — 自動質疑応答システム — は、カスタマー サポートの自動化、社内ナレッジ ベース、教育ツールなど、最も実用的な価値を持つ NLP アプリケーションの 1 つです。


1. QAの種類

タイプ答え方例
抽出コンテキストからテキストを抽出するSQuAD、「Apple CEO はティム・クック」
抽象的新しい回答を生成T5、GPT 生成応答
オープンドメインコンテキストを検索して答えますウィキペディア検索 + QA
クローズドドメイン文書内で回答FAQチャットボット

2. BERT を使用した抽出 QA

2.1 アイデア

Context: "Tim Cook là CEO của Apple từ năm 2011. Apple có trụ sở tại Cupertino."
Question: "Ai là CEO của Apple?"
Answer: "Tim Cook"  ← trích xuất từ context
         ↑ start     ↑ end

BERT は、コンテキスト内の回答の 開始位置 と 終了位置 を予測します。

2.2 実践

from transformers import pipeline

# Pre-trained QA
qa = pipeline("question-answering", model="deepset/roberta-base-squad2")

result = qa(
    question="What is the capital of France?",
    context="France is a country in Western Europe. Its capital is Paris, a major European city."
)
print(f"Answer: {result['answer']} (score: {result['score']:.4f})")
# Answer: Paris (score: 0.9834)

2.3 SQuaAD の微調整

from transformers import (
    AutoTokenizer,
    AutoModelForQuestionAnswering,
    Trainer,
    TrainingArguments,
)
from datasets import load_dataset

dataset = load_dataset("squad_v2")
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

def preprocess_qa(examples):
    questions = [q.strip() for q in examples["question"]]
    inputs = tokenizer(
        questions,
        examples["context"],
        max_length=384,
        truncation="only_second",
        stride=128,
        return_overflowing_tokens=True,
        return_offsets_mapping=True,
        padding="max_length",
    )

    offset_mapping = inputs.pop("offset_mapping")
    sample_map = inputs.pop("overflow_to_sample_mapping")
    answers = examples["answers"]
    start_positions = []
    end_positions = []

    for i, offset in enumerate(offset_mapping):
        sample_idx = sample_map[i]
        answer = answers[sample_idx]

        if len(answer["answer_start"]) == 0:
            start_positions.append(0)
            end_positions.append(0)
            continue

        start_char = answer["answer_start"][0]
        end_char = start_char + len(answer["text"][0])

        # Find token positions
        sequence_ids = inputs.sequence_ids(i)
        idx = 0
        while sequence_ids[idx] != 1:
            idx += 1
        context_start = idx
        while idx < len(sequence_ids) and sequence_ids[idx] == 1:
            idx += 1
        context_end = idx - 1

        if offset[context_start][0] > start_char or offset[context_end][1] < end_char:
            start_positions.append(0)
            end_positions.append(0)
        else:
            idx = context_start
            while idx <= context_end and offset[idx][0] <= start_char:
                idx += 1
            start_positions.append(idx - 1)

            idx = context_end
            while idx >= context_start and offset[idx][1] >= end_char:
                idx -= 1
            end_positions.append(idx + 1)

    inputs["start_positions"] = start_positions
    inputs["end_positions"] = end_positions
    return inputs

tokenized = dataset.map(preprocess_qa, batched=True, remove_columns=dataset["train"].column_names)

model = AutoModelForQuestionAnswering.from_pretrained("bert-base-uncased")

trainer = Trainer(
    model=model,
    args=TrainingArguments(
        output_dir="./qa-model",
        num_train_epochs=3,
        per_device_train_batch_size=16,
        learning_rate=3e-5,
    ),
    train_dataset=tokenized["train"],
)
trainer.train()

3. 検索拡張 QA

┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│   Question    │──▶│   Retriever   │──▶│    Reader     │
│ "CEO Apple?"  │    │ (bi-encoder)  │    │ (BERT QA)    │
└──────────────┘    │ Tìm top-k     │    │ Trích answer  │
                    │ documents     │    │ từ context    │
                    └──────────────┘    └──────────────┘
from sentence_transformers import SentenceTransformer, util
from transformers import pipeline

# 1. Retriever: tìm passages liên quan
retriever = SentenceTransformer("all-MiniLM-L6-v2")
passages = [
    "Tim Cook là CEO của Apple từ năm 2011.",
    "Google được sáng lập bởi Larry Page và Sergey Brin.",
    "Apple có trụ sở chính tại Cupertino, California.",
]
passage_embeddings = retriever.encode(passages, convert_to_tensor=True)

question = "Ai là CEO của Apple?"
q_embedding = retriever.encode(question, convert_to_tensor=True)
scores = util.cos_sim(q_embedding, passage_embeddings)[0]
top_idx = scores.argmax().item()

# 2. Reader: trích xuất answer
reader = pipeline("question-answering")
answer = reader(question=question, context=passages[top_idx])
print(f"Answer: {answer['answer']}")
# Answer: Tim Cook

概要

アプローチモデル利点制限事項
抽出BERT + QA 責任者正確、速い生成ではなく抽出のみ
抽象的なT5、GPT柔軟で自然幻覚が見える
検索 + QAバイエンコーダー + リーダースケーラブルなオープンドメイン2 段階の複雑さ

次の記事

レッスン 16: テキストの要約と機械翻訳 — テキストの要約と機械翻訳: 生成に関する 2 つの最も重要な問題です。