はじめに
質問応答 (QA) — 自動質疑応答システム — は、カスタマー サポートの自動化、社内ナレッジ ベース、教育ツールなど、最も実用的な価値を持つ NLP アプリケーションの 1 つです。
1. QAの種類
| タイプ | 答え方 | 例 |
|---|---|---|
| 抽出 | コンテキストからテキストを抽出する | SQuAD、「Apple CEO はティム・クック」 |
| 抽象的 | 新しい回答を生成 | T5、GPT 生成応答 |
| オープンドメイン | コンテキストを検索して答えます | ウィキペディア検索 + QA |
| クローズドドメイン | 文書内で回答 | FAQチャットボット |
2. BERT を使用した抽出 QA
2.1 アイデア
Context: "Tim Cook là CEO của Apple từ năm 2011. Apple có trụ sở tại Cupertino."
Question: "Ai là CEO của Apple?"
Answer: "Tim Cook" ← trích xuất từ context
↑ start ↑ end
BERT は、コンテキスト内の回答の 開始位置 と 終了位置 を予測します。
2.2 実践
from transformers import pipeline
# Pre-trained QA
qa = pipeline("question-answering", model="deepset/roberta-base-squad2")
result = qa(
question="What is the capital of France?",
context="France is a country in Western Europe. Its capital is Paris, a major European city."
)
print(f"Answer: {result['answer']} (score: {result['score']:.4f})")
# Answer: Paris (score: 0.9834)
2.3 SQuaAD の微調整
from transformers import (
AutoTokenizer,
AutoModelForQuestionAnswering,
Trainer,
TrainingArguments,
)
from datasets import load_dataset
dataset = load_dataset("squad_v2")
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
def preprocess_qa(examples):
questions = [q.strip() for q in examples["question"]]
inputs = tokenizer(
questions,
examples["context"],
max_length=384,
truncation="only_second",
stride=128,
return_overflowing_tokens=True,
return_offsets_mapping=True,
padding="max_length",
)
offset_mapping = inputs.pop("offset_mapping")
sample_map = inputs.pop("overflow_to_sample_mapping")
answers = examples["answers"]
start_positions = []
end_positions = []
for i, offset in enumerate(offset_mapping):
sample_idx = sample_map[i]
answer = answers[sample_idx]
if len(answer["answer_start"]) == 0:
start_positions.append(0)
end_positions.append(0)
continue
start_char = answer["answer_start"][0]
end_char = start_char + len(answer["text"][0])
# Find token positions
sequence_ids = inputs.sequence_ids(i)
idx = 0
while sequence_ids[idx] != 1:
idx += 1
context_start = idx
while idx < len(sequence_ids) and sequence_ids[idx] == 1:
idx += 1
context_end = idx - 1
if offset[context_start][0] > start_char or offset[context_end][1] < end_char:
start_positions.append(0)
end_positions.append(0)
else:
idx = context_start
while idx <= context_end and offset[idx][0] <= start_char:
idx += 1
start_positions.append(idx - 1)
idx = context_end
while idx >= context_start and offset[idx][1] >= end_char:
idx -= 1
end_positions.append(idx + 1)
inputs["start_positions"] = start_positions
inputs["end_positions"] = end_positions
return inputs
tokenized = dataset.map(preprocess_qa, batched=True, remove_columns=dataset["train"].column_names)
model = AutoModelForQuestionAnswering.from_pretrained("bert-base-uncased")
trainer = Trainer(
model=model,
args=TrainingArguments(
output_dir="./qa-model",
num_train_epochs=3,
per_device_train_batch_size=16,
learning_rate=3e-5,
),
train_dataset=tokenized["train"],
)
trainer.train()
3. 検索拡張 QA
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Question │──▶│ Retriever │──▶│ Reader │
│ "CEO Apple?" │ │ (bi-encoder) │ │ (BERT QA) │
└──────────────┘ │ Tìm top-k │ │ Trích answer │
│ documents │ │ từ context │
└──────────────┘ └──────────────┘
from sentence_transformers import SentenceTransformer, util
from transformers import pipeline
# 1. Retriever: tìm passages liên quan
retriever = SentenceTransformer("all-MiniLM-L6-v2")
passages = [
"Tim Cook là CEO của Apple từ năm 2011.",
"Google được sáng lập bởi Larry Page và Sergey Brin.",
"Apple có trụ sở chính tại Cupertino, California.",
]
passage_embeddings = retriever.encode(passages, convert_to_tensor=True)
question = "Ai là CEO của Apple?"
q_embedding = retriever.encode(question, convert_to_tensor=True)
scores = util.cos_sim(q_embedding, passage_embeddings)[0]
top_idx = scores.argmax().item()
# 2. Reader: trích xuất answer
reader = pipeline("question-answering")
answer = reader(question=question, context=passages[top_idx])
print(f"Answer: {answer['answer']}")
# Answer: Tim Cook
概要
| アプローチ | モデル | 利点 | 制限事項 |
|---|---|---|---|
| 抽出 | BERT + QA 責任者 | 正確、速い | 生成ではなく抽出のみ |
| 抽象的な | T5、GPT | 柔軟で自然 | 幻覚が見える |
| 検索 + QA | バイエンコーダー + リーダー | スケーラブルなオープンドメイン | 2 段階の複雑さ |
次の記事
レッスン 16: テキストの要約と機械翻訳 — テキストの要約と機械翻訳: 生成に関する 2 つの最も重要な問題です。