Chuyển đến nội dung chính

Bài 15: Question Answering — Hệ thống Hỏi Đáp Thông minh

QA types: extractive, abstractive, open-domain. SQuAD dataset và format. Fine-tune BERT cho extractive QA. Retrieval-Augmented QA. Cross-encoder vs bi-encoder cho retrieval. Hands-on xây QA system cho tiếng Việt.

🧠 AI & ML — Bài 14 Bài 15: Question Answering — Hệ thống Hỏi Đáp Thông minh

NLP từ Cơ bản đến Nâng cao: Làm chủ Xử lý Ngôn ngữ Tự nhiên

Phần 5: Bài toán NLP Ứng dụng — Hands-on Projects

xdev.asia

Giới thiệu

Question Answering (QA) — hệ thống hỏi đáp tự động — là một trong những ứng dụng NLP có giá trị thực tiễn cao nhất: customer support automation, internal knowledge base, educational tools.


1. Các loại QA

LoạiCách trả lờiVí dụ
ExtractiveTrích xuất đoạn text từ contextSQuAD, "CEO Apple là Tim Cook"
AbstractiveSinh câu trả lời mớiT5, GPT generate response
Open-domainTìm kiếm context rồi trả lờiWikipedia search + QA
Closed-domainTrả lời trong phạm vi tài liệuFAQ chatbot

2. Extractive QA với BERT

2.1 Ý tưởng

Context: "Tim Cook là CEO của Apple từ năm 2011. Apple có trụ sở tại Cupertino."
Question: "Ai là CEO của Apple?"
Answer: "Tim Cook"  ← trích xuất từ context
         ↑ start     ↑ end

BERT dự đoán start position và end position của câu trả lời trong context.

2.2 Hands-on

from transformers import pipeline

# Pre-trained QA
qa = pipeline("question-answering", model="deepset/roberta-base-squad2")

result = qa(
    question="What is the capital of France?",
    context="France is a country in Western Europe. Its capital is Paris, a major European city."
)
print(f"Answer: {result['answer']} (score: {result['score']:.4f})")
# Answer: Paris (score: 0.9834)

2.3 Fine-tune cho SQuAD

from transformers import (
    AutoTokenizer,
    AutoModelForQuestionAnswering,
    Trainer,
    TrainingArguments,
)
from datasets import load_dataset

dataset = load_dataset("squad_v2")
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

def preprocess_qa(examples):
    questions = [q.strip() for q in examples["question"]]
    inputs = tokenizer(
        questions,
        examples["context"],
        max_length=384,
        truncation="only_second",
        stride=128,
        return_overflowing_tokens=True,
        return_offsets_mapping=True,
        padding="max_length",
    )

    offset_mapping = inputs.pop("offset_mapping")
    sample_map = inputs.pop("overflow_to_sample_mapping")
    answers = examples["answers"]
    start_positions = []
    end_positions = []

    for i, offset in enumerate(offset_mapping):
        sample_idx = sample_map[i]
        answer = answers[sample_idx]

        if len(answer["answer_start"]) == 0:
            start_positions.append(0)
            end_positions.append(0)
            continue

        start_char = answer["answer_start"][0]
        end_char = start_char + len(answer["text"][0])

        # Find token positions
        sequence_ids = inputs.sequence_ids(i)
        idx = 0
        while sequence_ids[idx] != 1:
            idx += 1
        context_start = idx
        while idx < len(sequence_ids) and sequence_ids[idx] == 1:
            idx += 1
        context_end = idx - 1

        if offset[context_start][0] > start_char or offset[context_end][1] < end_char:
            start_positions.append(0)
            end_positions.append(0)
        else:
            idx = context_start
            while idx <= context_end and offset[idx][0] <= start_char:
                idx += 1
            start_positions.append(idx - 1)

            idx = context_end
            while idx >= context_start and offset[idx][1] >= end_char:
                idx -= 1
            end_positions.append(idx + 1)

    inputs["start_positions"] = start_positions
    inputs["end_positions"] = end_positions
    return inputs

tokenized = dataset.map(preprocess_qa, batched=True, remove_columns=dataset["train"].column_names)

model = AutoModelForQuestionAnswering.from_pretrained("bert-base-uncased")

trainer = Trainer(
    model=model,
    args=TrainingArguments(
        output_dir="./qa-model",
        num_train_epochs=3,
        per_device_train_batch_size=16,
        learning_rate=3e-5,
    ),
    train_dataset=tokenized["train"],
)
trainer.train()

3. Retrieval-Augmented QA

┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│   Question    │──▶│   Retriever   │──▶│    Reader     │
│ "CEO Apple?"  │    │ (bi-encoder)  │    │ (BERT QA)    │
└──────────────┘    │ Tìm top-k     │    │ Trích answer  │
                    │ documents     │    │ từ context    │
                    └──────────────┘    └──────────────┘
from sentence_transformers import SentenceTransformer, util
from transformers import pipeline

# 1. Retriever: tìm passages liên quan
retriever = SentenceTransformer("all-MiniLM-L6-v2")
passages = [
    "Tim Cook là CEO của Apple từ năm 2011.",
    "Google được sáng lập bởi Larry Page và Sergey Brin.",
    "Apple có trụ sở chính tại Cupertino, California.",
]
passage_embeddings = retriever.encode(passages, convert_to_tensor=True)

question = "Ai là CEO của Apple?"
q_embedding = retriever.encode(question, convert_to_tensor=True)
scores = util.cos_sim(q_embedding, passage_embeddings)[0]
top_idx = scores.argmax().item()

# 2. Reader: trích xuất answer
reader = pipeline("question-answering")
answer = reader(question=question, context=passages[top_idx])
print(f"Answer: {answer['answer']}")
# Answer: Tim Cook

Tổng kết

ApproachModelƯu điểmHạn chế
ExtractiveBERT + QA headChính xác, nhanhChỉ extract, không generate
AbstractiveT5, GPTLinh hoạt, tự nhiênCó thể hallucinate
Retrieval + QABi-encoder + ReaderScalable, open-domain2-stage complexity

Bài tiếp theo

Bài 16: Text Summarization & Machine Translation — Tóm tắt văn bản và dịch máy: hai bài toán generation quan trọng nhất.