Chuyển đến nội dung chính

Lesson 10: BERT — Bidirectional Encoder Representations from Transformers

BERT architecture: masked language modeling, next sentence prediction. Pre-training vs fine-tuning paradigm. BERT variants: RoBERTa, ALBERT, DistilBERT, PhoBERT (Vietnamese). Feature extraction vs fine-tuning. Demo classification with Hugging Face Transformers.

🧠 AI & ML — Lesson 9 Lesson 10: BERT — Bidirectional Encoder Representations from Transformers

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 4: Pre-trained Language Models — BERT, GPT & Beyond

xdev.asia

Introduction

BERT (Devlin et al., 2018) is the model that revolutionized NLP — demonstrating for the first time that a model pre-trained on large amounts of text can fine-tune almost any NLP task and reach state-of-the-art. BERT opens the era of transfer learning for NLP.


1. BERT Architecture

Encoder-only Transformer

BERT only uses Transformer's Encoder — processes text bidirectional (both left-right and right-left at the same time).

Input:  [CLS] The cat sat on the mat [SEP]
         │     │    │   │   │   │   │   │
         ▼     ▼    ▼   ▼   ▼   ▼   ▼   ▼
    ┌──────────────────────────────────────┐
    │         Transformer Encoder          │
    │            (12 layers)               │
    │         Self-Attention → FFN         │
    └──────────────────────────────────────┘
         │     │    │   │   │   │   │   │
         ▼     ▼    ▼   ▼   ▼   ▼   ▼   ▼
       T_CLS  T₁   T₂  T₃  T₄  T₅  T₆ T_SEP

[CLS] → Classification head
[SEP] → Separator between sentences

Pre-training Objectives

ObjectiveHow it works
MLM (Masked Language Modeling)Cover 15% of tokens, predict the covered word
NSP (Next Sentence Prediction)Predict whether sentence B is the next sentence from sentence A
MLM: "The [MASK] sat on the [MASK]" → "The cat sat on the mat"
NSP: Câu A + Câu B → IsNext / NotNext

2. Fine-tuning BERT for Classification

from transformers import BertTokenizer, BertForSequenceClassification
from transformers import Trainer, TrainingArguments
from datasets import load_dataset

# 1. Load pre-trained BERT
model_name = "bert-base-uncased"
tokenizer = BertTokenizer.from_pretrained(model_name)
model = BertForSequenceClassification.from_pretrained(
    model_name, num_labels=3
)

# 2. Tokenize dataset
def tokenize_fn(examples):
    return tokenizer(
        examples["text"],
        padding="max_length",
        truncation=True,
        max_length=128,
    )

dataset = load_dataset("emotion")
tokenized = dataset.map(tokenize_fn, batched=True)

# 3. Fine-tune
training_args = TrainingArguments(
    output_dir="./results",
    num_train_epochs=3,
    per_device_train_batch_size=16,
    learning_rate=2e-5,
    evaluation_strategy="epoch",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
    eval_dataset=tokenized["validation"],
)
trainer.train()

3. BERT Variants

ModelDifferenceParametersVietnamese
BERT-baseOriginal110MNot good
RoBERTaAbandon NSP, train longer125MNo
ALBERTFactorized embeddings12M–235MNo
DistilBERTDistilled BERT, 40% smaller66MNo
PhoBERTPre-trained for Vietnamese135MYes!
XLM-RoBERTaMultilingual, 100 languages ​​270MYes

PhoBERT for Vietnamese

from transformers import AutoTokenizer, AutoModelForSequenceClassification

# PhoBERT — BERT cho tiếng Việt
tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")
model = AutoModelForSequenceClassification.from_pretrained(
    "vinai/phobert-base-v2", num_labels=3
)

text = "Sản phẩm này rất tốt, tôi rất hài lòng"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
outputs = model(**inputs)
print(outputs.logits)

4. Feature Extraction vs Fine-tuning

How to useMeaningWhen to use
Feature extractionFreeze BERT, only train classifier headLittle data, limited compute
Fine-tuningTrain both BERT (small learning rate)Enough data, need high accuracy
LoRA/AdapterAdd small trainable layersBalance quality/cost
# Feature extraction: freeze BERT
for param in model.bert.parameters():
    param.requires_grad = False
# Chỉ train classification head

Summary

Key pointsDetails
BERTEncoder-only, bidirectional, MLM + NSP
Transfer learningPre-train → Fine-tune paradigm
PhoBERTBERT for Vietnamese
Fine-tuning2-5 epochs, lr=2e-5, batch=16-32

Next article

Lesson 11: GPT & Autoregressive Models — The other side of the Transformer: decoder-only, causal language modeling, and the path to ChatGPT.