Lesson 10: BERT — Bidirectional Encoder Representations from Transformers
BERT architecture: masked language modeling, next sentence prediction. Pre-training vs fine-tuning paradigm. BERT variants: RoBERTa, ALBERT, DistilBERT, PhoBERT (Vietnamese). Feature extraction vs fine-tuning. Demo classification with Hugging Face Transformers.
NLP from Basics to Advanced: Mastering Natural Language Processing
Part 4: Pre-trained Language Models — BERT, GPT & Beyond
xdev.asia
Introduction
BERT (Devlin et al., 2018) is the model that revolutionized NLP — demonstrating for the first time that a model pre-trained on large amounts of text can fine-tune almost any NLP task and reach state-of-the-art. BERT opens the era of transfer learning for NLP.
1. BERT Architecture
Encoder-only Transformer
BERT only uses Transformer's Encoder — processes text bidirectional (both left-right and right-left at the same time).
from transformers import AutoTokenizer, AutoModelForSequenceClassification
# PhoBERT — BERT cho tiếng Việt
tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")
model = AutoModelForSequenceClassification.from_pretrained(
"vinai/phobert-base-v2", num_labels=3
)
text = "Sản phẩm này rất tốt, tôi rất hài lòng"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
outputs = model(**inputs)
print(outputs.logits)
4. Feature Extraction vs Fine-tuning
How to use
Meaning
When to use
Feature extraction
Freeze BERT, only train classifier head
Little data, limited compute
Fine-tuning
Train both BERT (small learning rate)
Enough data, need high accuracy
LoRA/Adapter
Add small trainable layers
Balance quality/cost
# Feature extraction: freeze BERT
for param in model.bert.parameters():
param.requires_grad = False
# Chỉ train classification head
Summary
Key points
Details
BERT
Encoder-only, bidirectional, MLM + NSP
Transfer learning
Pre-train → Fine-tune paradigm
PhoBERT
BERT for Vietnamese
Fine-tuning
2-5 epochs, lr=2e-5, batch=16-32
Next article
Lesson 11: GPT & Autoregressive Models — The other side of the Transformer: decoder-only, causal language modeling, and the path to ChatGPT.