BERT 架構:掩碼語言建模、下一句預測。預訓練與微調範例。 BERT 變體:RoBERTa、ALBERT、DistilBERT、PhoBERT(越南語)。特徵提取與微調。使用 Hugging Face Transformer 進行演示分類。
🧠 人工智慧與機器學習 — 第 9 課
第 10 課:BERT-雙向編碼器
變形金剛的代表
NLP 從基礎到進階:掌握自然語言處理
第 4 部分:預訓練語言模型 — BERT、GPT 及其他
亞洲開發網
簡介
BERT (Devlin 等人,2018)是 NLP 革命性的模型 ——首次證明在大量文本上預先訓練的模型可以微調幾乎任何 NLP 任務並達到最先進的水平。 BERT 開啟了 NLP 的遷移學習 時代。
1. BERT 架構
僅編碼器變壓器
BERT 僅使用 Transformer 的 編碼器 — 處理文字 雙向 (同時左右和左右)。
Input: [CLS] The cat sat on the mat [SEP]
│ │ │ │ │ │ │ │
▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼
┌──────────────────────────────────────┐
│ Transformer Encoder │
│ (12 layers) │
│ Self-Attention → FFN │
└──────────────────────────────────────┘
│ │ │ │ │ │ │ │
▼ ▼ ▼ ▼ ▼ ▼ ▼ ▼
T_CLS T₁ T₂ T₃ T₄ T₅ T₆ T_SEP
[CLS] → Classification head
[SEP] → Separator between sentences
預訓練目標
目標 它是如何運作的 MLM(掩碼語言建模) 覆蓋15%的token,預測覆蓋的單字 NSP(下一句預測) 預測句子 B 是否是句子 A 的下一個句子
MLM: "The [MASK] sat on the [MASK]" → "The cat sat on the mat"
NSP: Câu A + Câu B → IsNext / NotNext
2. 微調 BERT 進行分類
from transformers import BertTokenizer, BertForSequenceClassification
from transformers import Trainer, TrainingArguments
from datasets import load_dataset
# 1. Load pre-trained BERT
model_name = "bert-base-uncased"
tokenizer = BertTokenizer.from_pretrained(model_name)
model = BertForSequenceClassification.from_pretrained(
model_name, num_labels=3
)
# 2. Tokenize dataset
def tokenize_fn(examples):
return tokenizer(
examples["text"],
padding="max_length",
truncation=True,
max_length=128,
)
dataset = load_dataset("emotion")
tokenized = dataset.map(tokenize_fn, batched=True)
# 3. Fine-tune
training_args = TrainingArguments(
output_dir="./results",
num_train_epochs=3,
per_device_train_batch_size=16,
learning_rate=2e-5,
evaluation_strategy="epoch",
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["validation"],
)
trainer.train()
3.BERT 變體
型號 差異 參數 越南語 BERT 基礎 原創 110M 不好 羅伯塔 放棄NSP,訓練更長 125M 沒有 阿爾伯特 因式分解嵌入 12M–235M 沒有 蒸餾伯特 蒸餾版 BERT,縮小 40% 66M 沒有 PhoBERT 越南語預訓練 135M **是的! ** XLM-羅伯塔 多語言,100種語言,100種語言 270M 是的
PhoBERT 越南語
from transformers import AutoTokenizer, AutoModelForSequenceClassification
# PhoBERT — BERT cho tiếng Việt
tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")
model = AutoModelForSequenceClassification.from_pretrained(
"vinai/phobert-base-v2", num_labels=3
)
text = "Sản phẩm này rất tốt, tôi rất hài lòng"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
outputs = model(**inputs)
print(outputs.logits)
4. 特徵提取與微調
如何使用 意義 何時使用 特徵提取 凍結BERT,只訓練分類器頭 資料少,計算有限 微調 訓練兩個 BERT(小學習率) 資料足夠,準確度要求高 LoRA/轉接器 新增小的可訓練層 平衡品質/成本
# Feature extraction: freeze BERT
for param in model.bert.parameters():
param.requires_grad = False
# Chỉ train classification head
總結
重點 詳情 伯特 僅編碼器、雙向、MLM + NSP 遷移學習 預訓練 → 微調範式 菲伯特 BERT 越南語 微調 2-5 epoch,lr=2e-5,batch=16-32
下一篇文章
第 11 課:GPT 與自回歸模型 — Transformer 的另一面:僅解碼器、因果語言建模以及通往 ChatGPT 的路徑。