Chuyển đến nội dung chính

第 10 課:BERT-Transformers 的雙向編碼器表示

BERT 架構:掩碼語言建模、下一句預測。預訓練與微調範例。 BERT 變體:RoBERTa、ALBERT、DistilBERT、PhoBERT(越南語)。特徵提取與微調。使用 Hugging Face Transformer 進行演示分類。

🧠 人工智慧與機器學習 — 第 9 課 第 10 課:BERT-雙向編碼器 變形金剛的代表

NLP 從基礎到進階:掌握自然語言處理

第 4 部分:預訓練語言模型 — BERT、GPT 及其他

亞洲開發網

簡介

BERT(Devlin 等人,2018)是 NLP 革命性的模型——首次證明在大量文本上預先訓練的模型可以微調幾乎任何 NLP 任務並達到最先進的水平。 BERT 開啟了 NLP 的遷移學習時代。


1. BERT 架構

僅編碼器變壓器

BERT 僅使用 Transformer 的 編碼器 — 處理文字 雙向(同時左右和左右)。

Input:  [CLS] The cat sat on the mat [SEP]
         │     │    │   │   │   │   │   │
         ▼     ▼    ▼   ▼   ▼   ▼   ▼   ▼
    ┌──────────────────────────────────────┐
    │         Transformer Encoder          │
    │            (12 layers)               │
    │         Self-Attention → FFN         │
    └──────────────────────────────────────┘
         │     │    │   │   │   │   │   │
         ▼     ▼    ▼   ▼   ▼   ▼   ▼   ▼
       T_CLS  T₁   T₂  T₃  T₄  T₅  T₆ T_SEP

[CLS] → Classification head
[SEP] → Separator between sentences

預訓練目標

目標它是如何運作的
MLM(掩碼語言建模)覆蓋15%的token,預測覆蓋的單字
NSP(下一句預測)預測句子 B 是否是句子 A 的下一個句子
MLM: "The [MASK] sat on the [MASK]" → "The cat sat on the mat"
NSP: Câu A + Câu B → IsNext / NotNext

2. 微調 BERT 進行分類

from transformers import BertTokenizer, BertForSequenceClassification
from transformers import Trainer, TrainingArguments
from datasets import load_dataset

# 1. Load pre-trained BERT
model_name = "bert-base-uncased"
tokenizer = BertTokenizer.from_pretrained(model_name)
model = BertForSequenceClassification.from_pretrained(
    model_name, num_labels=3
)

# 2. Tokenize dataset
def tokenize_fn(examples):
    return tokenizer(
        examples["text"],
        padding="max_length",
        truncation=True,
        max_length=128,
    )

dataset = load_dataset("emotion")
tokenized = dataset.map(tokenize_fn, batched=True)

# 3. Fine-tune
training_args = TrainingArguments(
    output_dir="./results",
    num_train_epochs=3,
    per_device_train_batch_size=16,
    learning_rate=2e-5,
    evaluation_strategy="epoch",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
    eval_dataset=tokenized["validation"],
)
trainer.train()

3.BERT 變體

型號差異參數越南語
BERT 基礎原創110M不好
羅伯塔放棄NSP,訓練更長125M沒有
阿爾伯特因式分解嵌入12M–235M沒有
蒸餾伯特蒸餾版 BERT,縮小 40%66M沒有
PhoBERT越南語預訓練135M**是的! **
XLM-羅伯塔多語言,100種語言,100種語言270M是的

PhoBERT 越南語

from transformers import AutoTokenizer, AutoModelForSequenceClassification

# PhoBERT — BERT cho tiếng Việt
tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")
model = AutoModelForSequenceClassification.from_pretrained(
    "vinai/phobert-base-v2", num_labels=3
)

text = "Sản phẩm này rất tốt, tôi rất hài lòng"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
outputs = model(**inputs)
print(outputs.logits)

4. 特徵提取與微調

如何使用意義何時使用
特徵提取凍結BERT,只訓練分類器頭資料少,計算有限
微調訓練兩個 BERT(小學習率)資料足夠,準確度要求高
LoRA/轉接器新增小的可訓練層平衡品質/成本
# Feature extraction: freeze BERT
for param in model.bert.parameters():
    param.requires_grad = False
# Chỉ train classification head

總結

重點詳情
伯特僅編碼器、雙向、MLM + NSP
遷移學習預訓練 → 微調範式
菲伯特BERT 越南語
微調2-5 epoch,lr=2e-5,batch=16-32

下一篇文章

第 11 課:GPT 與自回歸模型 — Transformer 的另一面:僅解碼器、因果語言建模以及通往 ChatGPT 的路徑。