Chuyển đến nội dung chính

第十二課:抱臉生態系-現代NLP實踐

Transformers 庫深入研究:管道、AutoModel、AutoTokenizer。模型中心:尋找並使用預先訓練的模型。數據集庫。用於快速微調的訓練器 API。 PEFT/LoRA 用於高效調整。加速多 GPU。演示空間。

🧠 人工智慧與機器學習 — 第 11 課 第十二課:抱臉生態系-實踐 現代自然語言處理

NLP 從基礎到進階:掌握自然語言處理

第 4 部分:預訓練語言模型 — BERT、GPT 及其他

亞洲開發網

簡介

Hugging Face 是「AI 的 GitHub」——一個共享模型、資料集和工具的平台,每個 NLP/AI 工程師都需要了解這些。圖書館 transformers 是全球最受歡迎的NLP練習工具。


1. Transformers 庫 — 快速入門

Pipeline API(5行程式碼)

from transformers import pipeline

# Sentiment Analysis
classifier = pipeline("sentiment-analysis")
result = classifier("NLP is amazing!")
print(result)  # [{'label': 'POSITIVE', 'score': 0.9998}]

# NER
ner = pipeline("ner", grouped_entities=True)

# Question Answering
qa = pipeline("question-answering")

# Summarization
summarizer = pipeline("summarization")

# Translation
translator = pipeline("translation_en_to_vi", model="Helsinki-NLP/opus-mt-en-vi")

# Zero-shot Classification
zero_shot = pipeline("zero-shot-classification")
result = zero_shot(
    "NLP giúp máy tính hiểu ngôn ngữ",
    candidate_labels=["technology", "sports", "politics"],
)
print(result)  # technology: 0.95

自動模型與自動標記器

from transformers import AutoTokenizer, AutoModelForSequenceClassification

model_name = "vinai/phobert-base-v2"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(
    model_name, num_labels=3
)

# Tokenize
inputs = tokenizer("NLP rất thú vị!", return_tensors="pt", padding=True)
print(inputs.keys())  # dict_keys(['input_ids', 'attention_mask'])

# Forward pass
outputs = model(**inputs)
print(outputs.logits.shape)  # torch.Size([1, 3])

2. 資料集庫

from datasets import load_dataset

# Load từ Hub
dataset = load_dataset("imdb")
print(dataset)
# DatasetDict({
#     train: Dataset({features: ['text', 'label'], num_rows: 25000})
#     test: Dataset({features: ['text', 'label'], num_rows: 25000})
# })

# Load CSV/JSON local
dataset = load_dataset("csv", data_files="data.csv")

# Map (preprocessing)
def tokenize_fn(examples):
    return tokenizer(examples["text"], truncation=True, padding="max_length")

tokenized = dataset.map(tokenize_fn, batched=True)

# Filter
short = dataset.filter(lambda x: len(x["text"]) < 200)

# Train/test split
split = dataset["train"].train_test_split(test_size=0.2)

3. Trainer API — 快速微調

from transformers import Trainer, TrainingArguments

training_args = TrainingArguments(
    output_dir="./results",
    num_train_epochs=3,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=32,
    learning_rate=2e-5,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    logging_dir="./logs",
    fp16=True,  # Mixed precision
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
    eval_dataset=tokenized["test"],
    tokenizer=tokenizer,
)

# Train!
trainer.train()

# Evaluate
results = trainer.evaluate()
print(results)

# Save
model.save_pretrained("./my-model")
tokenizer.save_pretrained("./my-model")

4. PEFT & LoRA — 高效能微調

from peft import LoraConfig, get_peft_model, TaskType

# Cấu hình LoRA
lora_config = LoraConfig(
    task_type=TaskType.SEQ_CLS,
    r=8,               # Rank
    lora_alpha=16,
    lora_dropout=0.1,
    target_modules=["query", "value"],
)

# Wrap model với LoRA
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# trainable params: 294,912 || all params: 109,482,240 || trainable%: 0.27%
# → Chỉ train 0.27% parameters!

5. 模型中心 — 尋找預訓練模型

任務熱門型號越南語
分類伯特基地,羅伯塔大vinai/phobert-base-v2
內爾dslim/bert-base-NERdslim/bert-base-NER
品質保證Deepset/羅伯塔基地小隊2—
翻譯赫爾辛基-NLP/opus-mt-*VietAI/envit5-翻譯
嵌入句子轉換器/*BAAI/bge-m3

總結

組件功能
pipeline()快速推理,5行程式碼
AutoModel載入任何預先訓練的模型
datasets載入、處理、快取資料集
Trainer使用內建最佳實踐微調
PEFT/LoRA高效率微調(0.1-1%參數)

下一篇文章

第 13 課:文本分類與情緒分析 — 最常應用的 NLP 問題:文本分類與情緒分析。