Introduction
Hugging Face is the "GitHub for AI" — a platform for sharing models, datasets, and tools that every NLP/AI engineer needs to know. Library transformers is the world's most popular NLP practice tool.
1. Transformers Library — Quick Start
Pipeline API (5 lines of code)
from transformers import pipeline
# Sentiment Analysis
classifier = pipeline("sentiment-analysis")
result = classifier("NLP is amazing!")
print(result) # [{'label': 'POSITIVE', 'score': 0.9998}]
# NER
ner = pipeline("ner", grouped_entities=True)
# Question Answering
qa = pipeline("question-answering")
# Summarization
summarizer = pipeline("summarization")
# Translation
translator = pipeline("translation_en_to_vi", model="Helsinki-NLP/opus-mt-en-vi")
# Zero-shot Classification
zero_shot = pipeline("zero-shot-classification")
result = zero_shot(
"NLP giúp máy tính hiểu ngôn ngữ",
candidate_labels=["technology", "sports", "politics"],
)
print(result) # technology: 0.95
AutoModel & AutoTokenizer
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_name = "vinai/phobert-base-v2"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(
model_name, num_labels=3
)
# Tokenize
inputs = tokenizer("NLP rất thú vị!", return_tensors="pt", padding=True)
print(inputs.keys()) # dict_keys(['input_ids', 'attention_mask'])
# Forward pass
outputs = model(**inputs)
print(outputs.logits.shape) # torch.Size([1, 3])
2. Datasets Library
from datasets import load_dataset
# Load từ Hub
dataset = load_dataset("imdb")
print(dataset)
# DatasetDict({
# train: Dataset({features: ['text', 'label'], num_rows: 25000})
# test: Dataset({features: ['text', 'label'], num_rows: 25000})
# })
# Load CSV/JSON local
dataset = load_dataset("csv", data_files="data.csv")
# Map (preprocessing)
def tokenize_fn(examples):
return tokenizer(examples["text"], truncation=True, padding="max_length")
tokenized = dataset.map(tokenize_fn, batched=True)
# Filter
short = dataset.filter(lambda x: len(x["text"]) < 200)
# Train/test split
split = dataset["train"].train_test_split(test_size=0.2)
3. Trainer API — Fast fine-tuning
from transformers import Trainer, TrainingArguments
training_args = TrainingArguments(
output_dir="./results",
num_train_epochs=3,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
learning_rate=2e-5,
weight_decay=0.01,
evaluation_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
logging_dir="./logs",
fp16=True, # Mixed precision
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["test"],
tokenizer=tokenizer,
)
# Train!
trainer.train()
# Evaluate
results = trainer.evaluate()
print(results)
# Save
model.save_pretrained("./my-model")
tokenizer.save_pretrained("./my-model")
4. PEFT & LoRA — Efficient Fine-tuning
from peft import LoraConfig, get_peft_model, TaskType
# Cấu hình LoRA
lora_config = LoraConfig(
task_type=TaskType.SEQ_CLS,
r=8, # Rank
lora_alpha=16,
lora_dropout=0.1,
target_modules=["query", "value"],
)
# Wrap model với LoRA
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# trainable params: 294,912 || all params: 109,482,240 || trainable%: 0.27%
# → Chỉ train 0.27% parameters!
5. Model Hub — Find Pre-trained Models
| Tasks | Popular Models | Vietnamese |
|---|---|---|
| Classification | bert-base, roberta-large | vinai/phobert-base-v2 |
| NER | dslim/bert-base-NER | NlpHUST/vibert4news-base-cased |
| QA | deepset/roberta-base-squad2 | — |
| Translation | Helsinki-NLP/opus-mt-* | VietAI/envit5-translation |
| Embeddings | sentence-transformers/* | BAAI/bge-m3 |
Summary
| Components | Function |
|---|---|
pipeline() | Quick inference, 5 lines of code |
AutoModel | Load any pre-trained model |
datasets | Load, process, cache datasets |
Trainer | Fine-tuning with best practices built-in |
PEFT/LoRA | Efficient fine-tuning (0.1-1% parameters) |
Next article
Lesson 13: Text Classification & Sentiment Analysis — The most commonly applied NLP problem: text classification and sentiment analysis.