1. Giới thiệu
Bài học đầu tiên trong series luyện thi NVIDIA DLI — Generative AI sẽ trang bị cho bạn nền tảng PyTorch vững chắc. Đây là framework chính được sử dụng trong toàn bộ khóa học DLI, từ Diffusion Models đến Large Language Models (LLMs).
Trong bài assessment của NVIDIA DLI, bạn sẽ phải viết code trực tiếp — không phải trắc nghiệm. Vì vậy, nắm vững các pattern cơ bản của PyTorch là điều kiện tiên quyết.
Exam tip: NVIDIA DLI assessment yêu cầu bạn viết và debug code PyTorch trực tiếp. Hãy chắc chắn bạn có thể viết training loop, nn.Module, và thao tác tensor mà không cần nhìn tài liệu.

2. PyTorch Tensors & Autograd
2.1 Tensor Basics
Tensor là cấu trúc dữ liệu cốt lõi của PyTorch — tương tự NumPy array nhưng có thể chạy trên GPU và hỗ trợ automatic differentiation.
import torch
# Tạo tensor từ list
x = torch.tensor([1.0, 2.0, 3.0])
# Tạo tensor với shape cụ thể
zeros = torch.zeros(3, 4) # shape: (3, 4)
ones = torch.ones(2, 3, 4) # shape: (2, 3, 4)
rand = torch.randn(64, 3, 32, 32) # batch of 64 RGB 32x32 images
# Kiểm tra shape và dtype
print(rand.shape) # torch.Size([64, 3, 32, 32])
print(rand.dtype) # torch.float32
print(rand.device) # cpu
2.2 Tensor Operations & Broadcasting
PyTorch hỗ trợ broadcasting tương tự NumPy — cho phép thực hiện phép tính giữa tensors có shape khác nhau.
# Reshape operations
x = torch.randn(2, 3, 4)
y = x.view(2, 12) # reshape thành (2, 12)
z = x.permute(0, 2, 1) # swap dims: (2, 4, 3)
w = x.unsqueeze(0) # thêm dim: (1, 2, 3, 4)
# Matrix multiplication
a = torch.randn(3, 4)
b = torch.randn(4, 5)
c = a @ b # shape: (3, 5)
# hoặc: c = torch.matmul(a, b)
# Broadcasting
x = torch.randn(64, 256) # (batch, features)
bias = torch.randn(256) # (features,)
result = x + bias # bias được broadcast: (64, 256)
| Operation | Syntax | Ghi chú |
|---|---|---|
| Reshape | x.view() / x.reshape() | view yêu cầu contiguous memory |
| Transpose | x.permute() / x.T | permute linh hoạt hơn cho nhiều dims |
| Add dim | x.unsqueeze(dim) | Thường dùng để chuẩn bị broadcasting |
| Remove dim | x.squeeze(dim) | Xóa dim có size = 1 |
| Matrix mul | a @ b | Equivalent to torch.matmul |
| Concat | torch.cat([a, b], dim=0) | Nối theo dim chỉ định |
2.3 Autograd — Automatic Differentiation
Autograd là hệ thống tự động tính gradient của PyTorch. Khi bạn set requires_grad=True, PyTorch sẽ theo dõi mọi phép tính trên tensor đó và xây dựng computational graph.
# Autograd cơ bản
x = torch.tensor([2.0, 3.0], requires_grad=True)
y = x ** 2 + 3 * x # y = x² + 3x
loss = y.sum() # scalar output
loss.backward() # tính gradient
print(x.grad) # dy/dx = 2x + 3 → tensor([7., 9.])
Computational Graph:
x (requires_grad=True)
│
├──→ x ** 2 ──→ + ──→ y ──→ sum() ──→ loss
│ ↑ │
└──→ 3 * x ─────┘ backward()
│
x.grad = 2x + 3
Exam tip: Trong DLI assessment, bạn có thể gặp lỗi "trying to backward through the graph a second time". Giải pháp: dùng
loss.backward(retain_graph=True)hoặc tính lại forward pass. Đây là lỗi rất phổ biến khi làm bài.
3. nn.Module & Building Networks
3.1 nn.Module Pattern
Mọi neural network trong PyTorch đều kế thừa từ nn.Module. Đây là pattern bắt buộc phải thuộc lòng:
import torch.nn as nn
class SimpleNet(nn.Module):
def __init__(self, input_dim, hidden_dim, output_dim):
super().__init__() # PHẢI gọi super().__init__()
self.fc1 = nn.Linear(input_dim, hidden_dim)
self.relu = nn.ReLU()
self.fc2 = nn.Linear(hidden_dim, output_dim)
def forward(self, x):
x = self.fc1(x)
x = self.relu(x)
x = self.fc2(x)
return x
# Sử dụng
model = SimpleNet(784, 256, 10)
x = torch.randn(32, 784) # batch of 32
output = model(x) # shape: (32, 10)
3.2 Common Layers
| Layer | Công dụng | Params chính |
|---|---|---|
nn.Linear(in, out) | Fully connected layer | in_features, out_features |
nn.Conv2d(in_ch, out_ch, k) | 2D convolution | in_channels, out_channels, kernel_size |
nn.BatchNorm2d(ch) | Batch normalization | num_features |
nn.GroupNorm(g, ch) | Group normalization | num_groups, num_channels |
nn.ReLU() | Activation function | inplace (optional) |
nn.Dropout(p) | Regularization | p = drop probability |
nn.Embedding(V, D) | Token embedding | num_embeddings, embedding_dim |
3.3 nn.Sequential — Quick Models
Với mô hình đơn giản, bạn có thể dùng nn.Sequential thay vì tạo class:
# Cách nhanh với nn.Sequential
model = nn.Sequential(
nn.Linear(784, 256),
nn.ReLU(),
nn.Dropout(0.2),
nn.Linear(256, 128),
nn.ReLU(),
nn.Linear(128, 10)
)
# Kiểm tra model
print(model)
# Đếm parameters
total_params = sum(p.numel() for p in model.parameters())
print(f"Total params: {total_params:,}")
3.4 Code: MLP cho MNIST
import torch
import torch.nn as nn
from torchvision import datasets, transforms
# Data
transform = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize((0.1307,), (0.3081,))
])
train_data = datasets.MNIST('./data', train=True, download=True,
transform=transform)
train_loader = torch.utils.data.DataLoader(train_data, batch_size=64,
shuffle=True)
# Model
class MNISTClassifier(nn.Module):
def __init__(self):
super().__init__()
self.flatten = nn.Flatten()
self.layers = nn.Sequential(
nn.Linear(28 * 28, 512),
nn.ReLU(),
nn.Dropout(0.2),
nn.Linear(512, 256),
nn.ReLU(),
nn.Dropout(0.2),
nn.Linear(256, 10)
)
def forward(self, x):
x = self.flatten(x) # (B, 1, 28, 28) → (B, 784)
return self.layers(x) # (B, 784) → (B, 10)
4. Training Loop Pattern
4.1 Loss Functions
| Loss Function | Dùng khi | Input shape |
|---|---|---|
nn.CrossEntropyLoss() | Multi-class classification | logits (B, C), labels (B,) |
nn.MSELoss() | Regression, diffusion noise prediction | (B, *) vs (B, *) |
nn.BCEWithLogitsLoss() | Binary / multi-label classification | logits (B, C), labels (B, C) |
nn.L1Loss() | Regression (robust to outliers) | (B, *) vs (B, *) |
Exam tip:
nn.CrossEntropyLossđã bao gồm softmax bên trong — KHÔNG cần thêm softmax ở output layer. Đây là lỗi mà nhiều người mắc phải trong assessment.nn.MSELosssẽ rất quan trọng khi bạn học Diffusion Models (predict noise).
4.2 Optimizers
# SGD — cơ bản, can thiệp learning rate thủ công
optimizer = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.9)
# Adam — phổ biến nhất, adaptive learning rate
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
# AdamW — Adam + weight decay đúng cách (dùng cho Transformers)
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-4, weight_decay=0.01)
| Optimizer | Khi nào dùng | Đặc điểm |
|---|---|---|
| SGD | CNNs, khi cần kiểm soát tỉ mỉ | Cần tune lr cẩn thận, thêm momentum |
| Adam | Mặc định cho hầu hết tasks | Hội tụ nhanh, ít cần tune |
| AdamW | Transformers, LLMs, Diffusion | Weight decay tách riêng, chuẩn hơn Adam |
4.3 Full Training Loop
Đây là pattern quan trọng nhất — bạn phải viết được training loop hoàn chỉnh trong assessment:
import torch
import torch.nn as nn
# Setup
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model = MNISTClassifier().to(device)
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
# Training loop
num_epochs = 10
for epoch in range(num_epochs):
model.train() # BẬT training mode (dropout, batchnorm active)
total_loss = 0
for batch_idx, (images, labels) in enumerate(train_loader):
images, labels = images.to(device), labels.to(device)
# Forward pass
outputs = model(images)
loss = criterion(outputs, labels)
# Backward pass
optimizer.zero_grad() # QUAN TRỌNG: reset gradients
loss.backward() # tính gradients
optimizer.step() # update weights
total_loss += loss.item()
avg_loss = total_loss / len(train_loader)
print(f"Epoch [{epoch+1}/{num_epochs}], Loss: {avg_loss:.4f}")
# Evaluation
model.eval() # TẮT dropout, batchnorm dùng running stats
with torch.no_grad(): # KHÔNG tính gradients → tiết kiệm memory
correct = 0
total = 0
for images, labels in test_loader:
images, labels = images.to(device), labels.to(device)
outputs = model(images)
_, predicted = torch.max(outputs, 1)
total += labels.size(0)
correct += (predicted == labels).sum().item()
print(f"Accuracy: {100 * correct / total:.2f}%")
Training Loop Flow:
┌─────────────────────────────────────────────┐
│ FOR EACH EPOCH │
│ ┌────────────────────────────────────────┐ │
│ │ FOR EACH BATCH │ │
│ │ │ │
│ │ 1. images, labels = batch.to(device) │ │
│ │ │ │ │
│ │ 2. outputs = model(images) FORWARD │ │
│ │ │ │ │
│ │ 3. loss = criterion(outputs, labels) │ │
│ │ │ │ │
│ │ 4. optimizer.zero_grad() RESET │ │
│ │ │ │ │
│ │ 5. loss.backward() BACKWARD │ │
│ │ │ │ │
│ │ 6. optimizer.step() UPDATE │ │
│ │ │ │ │
│ └──────────────┼─────────────────────────┘ │
│ ▼ │
│ Next Epoch │
└─────────────────────────────────────────────┘
Exam tip: Thứ tự
zero_grad() → backward() → step()là BẮT BUỘC. Quênzero_grad()sẽ khiến gradients bị tích lũy qua các batch — đây là bug #1 trong DLI assessment. Luôn nhớmodel.train()trước training vàmodel.eval()+torch.no_grad()trước evaluation.
4.4 GPU Acceleration
# Kiểm tra GPU
print(torch.cuda.is_available()) # True/False
print(torch.cuda.device_count()) # Số GPU
print(torch.cuda.get_device_name(0)) # Tên GPU
# Di chuyển model và data lên GPU
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model = model.to(device)
# TRONG training loop — data cũng phải lên cùng device
images = images.to(device)
labels = labels.to(device)
Exam tip: Lỗi phổ biến: model trên GPU nhưng data vẫn trên CPU (hoặc ngược lại). PyTorch sẽ báo lỗi "Expected all tensors to be on the same device". Luôn đảm bảo model và data cùng
device.
5. CNN Architecture
5.1 Convolutional Layers
Convolution trượt một kernel (filter) qua input image để trích xuất features. Mỗi kernel phát hiện một pattern cụ thể (edges, textures, shapes).
# Conv2d cơ bản
conv = nn.Conv2d(
in_channels=3, # RGB input
out_channels=32, # 32 filters → 32 feature maps
kernel_size=3, # 3×3 kernel
stride=1, # bước nhảy
padding=1 # zero-padding để giữ spatial size
)
# Output shape calculation:
# H_out = (H_in + 2*padding - kernel_size) / stride + 1
# Ví dụ: (32 + 2*1 - 3) / 1 + 1 = 32 (giữ nguyên size)
Convolution Operation:
Input (3 channels) Kernel (3×3) Output (1 feature map)
┌─────────────┐ ┌───────┐ ┌──────────┐
│ ■ ■ ■ ■ ■ ■│ × │ w w w │ = │ ○ ○ ○ ○ │
│ ■ ■ ■ ■ ■ ■│ │ w w w │ │ ○ ○ ○ ○ │
│ ■ ■ ■ ■ ■ ■│ │ w w w │ │ ○ ○ ○ ○ │
│ ■ ■ ■ ■ ■ ■│ └───────┘ │ ○ ○ ○ ○ │
│ ■ ■ ■ ■ ■ ■│ └──────────┘
│ ■ ■ ■ ■ ■ ■│ 32 kernels → 32 feature maps
└─────────────┘
Shape flow: (B, 3, 32, 32) → Conv2d(3, 32, 3, padding=1) → (B, 32, 32, 32)
5.2 Pooling Layers
Pooling giảm spatial dimensions, giúp giảm computation và tăng receptive field:
| Pooling | Cách hoạt động | Khi nào dùng |
|---|---|---|
nn.MaxPool2d(2) | Lấy giá trị max trong mỗi 2×2 window | Feature detection, CNNs thông thường |
nn.AvgPool2d(2) | Lấy trung bình trong mỗi 2×2 window | Smoother features |
nn.AdaptiveAvgPool2d((1,1)) | Pool về size cố định bất kể input | Trước fully-connected layer |
5.3 Batch Normalization vs Group Normalization
Đây là kiến thức cực kỳ quan trọng cho phần Diffusion Models ở các bài sau.
| Thuộc tính | BatchNorm | GroupNorm |
|---|---|---|
| Normalize theo | Batch dimension (N) | Channel groups (C/G) |
| Phụ thuộc batch size | Có — batch nhỏ thì noisy | Không — hoạt động tốt mọi batch size |
| Training vs Inference | Khác nhau (running stats) | Giống nhau |
| Phổ biến trong | CNNs truyền thống (ResNet) | Diffusion Models, nhóm nhỏ |
| Syntax | nn.BatchNorm2d(C) | nn.GroupNorm(G, C) |
# BatchNorm — normalize across batch
bn = nn.BatchNorm2d(64) # 64 channels
# GroupNorm — normalize within groups of channels
gn = nn.GroupNorm(
num_groups=32, # chia 64 channels thành 32 groups (2 ch/group)
num_channels=64
)
# Cả hai nhận input shape: (B, C, H, W)
x = torch.randn(8, 64, 16, 16)
print(bn(x).shape) # (8, 64, 16, 16)
print(gn(x).shape) # (8, 64, 16, 16)
BatchNorm vs GroupNorm:
BatchNorm: normalize theo ↓ (batch axis) GroupNorm: normalize theo → (channel groups)
┌────────────────────────┐ ┌────────────────────────┐
│ Sample 1: [c1 c2 c3 c4]│ │ Sample 1: [c1 c2│c3 c4]│
│ Sample 2: [c1 c2 c3 c4]│ ← normalize │ group1│group2 │ ← normalize
│ Sample 3: [c1 c2 c3 c4]│ mỗi column │ │ │ mỗi group
│ Sample 4: [c1 c2 c3 c4]│ │ Sample 2: [c1 c2│c3 c4]│
└────────────────────────┘ └────────────────────────┘
→ Diffusion Models dùng GroupNorm vì batch_size thường nhỏ
và noise level thay đổi → BatchNorm statistics không ổn định
Exam tip: Khi xây dựng U-Net cho Diffusion Models (bài sau), bạn sẽ luôn dùng GroupNorm thay vì BatchNorm. Lý do: diffusion training thường dùng batch size nhỏ, và mỗi sample có noise level khác nhau → BatchNorm statistics bị noisy. Nhớ quy tắc: Diffusion = GroupNorm.
5.4 Code: Simple CNN cho Image Classification
class SimpleCNN(nn.Module):
def __init__(self, num_classes=10):
super().__init__()
# Conv Block 1: (B, 1, 28, 28) → (B, 32, 14, 14)
self.block1 = nn.Sequential(
nn.Conv2d(1, 32, kernel_size=3, padding=1),
nn.BatchNorm2d(32),
nn.ReLU(),
nn.MaxPool2d(2)
)
# Conv Block 2: (B, 32, 14, 14) → (B, 64, 7, 7)
self.block2 = nn.Sequential(
nn.Conv2d(32, 64, kernel_size=3, padding=1),
nn.BatchNorm2d(64),
nn.ReLU(),
nn.MaxPool2d(2)
)
# Conv Block 3: (B, 64, 7, 7) → (B, 128, 1, 1)
self.block3 = nn.Sequential(
nn.Conv2d(64, 128, kernel_size=3, padding=1),
nn.BatchNorm2d(128),
nn.ReLU(),
nn.AdaptiveAvgPool2d((1, 1)) # global average pooling
)
# Classifier
self.classifier = nn.Linear(128, num_classes)
def forward(self, x):
x = self.block1(x) # (B, 32, 14, 14)
x = self.block2(x) # (B, 64, 7, 7)
x = self.block3(x) # (B, 128, 1, 1)
x = x.view(x.size(0), -1) # (B, 128)
x = self.classifier(x) # (B, 10)
return x
CNN Architecture Flow:
Input: (B, 1, 28, 28)
│
┌────▼────────────────────────┐
│ Conv2d(1→32, 3×3, pad=1) │
│ BatchNorm2d(32) │ Block 1
│ ReLU │
│ MaxPool2d(2) │
└────┬────────────────────────┘
│ (B, 32, 14, 14)
┌────▼────────────────────────┐
│ Conv2d(32→64, 3×3, pad=1) │
│ BatchNorm2d(64) │ Block 2
│ ReLU │
│ MaxPool2d(2) │
└────┬────────────────────────┘
│ (B, 64, 7, 7)
┌────▼────────────────────────┐
│ Conv2d(64→128, 3×3, pad=1) │
│ BatchNorm2d(128) │ Block 3
│ ReLU │
│ AdaptiveAvgPool2d(1,1) │
└────┬────────────────────────┘
│ (B, 128, 1, 1)
┌────▼────────────────────────┐
│ Flatten → (B, 128) │
│ Linear(128, 10) │ Classifier
└────┬────────────────────────┘
│ (B, 10)
▼
Output logits
6. Cheat Sheet
| Concept | Code Pattern | Ghi nhớ |
|---|---|---|
| Tạo model | class MyModel(nn.Module) | Luôn gọi super().__init__() |
| Forward pass | output = model(x) | Gọi model như function, không gọi .forward() trực tiếp |
| Training mode | model.train() | Bật dropout, batchnorm training stats |
| Eval mode | model.eval() + torch.no_grad() | Tắt dropout, dùng running stats |
| Training loop | zero_grad → forward → loss → backward → step | Thứ tự QUAN TRỌNG |
| GPU transfer | .to(device) | Cả model VÀ data |
| Classification loss | nn.CrossEntropyLoss() | Đã bao gồm softmax |
| Diffusion loss | nn.MSELoss() | Predict noise, tính MSE |
| Diffusion norm | nn.GroupNorm(G, C) | Không phụ thuộc batch size |
| Transformer optimizer | AdamW | Weight decay đúng cách |
7. Practice Questions
Các câu hỏi sau mô phỏng dạng coding assessment của NVIDIA DLI — bạn cần đọc code, tìm lỗi, và viết code hoàn chỉnh.
Q1: Fix the broken training loop
Đoạn code dưới đây có bug khiến model không hội tụ. Tìm và sửa lỗi:
for epoch in range(10):
model.train()
for images, labels in train_loader:
images, labels = images.to(device), labels.to(device)
outputs = model(images)
loss = criterion(outputs, labels)
loss.backward()
optimizer.step()
Xem đáp án Q1
Bug: Thiếu optimizer.zero_grad() trước loss.backward(). Không reset gradients sẽ khiến gradients tích lũy qua các batch, model không hội tụ hoặc hội tụ sai.
for epoch in range(10):
model.train()
for images, labels in train_loader:
images, labels = images.to(device), labels.to(device)
outputs = model(images)
loss = criterion(outputs, labels)
optimizer.zero_grad() # ← THÊM DÒNG NÀY
loss.backward()
optimizer.step()
Thứ tự chuẩn: zero_grad() → backward() → step(). Trong thực tế, zero_grad() có thể đặt trước forward cũng được, nhưng PHẢI có trước backward().
Q2: Implement a 3-layer CNN
Viết một CNN class nhận input RGB images (3, 64, 64) và output 5 classes. Yêu cầu:
- 3 convolutional blocks, mỗi block: Conv2d → BatchNorm2d → ReLU → MaxPool2d(2)
- Channels: 3 → 32 → 64 → 128
- Kết thúc bằng AdaptiveAvgPool2d + Linear
Xem đáp án Q2
class ThreeLayerCNN(nn.Module):
def __init__(self, num_classes=5):
super().__init__()
self.features = nn.Sequential(
# Block 1: (B, 3, 64, 64) → (B, 32, 32, 32)
nn.Conv2d(3, 32, kernel_size=3, padding=1),
nn.BatchNorm2d(32),
nn.ReLU(),
nn.MaxPool2d(2),
# Block 2: (B, 32, 32, 32) → (B, 64, 16, 16)
nn.Conv2d(32, 64, kernel_size=3, padding=1),
nn.BatchNorm2d(64),
nn.ReLU(),
nn.MaxPool2d(2),
# Block 3: (B, 64, 16, 16) → (B, 128, 8, 8)
nn.Conv2d(64, 128, kernel_size=3, padding=1),
nn.BatchNorm2d(128),
nn.ReLU(),
nn.MaxPool2d(2),
)
self.pool = nn.AdaptiveAvgPool2d((1, 1))
self.classifier = nn.Linear(128, num_classes)
def forward(self, x):
x = self.features(x) # (B, 128, 8, 8)
x = self.pool(x) # (B, 128, 1, 1)
x = x.view(x.size(0), -1) # (B, 128)
x = self.classifier(x) # (B, 5)
return x
# Verify
model = ThreeLayerCNN(num_classes=5)
x = torch.randn(4, 3, 64, 64)
print(model(x).shape) # torch.Size([4, 5])
Q3: Trace tensor shapes through a network
Cho model dưới đây và input shape (8, 1, 32, 32). Ghi lại shape tại mỗi bước:
model = nn.Sequential(
nn.Conv2d(1, 16, kernel_size=5, stride=2, padding=2), # Step A
nn.ReLU(),
nn.Conv2d(16, 32, kernel_size=3, stride=1, padding=0), # Step B
nn.ReLU(),
nn.AdaptiveAvgPool2d((4, 4)), # Step C
nn.Flatten(), # Step D
nn.Linear(32 * 4 * 4, 10), # Step E
)
Xem đáp án Q3
Công thức output size: H_out = (H_in + 2*padding - kernel_size) / stride + 1
Input: (8, 1, 32, 32)
Step A: Conv2d(1, 16, k=5, s=2, p=2)
H = (32 + 2*2 - 5) / 2 + 1 = 16
→ (8, 16, 16, 16)
Step B: Conv2d(16, 32, k=3, s=1, p=0)
H = (16 + 2*0 - 3) / 1 + 1 = 14
→ (8, 32, 14, 14)
Step C: AdaptiveAvgPool2d((4, 4))
→ (8, 32, 4, 4)
Step D: Flatten()
→ (8, 512) # 32 * 4 * 4 = 512
Step E: Linear(512, 10)
→ (8, 10)
Key insight: AdaptiveAvgPool2d luôn output size cố định bất kể input — rất hữu ích khi input size có thể thay đổi.
Q4: GroupNorm vs BatchNorm — khi nào dùng gì?
Bạn đang xây dựng một U-Net cho Diffusion Model. Mỗi block có Conv2d → ??? → SiLU. Bạn chọn normalization nào và tại sao? Viết code cho 1 block.
Xem đáp án Q4
Chọn GroupNorm. Lý do:
- Batch size nhỏ: Diffusion training thường dùng batch size 1-8 vì images lớn → BatchNorm statistics quá noisy
- Noise levels khác nhau: Mỗi sample trong batch có timestep (noise level) khác nhau → normalize across batch không hợp lý
- Inference consistency: GroupNorm hoạt động giống nhau ở train và inference
class DiffusionBlock(nn.Module):
def __init__(self, in_channels, out_channels, num_groups=32):
super().__init__()
self.conv = nn.Conv2d(in_channels, out_channels,
kernel_size=3, padding=1)
self.norm = nn.GroupNorm(num_groups, out_channels)
self.act = nn.SiLU() # SiLU phổ biến hơn ReLU trong diffusion
def forward(self, x):
x = self.conv(x)
x = self.norm(x)
x = self.act(x)
return x
# Ví dụ sử dụng
block = DiffusionBlock(64, 128, num_groups=32)
x = torch.randn(2, 64, 32, 32) # batch_size = 2, nhỏ!
print(block(x).shape) # (2, 128, 32, 32)
Quy tắc: Trong mọi kiến trúc Diffusion (U-Net, DiT), luôn dùng nn.GroupNorm. Activation thường là nn.SiLU() (Swish) thay vì ReLU.
Q5: Debug gradient issue — detach() vs torch.no_grad()
Đoạn code dưới đây có vấn đề gì? Output của feature_extractor không nên có gradient (freeze backbone), nhưng classifier vẫn cần train.
feature_extractor = pretrained_model.features
classifier = nn.Linear(512, 10).to(device)
optimizer = torch.optim.Adam(classifier.parameters(), lr=1e-3)
for images, labels in train_loader:
images, labels = images.to(device), labels.to(device)
# Extract features (should be frozen)
with torch.no_grad():
features = feature_extractor(images)
# Classify
outputs = classifier(features)
loss = criterion(outputs, labels)
optimizer.zero_grad()
loss.backward() # ← Có vấn đề?
optimizer.step()
Xem đáp án Q5
Vấn đề: Code này thực ra hoạt động đúng cho trường hợp này! torch.no_grad() ngăn gradient computation cho feature_extractor, và features tensor sẽ không có requires_grad. Gradient vẫn flow qua classifier bình thường.
Tuy nhiên, có 2 cách tiếp cận và bạn cần hiểu sự khác biệt:
# Cách 1: torch.no_grad() — KHÔNG tính gradient, tiết kiệm memory
with torch.no_grad():
features = feature_extractor(images)
# features.requires_grad = False
# Gradient KHÔNG flow ngược qua feature_extractor
# ✅ Dùng khi muốn freeze hoàn toàn, tiết kiệm GPU memory
# Cách 2: .detach() — tách tensor khỏi computational graph
features = feature_extractor(images).detach()
# features.requires_grad = False
# Feature extractor VẪN tính forward (tốn memory cho graph)
# nhưng gradient bị cắt tại .detach()
# ⚠️ Kém hiệu quả hơn vì vẫn build graph rồi mới cắt
# Cách 3: Freeze parameters — approach phổ biến nhất
for param in feature_extractor.parameters():
param.requires_grad = False
# ✅ Rõ ràng nhất, thường dùng trong fine-tuning
| Approach | Gradient flow | Memory | Khi nào dùng |
|---|---|---|---|
torch.no_grad() | Không tính graph | Tiết kiệm nhất | Inference, frozen features |
.detach() | Cắt tại điểm detach | Tốn hơn | Khi cần partial gradient flow |
| Freeze params | Không update weights | Vẫn build graph | Fine-tuning rõ ràng |
8. Kết luận
Bài 1 đã trang bị cho bạn toàn bộ nền tảng PyTorch cần thiết cho khóa NVIDIA DLI Generative AI. Hãy chắc chắn bạn có thể:
- Viết training loop hoàn chỉnh mà không cần nhìn tài liệu
- Tạo nn.Module class với
__init__vàforward - Tính tensor shapes qua từng layer
- Phân biệt GroupNorm vs BatchNorm — đặc biệt quan trọng cho Diffusion Models
- Debug các lỗi phổ biến: thiếu
zero_grad(), device mismatch, gradient issues
Bài tiếp theo: Bài 2 — Sequence Models & Attention Mechanism — nền tảng cho Transformers và LLMs.