Chuyển đến nội dung chính

Bài 1: PyTorch & Neural Network Fundamentals

PyTorch tensors, autograd, nn.Module. Build neural network from scratch. Training loop, loss functions, optimizers. GPU acceleration basics. CNN architecture, pooling, batch normalization.

1. Giới thiệu

Bài học đầu tiên trong series luyện thi NVIDIA DLI — Generative AI sẽ trang bị cho bạn nền tảng PyTorch vững chắc. Đây là framework chính được sử dụng trong toàn bộ khóa học DLI, từ Diffusion Models đến Large Language Models (LLMs).

Trong bài assessment của NVIDIA DLI, bạn sẽ phải viết code trực tiếp — không phải trắc nghiệm. Vì vậy, nắm vững các pattern cơ bản của PyTorch là điều kiện tiên quyết.

Exam tip: NVIDIA DLI assessment yêu cầu bạn viết và debug code PyTorch trực tiếp. Hãy chắc chắn bạn có thể viết training loop, nn.Module, và thao tác tensor mà không cần nhìn tài liệu.

Kiến trúc Deep Neural Network — Input Layer, Hidden Layers, Output Layer, Backpropagation
Kiến trúc Deep Neural Network — Input Layer, Hidden Layers, Output Layer, Backpropagation

2. PyTorch Tensors & Autograd

2.1 Tensor Basics

Tensor là cấu trúc dữ liệu cốt lõi của PyTorch — tương tự NumPy array nhưng có thể chạy trên GPU và hỗ trợ automatic differentiation.

import torch

# Tạo tensor từ list
x = torch.tensor([1.0, 2.0, 3.0])

# Tạo tensor với shape cụ thể
zeros = torch.zeros(3, 4)          # shape: (3, 4)
ones = torch.ones(2, 3, 4)         # shape: (2, 3, 4)
rand = torch.randn(64, 3, 32, 32)  # batch of 64 RGB 32x32 images

# Kiểm tra shape và dtype
print(rand.shape)   # torch.Size([64, 3, 32, 32])
print(rand.dtype)   # torch.float32
print(rand.device)  # cpu

2.2 Tensor Operations & Broadcasting

PyTorch hỗ trợ broadcasting tương tự NumPy — cho phép thực hiện phép tính giữa tensors có shape khác nhau.

# Reshape operations
x = torch.randn(2, 3, 4)
y = x.view(2, 12)        # reshape thành (2, 12)
z = x.permute(0, 2, 1)   # swap dims: (2, 4, 3)
w = x.unsqueeze(0)        # thêm dim: (1, 2, 3, 4)

# Matrix multiplication
a = torch.randn(3, 4)
b = torch.randn(4, 5)
c = a @ b                 # shape: (3, 5)
# hoặc: c = torch.matmul(a, b)

# Broadcasting
x = torch.randn(64, 256)  # (batch, features)
bias = torch.randn(256)   # (features,)
result = x + bias          # bias được broadcast: (64, 256)
OperationSyntaxGhi chú
Reshapex.view() / x.reshape()view yêu cầu contiguous memory
Transposex.permute() / x.Tpermute linh hoạt hơn cho nhiều dims
Add dimx.unsqueeze(dim)Thường dùng để chuẩn bị broadcasting
Remove dimx.squeeze(dim)Xóa dim có size = 1
Matrix mula @ bEquivalent to torch.matmul
Concattorch.cat([a, b], dim=0)Nối theo dim chỉ định

2.3 Autograd — Automatic Differentiation

Autograd là hệ thống tự động tính gradient của PyTorch. Khi bạn set requires_grad=True, PyTorch sẽ theo dõi mọi phép tính trên tensor đó và xây dựng computational graph.

# Autograd cơ bản
x = torch.tensor([2.0, 3.0], requires_grad=True)
y = x ** 2 + 3 * x       # y = x² + 3x
loss = y.sum()            # scalar output
loss.backward()           # tính gradient

print(x.grad)  # dy/dx = 2x + 3 → tensor([7., 9.])
Computational Graph:

  x (requires_grad=True)
  │
  ├──→ x ** 2 ──→ + ──→ y ──→ sum() ──→ loss
  │                ↑                        │
  └──→ 3 * x ─────┘                   backward()
                                            │
                                       x.grad = 2x + 3

Exam tip: Trong DLI assessment, bạn có thể gặp lỗi "trying to backward through the graph a second time". Giải pháp: dùng loss.backward(retain_graph=True) hoặc tính lại forward pass. Đây là lỗi rất phổ biến khi làm bài.

3. nn.Module & Building Networks

3.1 nn.Module Pattern

Mọi neural network trong PyTorch đều kế thừa từ nn.Module. Đây là pattern bắt buộc phải thuộc lòng:

import torch.nn as nn

class SimpleNet(nn.Module):
    def __init__(self, input_dim, hidden_dim, output_dim):
        super().__init__()  # PHẢI gọi super().__init__()
        self.fc1 = nn.Linear(input_dim, hidden_dim)
        self.relu = nn.ReLU()
        self.fc2 = nn.Linear(hidden_dim, output_dim)

    def forward(self, x):
        x = self.fc1(x)
        x = self.relu(x)
        x = self.fc2(x)
        return x

# Sử dụng
model = SimpleNet(784, 256, 10)
x = torch.randn(32, 784)  # batch of 32
output = model(x)          # shape: (32, 10)

3.2 Common Layers

LayerCông dụngParams chính
nn.Linear(in, out)Fully connected layerin_features, out_features
nn.Conv2d(in_ch, out_ch, k)2D convolutionin_channels, out_channels, kernel_size
nn.BatchNorm2d(ch)Batch normalizationnum_features
nn.GroupNorm(g, ch)Group normalizationnum_groups, num_channels
nn.ReLU()Activation functioninplace (optional)
nn.Dropout(p)Regularizationp = drop probability
nn.Embedding(V, D)Token embeddingnum_embeddings, embedding_dim

3.3 nn.Sequential — Quick Models

Với mô hình đơn giản, bạn có thể dùng nn.Sequential thay vì tạo class:

# Cách nhanh với nn.Sequential
model = nn.Sequential(
    nn.Linear(784, 256),
    nn.ReLU(),
    nn.Dropout(0.2),
    nn.Linear(256, 128),
    nn.ReLU(),
    nn.Linear(128, 10)
)

# Kiểm tra model
print(model)
# Đếm parameters
total_params = sum(p.numel() for p in model.parameters())
print(f"Total params: {total_params:,}")

3.4 Code: MLP cho MNIST

import torch
import torch.nn as nn
from torchvision import datasets, transforms

# Data
transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.1307,), (0.3081,))
])
train_data = datasets.MNIST('./data', train=True, download=True,
                            transform=transform)
train_loader = torch.utils.data.DataLoader(train_data, batch_size=64,
                                           shuffle=True)

# Model
class MNISTClassifier(nn.Module):
    def __init__(self):
        super().__init__()
        self.flatten = nn.Flatten()
        self.layers = nn.Sequential(
            nn.Linear(28 * 28, 512),
            nn.ReLU(),
            nn.Dropout(0.2),
            nn.Linear(512, 256),
            nn.ReLU(),
            nn.Dropout(0.2),
            nn.Linear(256, 10)
        )

    def forward(self, x):
        x = self.flatten(x)   # (B, 1, 28, 28) → (B, 784)
        return self.layers(x)  # (B, 784) → (B, 10)

4. Training Loop Pattern

4.1 Loss Functions

Loss FunctionDùng khiInput shape
nn.CrossEntropyLoss()Multi-class classificationlogits (B, C), labels (B,)
nn.MSELoss()Regression, diffusion noise prediction(B, *) vs (B, *)
nn.BCEWithLogitsLoss()Binary / multi-label classificationlogits (B, C), labels (B, C)
nn.L1Loss()Regression (robust to outliers)(B, *) vs (B, *)

Exam tip: nn.CrossEntropyLoss đã bao gồm softmax bên trong — KHÔNG cần thêm softmax ở output layer. Đây là lỗi mà nhiều người mắc phải trong assessment. nn.MSELoss sẽ rất quan trọng khi bạn học Diffusion Models (predict noise).

4.2 Optimizers

# SGD — cơ bản, can thiệp learning rate thủ công
optimizer = torch.optim.SGD(model.parameters(), lr=0.01, momentum=0.9)

# Adam — phổ biến nhất, adaptive learning rate
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)

# AdamW — Adam + weight decay đúng cách (dùng cho Transformers)
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-4, weight_decay=0.01)
OptimizerKhi nào dùngĐặc điểm
SGDCNNs, khi cần kiểm soát tỉ mỉCần tune lr cẩn thận, thêm momentum
AdamMặc định cho hầu hết tasksHội tụ nhanh, ít cần tune
AdamWTransformers, LLMs, DiffusionWeight decay tách riêng, chuẩn hơn Adam

4.3 Full Training Loop

Đây là pattern quan trọng nhất — bạn phải viết được training loop hoàn chỉnh trong assessment:

import torch
import torch.nn as nn

# Setup
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model = MNISTClassifier().to(device)
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)

# Training loop
num_epochs = 10
for epoch in range(num_epochs):
    model.train()  # BẬT training mode (dropout, batchnorm active)
    total_loss = 0

    for batch_idx, (images, labels) in enumerate(train_loader):
        images, labels = images.to(device), labels.to(device)

        # Forward pass
        outputs = model(images)
        loss = criterion(outputs, labels)

        # Backward pass
        optimizer.zero_grad()  # QUAN TRỌNG: reset gradients
        loss.backward()        # tính gradients
        optimizer.step()       # update weights

        total_loss += loss.item()

    avg_loss = total_loss / len(train_loader)
    print(f"Epoch [{epoch+1}/{num_epochs}], Loss: {avg_loss:.4f}")

# Evaluation
model.eval()  # TẮT dropout, batchnorm dùng running stats
with torch.no_grad():  # KHÔNG tính gradients → tiết kiệm memory
    correct = 0
    total = 0
    for images, labels in test_loader:
        images, labels = images.to(device), labels.to(device)
        outputs = model(images)
        _, predicted = torch.max(outputs, 1)
        total += labels.size(0)
        correct += (predicted == labels).sum().item()

    print(f"Accuracy: {100 * correct / total:.2f}%")
Training Loop Flow:

  ┌─────────────────────────────────────────────┐
  │              FOR EACH EPOCH                  │
  │  ┌────────────────────────────────────────┐  │
  │  │         FOR EACH BATCH                 │  │
  │  │                                        │  │
  │  │  1. images, labels = batch.to(device)  │  │
  │  │              │                         │  │
  │  │  2. outputs = model(images)   FORWARD  │  │
  │  │              │                         │  │
  │  │  3. loss = criterion(outputs, labels)  │  │
  │  │              │                         │  │
  │  │  4. optimizer.zero_grad()     RESET    │  │
  │  │              │                         │  │
  │  │  5. loss.backward()          BACKWARD  │  │
  │  │              │                         │  │
  │  │  6. optimizer.step()          UPDATE   │  │
  │  │              │                         │  │
  │  └──────────────┼─────────────────────────┘  │
  │                 ▼                             │
  │         Next Epoch                            │
  └─────────────────────────────────────────────┘

Exam tip: Thứ tự zero_grad() → backward() → step() là BẮT BUỘC. Quên zero_grad() sẽ khiến gradients bị tích lũy qua các batch — đây là bug #1 trong DLI assessment. Luôn nhớ model.train() trước training và model.eval() + torch.no_grad() trước evaluation.

4.4 GPU Acceleration

# Kiểm tra GPU
print(torch.cuda.is_available())        # True/False
print(torch.cuda.device_count())        # Số GPU
print(torch.cuda.get_device_name(0))    # Tên GPU

# Di chuyển model và data lên GPU
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
model = model.to(device)

# TRONG training loop — data cũng phải lên cùng device
images = images.to(device)
labels = labels.to(device)

Exam tip: Lỗi phổ biến: model trên GPU nhưng data vẫn trên CPU (hoặc ngược lại). PyTorch sẽ báo lỗi "Expected all tensors to be on the same device". Luôn đảm bảo model và data cùng device.

5. CNN Architecture

5.1 Convolutional Layers

Convolution trượt một kernel (filter) qua input image để trích xuất features. Mỗi kernel phát hiện một pattern cụ thể (edges, textures, shapes).

# Conv2d cơ bản
conv = nn.Conv2d(
    in_channels=3,    # RGB input
    out_channels=32,  # 32 filters → 32 feature maps
    kernel_size=3,    # 3×3 kernel
    stride=1,         # bước nhảy
    padding=1          # zero-padding để giữ spatial size
)

# Output shape calculation:
# H_out = (H_in + 2*padding - kernel_size) / stride + 1
# Ví dụ: (32 + 2*1 - 3) / 1 + 1 = 32 (giữ nguyên size)
Convolution Operation:

Input (3 channels)          Kernel (3×3)         Output (1 feature map)
┌─────────────┐            ┌───────┐              ┌──────────┐
│ ■ ■ ■ ■ ■ ■│   ×        │ w w w │     =        │ ○ ○ ○ ○  │
│ ■ ■ ■ ■ ■ ■│            │ w w w │              │ ○ ○ ○ ○  │
│ ■ ■ ■ ■ ■ ■│            │ w w w │              │ ○ ○ ○ ○  │
│ ■ ■ ■ ■ ■ ■│            └───────┘              │ ○ ○ ○ ○  │
│ ■ ■ ■ ■ ■ ■│                                   └──────────┘
│ ■ ■ ■ ■ ■ ■│         32 kernels → 32 feature maps
└─────────────┘

Shape flow: (B, 3, 32, 32) → Conv2d(3, 32, 3, padding=1) → (B, 32, 32, 32)

5.2 Pooling Layers

Pooling giảm spatial dimensions, giúp giảm computation và tăng receptive field:

PoolingCách hoạt độngKhi nào dùng
nn.MaxPool2d(2)Lấy giá trị max trong mỗi 2×2 windowFeature detection, CNNs thông thường
nn.AvgPool2d(2)Lấy trung bình trong mỗi 2×2 windowSmoother features
nn.AdaptiveAvgPool2d((1,1))Pool về size cố định bất kể inputTrước fully-connected layer

5.3 Batch Normalization vs Group Normalization

Đây là kiến thức cực kỳ quan trọng cho phần Diffusion Models ở các bài sau.

Thuộc tínhBatchNormGroupNorm
Normalize theoBatch dimension (N)Channel groups (C/G)
Phụ thuộc batch sizeCó — batch nhỏ thì noisyKhông — hoạt động tốt mọi batch size
Training vs InferenceKhác nhau (running stats)Giống nhau
Phổ biến trongCNNs truyền thống (ResNet)Diffusion Models, nhóm nhỏ
Syntaxnn.BatchNorm2d(C)nn.GroupNorm(G, C)
# BatchNorm — normalize across batch
bn = nn.BatchNorm2d(64)         # 64 channels

# GroupNorm — normalize within groups of channels
gn = nn.GroupNorm(
    num_groups=32,    # chia 64 channels thành 32 groups (2 ch/group)
    num_channels=64
)

# Cả hai nhận input shape: (B, C, H, W)
x = torch.randn(8, 64, 16, 16)
print(bn(x).shape)  # (8, 64, 16, 16)
print(gn(x).shape)  # (8, 64, 16, 16)
BatchNorm vs GroupNorm:

BatchNorm: normalize theo ↓ (batch axis)   GroupNorm: normalize theo → (channel groups)
┌────────────────────────┐                  ┌────────────────────────┐
│  Sample 1: [c1 c2 c3 c4]│                  │  Sample 1: [c1 c2│c3 c4]│
│  Sample 2: [c1 c2 c3 c4]│  ← normalize    │            group1│group2 │ ← normalize
│  Sample 3: [c1 c2 c3 c4]│    mỗi column   │                  │       │   mỗi group
│  Sample 4: [c1 c2 c3 c4]│                  │  Sample 2: [c1 c2│c3 c4]│
└────────────────────────┘                  └────────────────────────┘

→ Diffusion Models dùng GroupNorm vì batch_size thường nhỏ
  và noise level thay đổi → BatchNorm statistics không ổn định

Exam tip: Khi xây dựng U-Net cho Diffusion Models (bài sau), bạn sẽ luôn dùng GroupNorm thay vì BatchNorm. Lý do: diffusion training thường dùng batch size nhỏ, và mỗi sample có noise level khác nhau → BatchNorm statistics bị noisy. Nhớ quy tắc: Diffusion = GroupNorm.

5.4 Code: Simple CNN cho Image Classification

class SimpleCNN(nn.Module):
    def __init__(self, num_classes=10):
        super().__init__()

        # Conv Block 1: (B, 1, 28, 28) → (B, 32, 14, 14)
        self.block1 = nn.Sequential(
            nn.Conv2d(1, 32, kernel_size=3, padding=1),
            nn.BatchNorm2d(32),
            nn.ReLU(),
            nn.MaxPool2d(2)
        )

        # Conv Block 2: (B, 32, 14, 14) → (B, 64, 7, 7)
        self.block2 = nn.Sequential(
            nn.Conv2d(32, 64, kernel_size=3, padding=1),
            nn.BatchNorm2d(64),
            nn.ReLU(),
            nn.MaxPool2d(2)
        )

        # Conv Block 3: (B, 64, 7, 7) → (B, 128, 1, 1)
        self.block3 = nn.Sequential(
            nn.Conv2d(64, 128, kernel_size=3, padding=1),
            nn.BatchNorm2d(128),
            nn.ReLU(),
            nn.AdaptiveAvgPool2d((1, 1))  # global average pooling
        )

        # Classifier
        self.classifier = nn.Linear(128, num_classes)

    def forward(self, x):
        x = self.block1(x)     # (B, 32, 14, 14)
        x = self.block2(x)     # (B, 64, 7, 7)
        x = self.block3(x)     # (B, 128, 1, 1)
        x = x.view(x.size(0), -1)  # (B, 128)
        x = self.classifier(x)     # (B, 10)
        return x
CNN Architecture Flow:

Input: (B, 1, 28, 28)
         │
    ┌────▼────────────────────────┐
    │ Conv2d(1→32, 3×3, pad=1)   │
    │ BatchNorm2d(32)             │  Block 1
    │ ReLU                        │
    │ MaxPool2d(2)                │
    └────┬────────────────────────┘
         │ (B, 32, 14, 14)
    ┌────▼────────────────────────┐
    │ Conv2d(32→64, 3×3, pad=1)  │
    │ BatchNorm2d(64)             │  Block 2
    │ ReLU                        │
    │ MaxPool2d(2)                │
    └────┬────────────────────────┘
         │ (B, 64, 7, 7)
    ┌────▼────────────────────────┐
    │ Conv2d(64→128, 3×3, pad=1) │
    │ BatchNorm2d(128)            │  Block 3
    │ ReLU                        │
    │ AdaptiveAvgPool2d(1,1)      │
    └────┬────────────────────────┘
         │ (B, 128, 1, 1)
    ┌────▼────────────────────────┐
    │ Flatten → (B, 128)          │
    │ Linear(128, 10)             │  Classifier
    └────┬────────────────────────┘
         │ (B, 10)
         ▼
      Output logits

6. Cheat Sheet

ConceptCode PatternGhi nhớ
Tạo modelclass MyModel(nn.Module)Luôn gọi super().__init__()
Forward passoutput = model(x)Gọi model như function, không gọi .forward() trực tiếp
Training modemodel.train()Bật dropout, batchnorm training stats
Eval modemodel.eval() + torch.no_grad()Tắt dropout, dùng running stats
Training loopzero_grad → forward → loss → backward → stepThứ tự QUAN TRỌNG
GPU transfer.to(device)Cả model VÀ data
Classification lossnn.CrossEntropyLoss()Đã bao gồm softmax
Diffusion lossnn.MSELoss()Predict noise, tính MSE
Diffusion normnn.GroupNorm(G, C)Không phụ thuộc batch size
Transformer optimizerAdamWWeight decay đúng cách

7. Practice Questions

Các câu hỏi sau mô phỏng dạng coding assessment của NVIDIA DLI — bạn cần đọc code, tìm lỗi, và viết code hoàn chỉnh.

Q1: Fix the broken training loop

Đoạn code dưới đây có bug khiến model không hội tụ. Tìm và sửa lỗi:

for epoch in range(10):
    model.train()
    for images, labels in train_loader:
        images, labels = images.to(device), labels.to(device)
        outputs = model(images)
        loss = criterion(outputs, labels)
        loss.backward()
        optimizer.step()
Xem đáp án Q1

Bug: Thiếu optimizer.zero_grad() trước loss.backward(). Không reset gradients sẽ khiến gradients tích lũy qua các batch, model không hội tụ hoặc hội tụ sai.

for epoch in range(10):
    model.train()
    for images, labels in train_loader:
        images, labels = images.to(device), labels.to(device)
        outputs = model(images)
        loss = criterion(outputs, labels)

        optimizer.zero_grad()  # ← THÊM DÒNG NÀY
        loss.backward()
        optimizer.step()

Thứ tự chuẩn: zero_grad() → backward() → step(). Trong thực tế, zero_grad() có thể đặt trước forward cũng được, nhưng PHẢI có trước backward().

Q2: Implement a 3-layer CNN

Viết một CNN class nhận input RGB images (3, 64, 64) và output 5 classes. Yêu cầu:

  • 3 convolutional blocks, mỗi block: Conv2d → BatchNorm2d → ReLU → MaxPool2d(2)
  • Channels: 3 → 32 → 64 → 128
  • Kết thúc bằng AdaptiveAvgPool2d + Linear
Xem đáp án Q2
class ThreeLayerCNN(nn.Module):
    def __init__(self, num_classes=5):
        super().__init__()
        self.features = nn.Sequential(
            # Block 1: (B, 3, 64, 64) → (B, 32, 32, 32)
            nn.Conv2d(3, 32, kernel_size=3, padding=1),
            nn.BatchNorm2d(32),
            nn.ReLU(),
            nn.MaxPool2d(2),

            # Block 2: (B, 32, 32, 32) → (B, 64, 16, 16)
            nn.Conv2d(32, 64, kernel_size=3, padding=1),
            nn.BatchNorm2d(64),
            nn.ReLU(),
            nn.MaxPool2d(2),

            # Block 3: (B, 64, 16, 16) → (B, 128, 8, 8)
            nn.Conv2d(64, 128, kernel_size=3, padding=1),
            nn.BatchNorm2d(128),
            nn.ReLU(),
            nn.MaxPool2d(2),
        )
        self.pool = nn.AdaptiveAvgPool2d((1, 1))
        self.classifier = nn.Linear(128, num_classes)

    def forward(self, x):
        x = self.features(x)        # (B, 128, 8, 8)
        x = self.pool(x)            # (B, 128, 1, 1)
        x = x.view(x.size(0), -1)   # (B, 128)
        x = self.classifier(x)      # (B, 5)
        return x

# Verify
model = ThreeLayerCNN(num_classes=5)
x = torch.randn(4, 3, 64, 64)
print(model(x).shape)  # torch.Size([4, 5])

Q3: Trace tensor shapes through a network

Cho model dưới đây và input shape (8, 1, 32, 32). Ghi lại shape tại mỗi bước:

model = nn.Sequential(
    nn.Conv2d(1, 16, kernel_size=5, stride=2, padding=2),  # Step A
    nn.ReLU(),
    nn.Conv2d(16, 32, kernel_size=3, stride=1, padding=0), # Step B
    nn.ReLU(),
    nn.AdaptiveAvgPool2d((4, 4)),                           # Step C
    nn.Flatten(),                                            # Step D
    nn.Linear(32 * 4 * 4, 10),                              # Step E
)
Xem đáp án Q3

Công thức output size: H_out = (H_in + 2*padding - kernel_size) / stride + 1

Input:  (8, 1, 32, 32)

Step A: Conv2d(1, 16, k=5, s=2, p=2)
        H = (32 + 2*2 - 5) / 2 + 1 = 16
        → (8, 16, 16, 16)

Step B: Conv2d(16, 32, k=3, s=1, p=0)
        H = (16 + 2*0 - 3) / 1 + 1 = 14
        → (8, 32, 14, 14)

Step C: AdaptiveAvgPool2d((4, 4))
        → (8, 32, 4, 4)

Step D: Flatten()
        → (8, 512)       # 32 * 4 * 4 = 512

Step E: Linear(512, 10)
        → (8, 10)

Key insight: AdaptiveAvgPool2d luôn output size cố định bất kể input — rất hữu ích khi input size có thể thay đổi.

Q4: GroupNorm vs BatchNorm — khi nào dùng gì?

Bạn đang xây dựng một U-Net cho Diffusion Model. Mỗi block có Conv2d → ??? → SiLU. Bạn chọn normalization nào và tại sao? Viết code cho 1 block.

Xem đáp án Q4

Chọn GroupNorm. Lý do:

  1. Batch size nhỏ: Diffusion training thường dùng batch size 1-8 vì images lớn → BatchNorm statistics quá noisy
  2. Noise levels khác nhau: Mỗi sample trong batch có timestep (noise level) khác nhau → normalize across batch không hợp lý
  3. Inference consistency: GroupNorm hoạt động giống nhau ở train và inference
class DiffusionBlock(nn.Module):
    def __init__(self, in_channels, out_channels, num_groups=32):
        super().__init__()
        self.conv = nn.Conv2d(in_channels, out_channels,
                              kernel_size=3, padding=1)
        self.norm = nn.GroupNorm(num_groups, out_channels)
        self.act = nn.SiLU()  # SiLU phổ biến hơn ReLU trong diffusion

    def forward(self, x):
        x = self.conv(x)
        x = self.norm(x)
        x = self.act(x)
        return x

# Ví dụ sử dụng
block = DiffusionBlock(64, 128, num_groups=32)
x = torch.randn(2, 64, 32, 32)  # batch_size = 2, nhỏ!
print(block(x).shape)  # (2, 128, 32, 32)

Quy tắc: Trong mọi kiến trúc Diffusion (U-Net, DiT), luôn dùng nn.GroupNorm. Activation thường là nn.SiLU() (Swish) thay vì ReLU.

Q5: Debug gradient issue — detach() vs torch.no_grad()

Đoạn code dưới đây có vấn đề gì? Output của feature_extractor không nên có gradient (freeze backbone), nhưng classifier vẫn cần train.

feature_extractor = pretrained_model.features
classifier = nn.Linear(512, 10).to(device)
optimizer = torch.optim.Adam(classifier.parameters(), lr=1e-3)

for images, labels in train_loader:
    images, labels = images.to(device), labels.to(device)

    # Extract features (should be frozen)
    with torch.no_grad():
        features = feature_extractor(images)

    # Classify
    outputs = classifier(features)
    loss = criterion(outputs, labels)

    optimizer.zero_grad()
    loss.backward()    # ← Có vấn đề?
    optimizer.step()
Xem đáp án Q5

Vấn đề: Code này thực ra hoạt động đúng cho trường hợp này! torch.no_grad() ngăn gradient computation cho feature_extractor, và features tensor sẽ không có requires_grad. Gradient vẫn flow qua classifier bình thường.

Tuy nhiên, có 2 cách tiếp cận và bạn cần hiểu sự khác biệt:

# Cách 1: torch.no_grad() — KHÔNG tính gradient, tiết kiệm memory
with torch.no_grad():
    features = feature_extractor(images)
# features.requires_grad = False
# Gradient KHÔNG flow ngược qua feature_extractor
# ✅ Dùng khi muốn freeze hoàn toàn, tiết kiệm GPU memory

# Cách 2: .detach() — tách tensor khỏi computational graph
features = feature_extractor(images).detach()
# features.requires_grad = False
# Feature extractor VẪN tính forward (tốn memory cho graph)
# nhưng gradient bị cắt tại .detach()
# ⚠️ Kém hiệu quả hơn vì vẫn build graph rồi mới cắt

# Cách 3: Freeze parameters — approach phổ biến nhất
for param in feature_extractor.parameters():
    param.requires_grad = False
# ✅ Rõ ràng nhất, thường dùng trong fine-tuning
ApproachGradient flowMemoryKhi nào dùng
torch.no_grad()Không tính graphTiết kiệm nhấtInference, frozen features
.detach()Cắt tại điểm detachTốn hơnKhi cần partial gradient flow
Freeze paramsKhông update weightsVẫn build graphFine-tuning rõ ràng

8. Kết luận

Bài 1 đã trang bị cho bạn toàn bộ nền tảng PyTorch cần thiết cho khóa NVIDIA DLI Generative AI. Hãy chắc chắn bạn có thể:

  • Viết training loop hoàn chỉnh mà không cần nhìn tài liệu
  • Tạo nn.Module class với __init__ và forward
  • Tính tensor shapes qua từng layer
  • Phân biệt GroupNorm vs BatchNorm — đặc biệt quan trọng cho Diffusion Models
  • Debug các lỗi phổ biến: thiếu zero_grad(), device mismatch, gradient issues

Bài tiếp theo: Bài 2 — Sequence Models & Attention Mechanism — nền tảng cho Transformers và LLMs.