Chuyển đến nội dung chính

Lesson 3: Deep Learning & Neural Networks Essentials

Neural networks fundamentals. PyTorch basics. CNN, RNN overview. Training loops, loss functions, optimizers. GPU training. Transfer learning concepts. Model serialization.

If traditional ML is "you choose features yourself", then Deep Learning is "the model learns features from raw data". Neural networks are the foundation for all modern LLMs. This lesson will give you a solid PyTorch foundation: from tensor to training loop, CNN/RNN overview to transfer learning — prepare for Lesson 4 (NLP & Transformers).

1. From traditional ML to Deep Learning

Traditional ML works well with structured data and manual features. When data is images, text, audio — you need Deep Learning.

ML truyền thống:  Raw Data ──▶ [Feature Engineering] ──▶ Model ──▶ Output
                                 (con người thiết kế)
Deep Learning:    Raw Data ──▶ [Neural Network] ──▶ Output
                                (tự học features qua nhiều layers)
CriteriaTraditional MLDeep Learning
Data typeStructured (tabular)Unstructured (image, text, audio)
Data size100 — 10K samples10K — millions
Feature engineeringManually, requires domain knowledgeAutomatically pass layers
InterpretabilityHigh (SHAP, decision tree)Low (black box)
ComputeEnough CPURequires GPU/TPU
Toolsscikit-learn, XGBoostPyTorch, TensorFlow

Practical tip: With tabular data, XGBoost still beats DL in most cases. Only use neural networks when data is unstructured or needs end-to-end learning.

1.1. Environment setup

pip install torch torchvision torchaudio
# CUDA 12.x: thêm --index-url https://download.pytorch.org/whl/cu121
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader
import torchvision.transforms as transforms

device = torch.device(
    "cuda" if torch.cuda.is_available()
    else "mps" if torch.backends.mps.is_available()
    else "cpu"
)
print(f"Using device: {device}")

2. Neural Network Fundamentals

2.1. Neurons, Layers, Architecture

A neuron: receives inputs → multiplies weights → adds bias → through activation function.

x₁ ──w₁──┐
          ├──▶ Σ(wᵢxᵢ + b) ──▶ f(z) ──▶ output
x₂ ──w₂──┤      (linear)     (activation)
          │
x₃ ──w₃──┘

Neural network = stack of multiple layers. Each layer has many neurons. Input → Hidden layers → Output.

2.2. Activation Functions

Activation creates non-linearity — without it, the network is just matrix multiplication.

ActivationRecipeUse whenNotes
ReLUmax(0, x)Hidden layers (default)Fast, can kill neurons
GELUx·Φ(x)Transformer modelsUsed in BERT/GPT
Sigmoid1/(1+e⁻ˣ)Binary outputOutput ∈ (0,1)
Softmaxeˣⁱ/ΣeˣʲMulti-class outputProbability distribution
import torch.nn.functional as F
x = torch.tensor([-2.0, -1.0, 0.0, 1.0, 2.0])
print(f"ReLU:    {F.relu(x)}")         # [0, 0, 0, 1, 2]
print(f"Sigmoid: {torch.sigmoid(x)}")  # [0.12, 0.27, 0.5, 0.73, 0.88]

Practical tip: 95% of cases → ReLU for hidden layers. GELU for Transformer. Sigmoid for binary output, Softmax for multi-class.

3. PyTorch from Scratch

3.1. Tensors

Tensor = multidimensional array, similar to NumPy but runs on GPU and supports autograd.

a = torch.tensor([1.0, 2.0, 3.0])       # từ list
b = torch.randn(3, 4)                    # random normal 3×4
c = torch.from_numpy(np.array([1, 2]))   # từ NumPy (zero-copy)

x, y = torch.randn(3, 4), torch.randn(3, 4)
z = x + y              # element-wise
m = x @ y.T            # matrix multiply
s = x.sum(dim=1)       # sum theo axis
t_gpu = x.to(device)   # chuyển lên GPU

3.2. Autograd

Autograd automatically calculates the gradient for backpropagation:

x = torch.tensor([2.0, 3.0], requires_grad=True)
y = x ** 2 + 3 * x
loss = y.sum()
loss.backward()
print(x.grad)  # tensor([7., 9.])  — dy/dx = 2x + 3

3.3. nn.Module

All legacy PyTorch models nn.Module:

class SimpleClassifier(nn.Module):
    def __init__(self, input_dim, hidden_dim, output_dim):
        super().__init__()
        self.network = nn.Sequential(
            nn.Linear(input_dim, hidden_dim),
            nn.ReLU(),
            nn.Dropout(0.3),
            nn.Linear(hidden_dim, hidden_dim // 2),
            nn.ReLU(),
            nn.Dropout(0.3),
            nn.Linear(hidden_dim // 2, output_dim),
        )

    def forward(self, x):
        return self.network(x)

model = SimpleClassifier(784, 256, 10)
total_params = sum(p.numel() for p in model.parameters())
print(f"Total parameters: {total_params:,}")  # 235,146

4. Detailed Training Loop

Training loop — the heart of deep learning — always boils down to 5 steps:

for each epoch:
  for each batch:
    1. Forward:  output = model(data)
    2. Loss:     loss = criterion(output, target)
    3. Zero:     optimizer.zero_grad()
    4. Backward: loss.backward()
    5. Update:   optimizer.step()

4.1. MNIST Example — Full Loop

from torchvision import datasets, transforms

transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.1307,), (0.3081,)),
])
train_dataset = datasets.MNIST("./data", train=True, download=True, transform=transform)
test_dataset = datasets.MNIST("./data", train=False, transform=transform)
train_loader = DataLoader(train_dataset, batch_size=64, shuffle=True)
test_loader = DataLoader(test_dataset, batch_size=1000)

class MNISTNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.net = nn.Sequential(
            nn.Flatten(),
            nn.Linear(28*28, 512), nn.ReLU(), nn.Dropout(0.2),
            nn.Linear(512, 256), nn.ReLU(), nn.Dropout(0.2),
            nn.Linear(256, 10),
        )
    def forward(self, x):
        return self.net(x)

model = MNISTNet().to(device)
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=1e-3)

def train_one_epoch(model, loader, criterion, optimizer):
    model.train()
    total_loss, correct, total = 0, 0, 0
    for data, target in loader:
        data, target = data.to(device), target.to(device)
        output = model(data)
        loss = criterion(output, target)
        optimizer.zero_grad()   # QUAN TRỌNG — quên = gradients accumulate
        loss.backward()
        optimizer.step()
        total_loss += loss.item()
        correct += output.argmax(1).eq(target).sum().item()
        total += target.size(0)
    return total_loss / len(loader), 100.0 * correct / total

def evaluate(model, loader, criterion):
    model.eval()
    total_loss, correct, total = 0, 0, 0
    with torch.no_grad():
        for data, target in loader:
            data, target = data.to(device), target.to(device)
            output = model(data)
            total_loss += criterion(output, target).item()
            correct += output.argmax(1).eq(target).sum().item()
            total += target.size(0)
    return total_loss / len(loader), 100.0 * correct / total

for epoch in range(1, 11):
    t_loss, t_acc = train_one_epoch(model, train_loader, criterion, optimizer)
    v_loss, v_acc = evaluate(model, test_loader, criterion)
    print(f"Epoch {epoch:2d} | Train {t_loss:.4f} ({t_acc:.1f}%) | Val {v_loss:.4f} ({v_acc:.1f}%)")

Practical tip: val_loss increase + train_loss reduce → overfitting. Reduce model size, increase dropout, or use early stopping.

5. Loss Functions — Choose the right loss

LossUsed forInput → TargetPyTorch
CrossEntropyLossMulti-classRaw logits → Class indexnn.CrossEntropyLoss()
BCEWithLogitsLossBinary / Multi-labelRaw logits → 0/1nn.BCEWithLogitsLoss()
MSELossRegressionPredicted → Actualnn.MSELoss()
HuberLossRegression (robust)Predicted → Actualnn.HuberLoss()
# CrossEntropy — KHÔNG cần softmax trước đó (đã tích hợp)
logits = torch.tensor([[2.0, 1.0, 0.1]])
target = torch.tensor([0])
print(nn.CrossEntropyLoss()(logits, target))  # 0.4170

# BCEWithLogits — binary/multi-label
logits_bin = torch.tensor([0.5, -1.0, 2.0])
target_bin = torch.tensor([1.0, 0.0, 1.0])
print(nn.BCEWithLogitsLoss()(logits_bin, target_bin))  # 0.3364

Practical tip: CrossEntropyLoss softmax included. DO NOT use softmax in forward() then passed in — will double softmax, train extremely slow.

6. Optimizers — SGD, Adam, AdamW

OptimizerAdvantagesDisadvantagesUse when
SGD+MomentumGeneralize wellSlow convergence, sensitive lrVision (CNN)
AdamFast convergence, less tuningGeneralize is worse than SGDDefault/prototype
AdamWFix weight decay, best for Transformers—Best default
optimizer_sgd = optim.SGD(model.parameters(), lr=0.01, momentum=0.9, weight_decay=1e-4)
optimizer_adam = optim.Adam(model.parameters(), lr=1e-3)
optimizer_adamw = optim.AdamW(model.parameters(), lr=5e-5, weight_decay=0.01)

# Learning rate schedulers
scheduler = optim.lr_scheduler.CosineAnnealingLR(optimizer_adamw, T_max=100)
Optimizer Decision Tree:
├── Vision (CNN)            ──▶ SGD + Momentum (lr=0.01-0.1)
├── NLP / Transformer       ──▶ AdamW (lr=1e-5 to 5e-5)
├── General / Prototype     ──▶ Adam (lr=1e-3)
└── Fine-tuning pretrained  ──▶ AdamW + low lr

Practical tip: Don't know what to choose → AdamW lr=1e-3, weight_decay=0.01. Safe default for almost every problem.

7. CNN Overview — Convolutional Neural Networks

CNN specializes in processing spatial data (images, videos) using convolution filters that slide over the input.

Input(3×32×32) ──▶ Conv+BN+ReLU ──▶ Pool ──▶ Conv+BN+ReLU ──▶ Pool ──▶ FC
                   (3→32, 3×3)     (2×2)    (32→64, 3×3)     (2×2)   (→10)
Feature maps: spatial giảm dần, channels tăng dần
ComponentsRolePyTorch
Conv2dExtract local featuresnn.Conv2d(in, out, kernel_size)
MaxPool2dReduce spatial sizenn.MaxPool2d(2)
BatchNorm2dStabilize trainingnn.BatchNorm2d(channels)
class SimpleCNN(nn.Module):
    def __init__(self, num_classes=10):
        super().__init__()
        self.features = nn.Sequential(
            nn.Conv2d(3, 32, 3, padding=1), nn.BatchNorm2d(32), nn.ReLU(), nn.MaxPool2d(2),
            nn.Conv2d(32, 64, 3, padding=1), nn.BatchNorm2d(64), nn.ReLU(), nn.MaxPool2d(2),
            nn.Conv2d(64, 128, 3, padding=1), nn.BatchNorm2d(128), nn.ReLU(),
            nn.AdaptiveAvgPool2d(1),
        )
        self.classifier = nn.Sequential(nn.Flatten(), nn.Dropout(0.5), nn.Linear(128, num_classes))

    def forward(self, x):
        return self.classifier(self.features(x))

7.1. ResNet — Skip Connections

ResNet solves vanishing gradients using skip connections: output = F(x) + x. Gradients flow directly through the identity path, allowing for training very deep networks (50-152 layers).

Practical tip: No one in production wrote CNN from scratch. Use pretrained ResNet, EfficientNet and then fine-tune (see part 9).

8. RNN/LSTM/GRU — Sequence Modeling

RNN processes sequential data (text, time series). Each step receives input + hidden state from the previous step.

x₁       x₂       x₃       x₄
 │        │        │        │
 ▼        ▼        ▼        ▼
[RNN]─h₁─[RNN]─h₂─[RNN]─h₃─[RNN]─h₄──▶ output

Vấn đề: vanishing gradient khi sequence dài → LSTM/GRU
ModelGatesAdvantagesUse cases
RNNNoFast, simpleShort sequences
LSTM3 (forget, input, output)Long-range dependenciesText, speech
GRU2 (reset, update)Almost equal to LSTM, fasterTrade-off speed
class LSTMClassifier(nn.Module):
    def __init__(self, vocab_size, embed_dim, hidden_dim, output_dim):
        super().__init__()
        self.embedding = nn.Embedding(vocab_size, embed_dim)
        self.lstm = nn.LSTM(embed_dim, hidden_dim, num_layers=2,
                            batch_first=True, dropout=0.3, bidirectional=True)
        self.fc = nn.Linear(hidden_dim * 2, output_dim)
        self.dropout = nn.Dropout(0.3)

    def forward(self, x):
        embedded = self.dropout(self.embedding(x))
        _, (hidden, _) = self.lstm(embedded)
        hidden_cat = torch.cat([hidden[-2], hidden[-1]], dim=1)
        return self.fc(self.dropout(hidden_cat))

Practical tip: LSTM/GRU has been almost completely replaced by Transformers in NLP. Lesson 4 will delve into Transformers — the architecture behind every modern LLM.

9. Transfer Learning

Get the model trained on a large dataset → fine-tune for a specific problem. This is the most effective way for production.

StrategyWhenHow to do
Feature ExtractionLittle data (<1K), similar domainsFreeze backbone, only train new FC
Partial Fine-tuningAverage data (1K-10K)Unfreeze the last few layers + FC
Full Fine-tuningLots of data (>10K), other domainsUnfreeze all, low lr
from torchvision import models

weights = models.ResNet18_Weights.IMAGENET1K_V1
model_tl = models.resnet18(weights=weights)

# Strategy 1: Feature Extraction — freeze backbone
for param in model_tl.parameters():
    param.requires_grad = False

num_features = model_tl.fc.in_features  # 512
model_tl.fc = nn.Sequential(nn.Dropout(0.3), nn.Linear(num_features, 5))
model_tl = model_tl.to(device)
optimizer = optim.Adam(model_tl.fc.parameters(), lr=1e-3)

# Strategy 2: Fine-tuning — unfreeze last layers, differential lr
for name, param in model_tl.named_parameters():
    if "layer4" in name or "fc" in name:
        param.requires_grad = True

optimizer = optim.AdamW([
    {"params": model_tl.layer4.parameters(), "lr": 1e-5},
    {"params": model_tl.fc.parameters(), "lr": 1e-3},
], weight_decay=0.01)

# Data preprocessing PHẢI dùng transform giống lúc pretrain
preprocess = weights.transforms()

Practical tip: Transfer Learning is the most important technique. Part 2 will use this same technique to fine-tune LLMs (BERT, GPT).

10. GPU Training & Performance

10.1. Efficient DataLoader

train_loader = DataLoader(
    train_dataset,
    batch_size=64,
    shuffle=True,
    num_workers=4,           # parallel data loading
    pin_memory=True,         # tăng tốc CPU→GPU transfer
    persistent_workers=True,
)

10.2. Mixed Precision (AMP)

Use float16 for most computation → speed up 2-3x, reduce VRAM ~50%:

from torch.amp import autocast, GradScaler

scaler = GradScaler("cuda")
for data, target in train_loader:
    data, target = data.to(device), target.to(device)
    with autocast("cuda"):
        output = model(data)
        loss = criterion(output, target)
    optimizer.zero_grad()
    scaler.scale(loss).backward()
    scaler.step(optimizer)
    scaler.update()
EngineeringSpeedupWhen to use
Mixed Precision (AMP)2-3x, ~50% VRAMAlways on GPU
pin_memory=True10-30% transferWhen using GPU
num_workers > 02-4x loadingWhen CPU idle
torch.compile()10-40%PyTorch 2.0+

Practical tip: Always on torch.amp + pin_memory=True on GPU — "free performance", does not affect accuracy.

11. Model Saving/Loading & Export

11.1. state_dict (PyTorch standard)

# Save checkpoint (training continuation)
torch.save({
    "model_state_dict": model.state_dict(),
    "optimizer_state_dict": optimizer.state_dict(),
    "epoch": epoch,
}, "checkpoint.pth")

# Load
ckpt = torch.load("checkpoint.pth", map_location=device, weights_only=True)
model.load_state_dict(ckpt["model_state_dict"])

# Production: chỉ save weights
torch.save(model.state_dict(), "model_weights.pth")

11.2. TorchScript & ONNX Export

# TorchScript — portable, chạy không cần Python
model.eval()
scripted = torch.jit.script(model)
scripted.save("model_scripted.pt")

# ONNX — cross-framework (ONNX Runtime, TensorRT, CoreML)
dummy = torch.randn(1, 1, 28, 28).to(device)
torch.onnx.export(model, dummy, "model.onnx",
    input_names=["input"], output_names=["output"],
    dynamic_axes={"input": {0: "batch"}, "output": {0: "batch"}},
    opset_version=17)
FormatUse CasePortable
state_dict (.pth)Training, PyTorch inferencePyTorch only
TorchScript (.pt)Production, mobile, C++No need for Python
ONNX (.onnx)Cross-framework deploymentAll frameworks
TensorRTNVIDIA GPU optimized inferenceNVIDIA only, fastest

Practical tip: Popular workflow: train PyTorch → save state_dict → export ONNX → deploy using ONNX Runtime.

Summary

  • ✅ Neural Network fundamentals: neurons, layers, activations (ReLU, GELU, Sigmoid)
  • ✅ PyTorch core: tensors, autograd, nn.Module — the main framework for DL
  • ✅ Training loop: forward → loss → zero_grad → backward → step
  • ✅ Loss functions: CrossEntropy, BCE, MSE — choose the right loss for the right task
  • ✅ Optimizers: AdamW is safe default, SGD for vision fine-tuning
  • ✅ CNN: Conv2d, pooling, ResNet skip connections
  • ✅ RNN/LSTM/GRU: sequence modeling — a stepping stone before Transformers
  • ✅ Transfer Learning: pretrained + fine-tuning — the most important technique
  • ✅ GPU training: AMP, DataLoader, pin_memory
  • ✅ Export model: state_dict, TorchScript, ONNX

Next → Lesson 4: NLP & Transformer Architecture — Self-Attention, Transformer architecture, BERT/GPT foundations. Direct foundation for understanding LLMs in Part 2.

Exercises

Exercise 1: MNIST CNN (Basic)

Train CNN for MNIST, achieving > 99% accuracy:

  • Use SimpleCNN From part 7, add learning rate scheduler
  • Turn on mixed precision (AMP), save best checkpoint accordingly val_accuracy

Exercise 2: Transfer Learning CIFAR-10 (Medium)

Fine-tune ResNet-18 for CIFAR-10:

  1. Try both feature extraction vs fine-tuning, compare accuracy + training time
  2. Export the best model to ONNX

Exercise 3: LSTM Text Classification (Advanced)

Build bidirectional LSTM for sentiment analysis (IMDB dataset):

  1. Implement tokenization + vocabulary building
  2. Compare with simple MLP on the same data
  3. Bonus: try GRU — compare speed vs accuracy

Exercise 4: Training Dashboard (Real combat)

Training script complete with Tensorboard logging, early stopping (patience=5), model checkpointing, config dataclass.