Introduction
LoRA and QLoRA are the most economical fine-tune techniques — only updating 0.1–1% parameters, running on free GPUs (Colab T4).
1. LoRA — Intuition
Thay vì update TOÀN BỘ weight matrix W (d × d):
W_new = W + ΔW
LoRA decompose ΔW thành 2 matrices nhỏ:
ΔW = A × B (A: d × r, B: r × d, với r << d)
Ví dụ:
W: 4096 × 4096 = 16.7M params → cập nhật TẤT CẢ
LoRA (r=16): 4096×16 + 16×4096 = 131K params → cập nhật 0.8%
2. QLoRA — Quantize + LoRA
QLoRA = 4-bit quantization base model + LoRA adapters
→ Giảm VRAM từ 32GB → 6GB
→ Chạy được trên Google Colab T4 (16GB)!
3. Hands-on with Unsloth
from unsloth import FastLanguageModel
# Load model với 4-bit quantization
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/Meta-Llama-3.1-8B-Instruct",
max_seq_length=2048,
load_in_4bit=True,
)
# Add LoRA adapters
model = FastLanguageModel.get_peft_model(
model, r=16, lora_alpha=16,
target_modules=["q_proj","k_proj","v_proj","o_proj"],
lora_dropout=0,
)
# Train
from trl import SFTTrainer
trainer = SFTTrainer(
model=model,
dataset=dataset,
max_seq_length=2048,
args=TrainingArguments(
per_device_train_batch_size=2,
num_train_epochs=3,
learning_rate=2e-4,
output_dir="outputs",
),
)
trainer.train()
Summary
- LoRA: only update ~1% parameters → save GPU and time
- QLoRA: added quantization → runs on consumer GPU
- Unsloth: 2x faster than standard LoRA training
- Cost: $0 on Colab, or ~$1–3/hour cloud GPU
Exercises
- Fine-tune LLaMA 3 8B on Google Colab (QLoRA)
- Output comparison: LoRA FT vs API FT (Gemini/OpenAI)
- Try rank r=8 vs r=16 vs r=32 — compare quality
- Merge LoRA adapters and export the complete model