Chuyển đến nội dung chính

Lesson 5: LLM Deep Dive — LLaMA, Mistral, Qwen, Phi

Detailed comparison of open-source LLMs: LLaMA 3, Mistral, Qwen 2.5, Phi-3/4. Architecture differences, benchmarks, use cases. Run locally with Ollama, vLLM. Commercial models: GPT-4, Claude, Gemini.

In 2023, LLaMA leak causes open-source AI to explode. In 2025, you can run model 70B on a laptop. The race between Meta, Mistral, Alibaba, Microsoft has completely changed the landscape. This article dives deep into the architecture, benchmarks, and how to actually run these models — from ollama run Go to serving production with vLLM.

1. Landscape LLM 2024-2026 — Open-Source vs Commercial

LLM Evolution Timeline:

2023 Q1         2023 Q3          2024 Q1-Q2        2024 Q3-2025      2025-2026
┌─────────┐   ┌──────────────┐  ┌──────────────┐  ┌─────────────┐  ┌─────────────┐
│ LLaMA 1 │──▶│ LLaMA 2      │─▶│ LLaMA 3      │─▶│ LLaMA 3.1-  │─▶│ LLaMA 4     │
│ GPT-4   │   │ Mistral 7B   │  │ Mixtral 8x22B│  │ 3.3, Qwen2.5│  │ Qwen 3      │
│ (leaked)│   │ Qwen 1,Phi-2 │  │ Phi-3,Claude3│  │ Mistral Lg  │  │ Phi-4       │
└─────────┘   └──────────────┘  └──────────────┘  └─────────────┘  └─────────────┘
GPT-4 chỉ     Open-source       Gap thu hẹp       Open ≈ Closed    Open dẫn đầu
qua API        bùng nổ          coding, math      nhiều task        nhiều benchmark
CriteriaOpen-WeightClosed-Source
RepresentationLLaMA, Mistral, Qwen, PhiGPT-4o, Claude 4, Gemini 2
CostInfra cost, no per-token feePay-per-token
CustomizationFine-tune, merge, quantizeLimited (system prompt, RAG)
PrivacyData stays on-premiseData sent via API
Best forProduction is autonomous, domain-specificFast prototype, SOTA quality

1.1. Terminology needs to be distinguished

  • Open-source: Code + weights + training data (rare — only OLMo, BLOOM)
  • Open-weight: Weights released, no training data (LLaMA, Mistral, Qwen)
  • Open-access: Free use via API but does not download weights
  • Proprietary: Nothing — API only (GPT-4, Claude)

2. Architecture Patterns — Decoder-Only, MoE, GQA, RoPE, SWA

2.1. Decoder-Only + GQA

Decoder-Only (GPT, LLaMA, Mistral):

Input tokens → Embeddings + RoPE
        │
        ▼
┌───────────────────────────────────────┐
│  Masked Self-Attention (causal)  │◄─ Chỉ nhìn tokens TRƯỚC
├───────────────────────────────────────┤
│  Feed-Forward (SwiGLU)           │
├───────────────────────────────────────┤
│  RMSNorm (pre-norm)              │
└───────────────────────────────────────┘
        │  × N layers (32-80)
        ▼
  Linear → Softmax → Next token

Grouped-Query Attention (GQA) — giảm KV cache:

MHA (cũ):        GQA (LLaMA 3):       MQA:
Q1 Q2 Q3 Q4      Q1 Q2 | Q3 Q4        Q1 Q2 Q3 Q4
│  │  │  │         \ | /   \ | /         \ |  | /
K1 K2 K3 K4        K1,2    K3,4           K_shared
V1 V2 V3 V4        V1,2    V3,4           V_shared
KV cache: 4×d     KV cache: 2×d         KV cache: 1×d

Key insight: GQA reduces cache KV by 4-8× but quality is nearly unchanged - the reason model 70B can run on consumer GPUs.

2.2. Rotary Position Embedding (RoPE)

RoPE encodes position by rotating the vector in complex space. Advantages: natural relative position, easily extend context length. Used in LLaMA, Mistral, Qwen — is the standard position encoding for all modern LLMs.

2.3. Sliding Window Attention (SWA) & MoE

SWA (Mistral): mỗi layer attend W tokens gần nhất
  Layer 1: window = 4096 → mỗi token "thấy" 4K tokens
  Layer N: effective range = N × W = 32 × 4096 = 131K tokens!
  Memory: O(n×w) thay vì O(n²) — tiết kiệm VRAM lớn

MoE (Mixtral 8×7B): sparse — chỉ activate 2/8 experts mỗi token
  Input → Router (gating) → top-2 experts → weighted sum
  Total params: 47B | Active: ~13B/token
  Speed ≈ 13B dense | Quality ≈ 70B dense
PatternUsed inBenefits
GQALLaMA 3, MistralCache KV reduced by 4-8×
RoPEMost LLMsExtend context is easy
SWAMistral 7BO(n×w) memory
MoEMixtral, Qwen-MoEFast inference, scalable
SwiGLULLaMA, MistralBetter activation than ReLU

3. LLaMA Family (Meta) — Open-Source Pillar

LLaMA 1 (Feb'23)      LLaMA 2 (Jul'23)      LLaMA 3 (Apr'24)
├─ 7B-65B              ├─ 7B, 13B, 70B       ├─ 8B, 70B (15T tokens!)
├─ Research only       ├─ Commercial license  ├─ GQA, RoPE, 128K vocab
└─ Leaked → bùng nổ   └─ RLHF, 4K ctx       └─ 128K ctx (3.1), tool use

LLaMA 3.2 (Sep'24): 1B/3B (edge) + 11B/90B (vision!)
LLaMA 3.3 (Dec'24): 70B cải thiện multilingual, instruction following
FeaturesLLaMA 3 8BLLaMA 3 70BLLaMA 3.1 405B
Layers3280126
Hidden dim4096819216384
KV heads (GQA)8816
Context8K→128K8K→128K128K
Training tokens15T+15T+15T+

Why is LLaMA important? The largest open-source AI Ecosystem. Most fine-tuned models are based on LLaMA base: Alpaca (Stanford), Vicuna (LMSYS), CodeLlama, Llama-Guard (safety), WizardLM. Key innovations: RMSNorm (pre-norm), SwiGLU activation, RoPE, GQA, 128K vocab. License: Llama Community — free under 700M MAU.

4. Mistral Family — Efficiency Is King

Mistral AI (French startup, ex-DeepMind/Meta) — philosophy: small model, big performance.

ModelSizeTypeContextHighlights
Mistral 7B7BDense32KSWA, beats LLaMA 2 13B
Mixtral 8×7B47B (13B active)MoE32KFirst MoE open
Mixtral 8×22B141B (39B active)MoE64KStrongest open MoE
Mistral Large~120BDense128KNear GPT-4
Codestral22BDense32KCode-specialized

Mistral 7B innovations: SWA (window=4096) + GQA (8 KV heads) + Rolling Buffer Cache (fixed memory, no grow) + Pre-fill Chunking → 7B model surpasses LLaMA 2 13B on most benchmarks. Mixtral 8×7B proves MoE practical: speed is 13B dense, quality is nearly 70B dense.

5. Qwen Family (Alibaba) — Multilingual Champion

Qwen 2.5 (Sep 2024) continuously ranks at the top of the Open LLM Leaderboard — especially strong multilingual (29+ languages, CJK, Vietnamese).

FeaturesDetails
Sizes0.5B → 72B (both MoE variants)
Training18T tokens, curated multilingual
Context128K native (YARN RoPE)
CodingQwen2.5-Coder — top open coding model
LicenseApache 2.0 (most widespread)
Tool useNative function calling, JSON mode

QwQ (32B) — OpenAI o1 style reasoning model but open-weight. "Think" many steps before answering, especially math/code. The model asks itself questions, verifies each step, then outputs the final result.

QwQ reasoning flow:
  Prompt: "Sum of primes < 20?"
  → "Let me think step by step..."
  → "Primes: 2, 3, 5, 7, 11, 13, 17, 19"
  → "Wait, is 9 prime? 9 = 3×3, no."
  → "Sum = 2+3+5+7+11+13+17+19 = 77"
  → "The answer is 77."
  
  Self-verification → fewer errors on complex tasks

6. Phi Family (Microsoft) — Small Models, Big Performance

Philosophy: data quality > model size — train on "textbook-quality" + synthetic data.

Phi Philosophy:  Web crawl (noisy) × Massive scale → okay model
                 vs
                 Curated data × Modest scale → GREAT model
ModelSizeContextMMLUHighlights
Phi-22.7B2K56.7Beats some 7B models
Phi-3 Mini3.8B128K68.8On-device, extended ctx
Phi-3 Medium14B128K78.0Near GPT-3.5
Phi-3.5 MoE42B (6.6B active)128K78.9Sparse MoE
Phi-414B16K81.414B competes with 70B!

Phi-4 achieved MMLU 81.4, MATH 80.4, HumanEval 82.6 — equal to LLaMA 3.1 70B on reasoning tasks. Secret: synthetic data generation pipeline creates high-quality training data from larger LLMs, filters and validates before training.

7. Commercial Models — GPT-4o, Claude, Gemini

ModelContextStrengthsPricing (1M tok)
GPT-4o128KGeneral-purpose, fast$2.5 in / $10 out
GPT-4o-mini128KBudget-friendly$0.15 / $0.60
o1 / o3200KDeep reasoning$15 / $60
Claude 3.5 Sonnets200KCoding, analysis$3 / $15
Claude 4 Opus200KAgentic, tool use$15 / $75
Gemini 2 Flash1MSpeed, huge context$0.075 / $0.30
Gemini 2 Pro2MLong context king$1.25 / $10

8. Mega Benchmark Table

ModelMMLUHumanEvalGSM8KArena ELO
GPT-4o88.790.297.81285
Claude 4 Opus88.192.096.51290
Gemini 2 Pro87.585.096.01270
LLaMA 3.1 405B85.289.096.81210
Qwen 2.5 72B83.186.495.81190
Phi-4 (14B)81.482.695.31150
LLaMA 3.3 70B82.084.595.11180
Mistral Large81.282.093.01160
Phi-3 Mini (3.8B)68.858.582.51010

Benchmark explanation: MMLU = knowledge 57 subjects, HumanEval = Python code gene, GSM8K = math reasoning, Arena ELO = real human preference (most reliable).

Caution: Benchmarks can be "gamed" — training data contamination, benchmark-specific tuning. Arena ELO from real users is the most reliable.

9. Run LLM Local — Ollama, vLLM, llama.cpp

The most important part: hands-on. You will learn how to run LLMs right on your computer.

9.1. Ollama — 5 minutes setup

# Install
curl -fsSL https://ollama.com/install.sh | sh  # Linux
brew install ollama                              # macOS

# Chạy models
ollama run llama3.1          # LLaMA 3.1 8B (~4.7GB)
ollama run mistral           # Mistral 7B
ollama run qwen2.5           # Qwen 2.5 7B
ollama run phi4              # Phi-4 14B (~8GB)
ollama run qwen2.5-coder     # Coding optimized
ollama run llama3.1:70b      # 70B — cần ~40GB VRAM

9.2. Ollama API — OpenAI-Compatible

from openai import OpenAI

# Ollama chạy OpenAI-compatible API tại localhost:11434
client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama",  # Không cần real key
)

response = client.chat.completions.create(
    model="llama3.1",
    messages=[
        {"role": "system", "content": "Trả lời ngắn gọn bằng tiếng Việt."},
        {"role": "user", "content": "GQA là gì trong Transformer?"},
    ],
    temperature=0.7,
)
print(response.choices[0].message.content)

9.3. vLLM — Production Serving

vLLM (UC Berkeley) is the fastest inference engine — uses PagedAttention for efficient KV cache management, supporting continuous batching.

pip install vllm

# Serve với OpenAI-compatible API
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.1-8B-Instruct \
    --port 8000 \
    --gpu-memory-utilization 0.9
# Gọi vLLM — same OpenAI format
client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")
response = client.chat.completions.create(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[{"role": "user", "content": "Hello!"}],
)

9.4. llama.cpp — CPU Inference & GGUF

llama.cpp (Georgi Gerganov) allows running LLM on pure CPU or mixed CPU/GPU. Format GGUF is the standard for quantized models.

brew install llama.cpp  # macOS

# Download GGUF model từ Hugging Face
huggingface-cli download \
  bartowski/Meta-Llama-3.1-8B-Instruct-GGUF \
  Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
  --local-dir ./models/

# Chạy interactive chat
llama-cli -m ./models/Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
  -c 4096 --chat-template llama3 -i

9.5. Hugging Face Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_name = "microsoft/Phi-4"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name, torch_dtype=torch.bfloat16, device_map="auto",
)

messages = [{"role": "user", "content": "Explain MoE briefly."}]
inputs = tokenizer.apply_chat_template(
    messages, return_tensors="pt", add_generation_prompt=True
).to(model.device)
outputs = model.generate(inputs, max_new_tokens=300, temperature=0.7, do_sample=True)
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))

9.6. Compare Tools

ToolsBest ForSpeed ​​EaseGPU?
OllamaDev, experimentGood⭐⭐⭐⭐⭐Optional
vLLMProduction serving⭐⭐⭐⭐⭐⭐⭐⭐Yes
llama.cppCPU/edgeGood⭐⭐⭐Optional
TransformersFine-tune, researchGood⭐⭐⭐⭐Recommended

10. Quantization — GGUF, GPTQ, AWQ

Quantization reduces the precision of model weights: FP16 (16-bit) → INT4 (4-bit). Result: 3-4× smaller model, faster inference, but may reduce quality. This is a required technique when running large models on consumer hardware.

LLaMA 3.1 8B (FP16): ~16 GB
        │
  ┌─────┼─────────────────┐
  ▼     ▼                 ▼
GGUF   GPTQ              AWQ
(CPU+) (GPU)              (GPU)
Q4_K_M: ~4.9GB   INT4: ~4.5GB   INT4: ~4.3GB
Best: llama.cpp   Best: HF        Best: vLLM
      Ollama            AutoGPTQ         Fastest
QuantBitsSize (8B)QualityUse Case
Q3_K_M3-4~3.9 GB~90%Limited VRAM
Q4_K_M4-5~4.9 GB~95%Sweet spot
Q5_K_M5-6~5.7 GB~97%Better quality
Q8_08~8.5 GB~99.5%Near-lossless
FP1616~16 GB100%Baseline
# Tạo custom model với Ollama Modelfile
cat > Modelfile << 'EOF'
FROM llama3.1:8b-instruct-q4_K_M
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
SYSTEM "Bạn là trợ lý AI, trả lời bằng tiếng Việt, ngắn gọn kèm ví dụ."
EOF

ollama create my-assistant -f Modelfile
ollama run my-assistant "Giải thích MoE là gì?"

Rule of thumb: Use Q4_K_M for most use cases. Need more quality: Q5_K_M. VRAM limit: Q3_K_M.

11. Choose the Right Model — Decision Tree

There is no "best model" — there is only the most suitable model for the use case, budget, and hardware.

START: Bạn cần gì?
├─► Prototype nhanh → GPT-4o-mini / Claude 3.5 Sonnet (API)
├─► Production on-premise
│   ├─ GPU (A100+)  → vLLM + LLaMA 3.1 70B / Qwen 2.5 72B
│   └─ No GPU       → Ollama + Phi-4 (Q4) / LLaMA 3.1 8B (Q4)
├─► Coding → Qwen2.5-Coder (local) / Claude 4 (API)
├─► Multilingual (CJK, Vietnamese) → Qwen 2.5
├─► Mobile/Edge (<4GB) → Phi-3 Mini / LLaMA 3.2 1B-3B
├─► Math/Reasoning → QwQ 32B / o1-mini (API) / Phi-4
└─► Longest context → Gemini 2 (1M-2M) / Claude 4 (200K)

11.1. Hardware Requirements

Model SizeVRAM (FP16)VRAM (Q4)RAM (CPU)
1-3B4 GB2 GB4 GB
7-8B16 GB6 GB8 GB
13-14B28 GB10 GB16 GB
70-72B140 GB42 GB64 GB
405B810 GB230 GB256 GB+
MacBook M-series (8GB):   → Phi-3 Mini, LLaMA 3.2 3B (Q4)
MacBook M-series (16GB):  → LLaMA 3.1 8B (Q4), Mistral 7B
MacBook M-series (32GB):  → Phi-4, Qwen 2.5 14B (Q4)
RTX 4090 (24GB):          → LLaMA 3.1 8B (FP16), 70B (Q4)
A100 (80GB):              → LLaMA 3.1 70B (FP16)
8× H100:                  → LLaMA 3.1 405B (FP16)

Summary

Key Takeaways:

1. ARCHITECTURE: Decoder-only + GQA + RoPE = standard recipe
   MoE = quality lớn compute nhỏ | SWA = memory-efficient

2. OPEN-SOURCE FAMILIES:
   LLaMA (Meta) → ecosystem lớn nhất
   Mistral → efficiency king, SWA + MoE pioneer
   Qwen (Alibaba) → best multilingual, Apache 2.0
   Phi (Microsoft) → small model big performance

3. CHẠY LOCAL: Ollama (5 phút) → vLLM (production) → llama.cpp (CPU)

4. QUANTIZATION: Q4_K_M = sweet spot (3× nhỏ, ~95% quality)

5. Không có "best model" — chỉ có "best fit" cho context cụ thể

Exercises

Exercise 1: Run & Compare Models (30 minutes)

  1. Install Ollama, pull 3 models: llama3.1, qwen2.5, phi4
  2. Ask the same question for all 3 (eg: "Explain Transformer in 5 sentences")
  3. Compare response quality, speed, style

Exercise 2: Ollama API Integration (30 minutes)

  1. Write a Python script using OpenAI SDK to call Ollama local
  2. Implement compare_models(prompt, models) — send same prompt to multiple models, return result + elapsed time
  3. Measure tokens/second for each model

Exercise 3: Quantization Hands-on (20 minutes)

  1. Create a custom Modelfile with Vietnamese system prompt
  2. Build model: ollama create my-assistant -f Modelfile
  3. Measure VRAM/RAM usage when running the model

Exercise 4: Model Selection (20 minutes)

Choose model + deployment for each scenario, explain why:

  1. Healthcare chatbot — data must be on-premise, low budget, needs to be accurate
  2. Mobile AI assistant — Android 4GB RAM, offline mode
  3. Research lab — analyze Chinese + English papers, long context (>50K tokens)
  4. E-commerce support — 1000+ concurrent requests/minute, fast response
  5. Code review tool — review PRs for a team of 50 devs, many languages

Exercise 5: Benchmark Research (20 minutes)

  1. Access Open LLM Leaderboard
  2. Compare top 5 models — which model leads which benchmark?
  3. Find "hidden gem": model <15B but high rank
  4. Write a short report analyzing trends: MoE vs Dense, small vs large