Introduction
Ollama registry has hundreds of models. So which model to choose? This article is a real-life cheat sheet: comparing the most popular models, helping you choose the right model for each task and Mac configuration.
1. Understand Model Naming Convention
When you see qwen2.5:14b-instruct-q4_K_M, here's how to read:
qwen2.5 : 14b - instruct - q4_K_M
└── Family └── Size └── Variant └── Quantization
- Family: Original model name (Llama, Gemma, Qwen, Mistral...)
- Size: Number of parameters (1B, 3B, 7B, 8B, 14B, 32B, 70B...)
- Variant:
instruct(chat),base(raw),code(coding) - Quantization: Model compression level (Q4, Q5, Q8, F16)
2. Quantization — Which Q to choose?
Quantization is a technique to reduce model size by reducing the precision of weights:
| Quantization | Bits/weight | Size ratio | Quality | Speed |
|---|---|---|---|---|
| F16 | 16 bits | 100% (baseline) | Best | Slowest |
| Q8_0 | 8 bits | ~50% | Almost equal to F16 | Faster |
| Q6_K | 6 bits | ~37% | Very good | Fast |
| Q5_K_M | 5 bits | ~31% | Good | Fast |
| Q4_K_M | 4 bits | ~25% | Pretty good | Fastest |
| Q3_K_M | 3 bits | ~19% | Significant reduction | Very fast |
| Q2_K | 2 bits | ~12% | Poor | Extremely fast |
💡 Recommended: Q4_K_M is the sweet spot for most cases. Reduced size by ~75% but still good quality. Q5_K_M if you want a little higher quality.
Real world example: Llama 3.2 8B
| Quantization | Size on disk | RAM needed | Tokens/s (M3 Pro) |
|---|---|---|---|
| F16 | 16 GB | ~18 GB | ~12 tok/s |
| Q8_0 | 8.5 GB | ~10 GB | ~24 tok/s |
| Q4_K_M | 4.9 GB | ~6.5 GB | ~32 tok/s |
3. Comprehensive model comparison table
Small model group (1B-4B) — MacBook Air 8GB
| Model | Size (Q4) | RAM min | Code | Vietnamese | Overview |
|---|---|---|---|---|---|
| Llama 3.2 3B | 2.0 GB | 4 GB | ★★★☆ | ★★☆☆ | Jack of all trades |
| Gemma 3 4B | 3.3 GB | 5 GB | ★★★☆ | ★★★☆ | Good multilingual |
| Phi-4 Mini 3.8B | 2.5 GB | 4.5 GB | ★★★★ | ★★☆☆ | Code/math is very powerful |
| Qwen 2.5 3B | 1.9 GB | 4 GB | ★★★☆ | ★★★☆ | Good balance |
Medium model group (7B-14B) — MacBook 16-24GB
| Model | Size (Q4) | RAM min | Code | Vietnamese | Overview |
|---|---|---|---|---|---|
| Llama 3.2 8B | 4.9 GB | 7 GB | ★★★★ | ★★★☆ | Best all round |
| Gemma 3 12B | 8.1 GB | 10 GB | ★★★★ | ★★★★ | Multilingual champion |
| Qwen 2.5 14B | 9.0 GB | 11 GB | ★★★★ | ★★★★★ | Best Vietnamese |
| Mistral 7B | 4.1 GB | 6 GB | ★★★★ | ★★★☆ | Code/reasoning solid |
| DeepSeek Coder V2 16B | 10.2 GB | 13 GB | ★★★★★ | ★★☆☆ | Coding beast |
Large model group (30B+) — MacBook 32GB+
| Model | Size (Q4) | RAM min | Code | Vietnamese | Overview |
|---|---|---|---|---|---|
| Qwen 2.5 32B | 18 GB | 22 GB | ★★★★★ | ★★★★★ | Best local model |
| Llama 3.3 70B | 40 GB | 48 GB | ★★★★★ | ★★★★ | Needs 64GB+ RAM |
| DeepSeek V3 (distill 32B) | 19 GB | 23 GB | ★★★★★ | ★★★☆ | Reasoning king |
4. Choose model according to use case
Smart Chatbot / Q&A
# Tiếng Việt tốt nhất
ollama run qwen2.5:14b
# Cân bằng nhất
ollama run llama3.2
# RAM ít (8GB)
ollama run gemma3:4b
Code writing / Code review
# Coding chuyên sâu
ollama run deepseek-coder-v2:16b
# Cân bằng code + chat
ollama run qwen2.5-coder:14b
# RAM ít
ollama run phi4-mini
Summary / Writing
# Tiếng Việt
ollama run qwen2.5:14b
# Tiếng Anh
ollama run llama3.2
Image analysis (Vision)
# Vision tốt nhất
ollama run gemma3:12b # Có vision built-in
# Nhẹ hơn
ollama run llava:7b
5. Select model according to Mac RAM
8 GB RAM (MacBook Air M1/M2 base)
# Chỉ nên dùng model 3-4B
ollama run llama3.2:3b # 2.0 GB, chạy tốt
ollama run phi4-mini # 2.5 GB, code tốt
ollama run gemma3:1b # 1.0 GB, cực nhẹ
⚠️ With 8GB, close Safari before running model 3B. macOS requires ~4GB for the system.
16 GB RAM
# Sweet spot
ollama run llama3.2 # 8B, 4.9 GB
ollama run qwen2.5:7b # 7B, 4.7 GB
ollama run mistral # 7B, 4.1 GB
24-36 GB RAM
# Mở rộng lên 12-14B
ollama run gemma3:12b # 8.1 GB
ollama run qwen2.5:14b # 9.0 GB, khuyến nghị
ollama run deepseek-coder-v2:16b # 10.2 GB
48-64 GB+ RAM
# Model lớn, chất lượng gần cloud
ollama run qwen2.5:32b # 18 GB
ollama run llama3.3:70b # 40 GB, cần 48GB+ RAM
6. Realistic benchmark
Run a simple benchmark:
# Đo thời gian generate
time ollama run llama3.2 "Viết function fibonacci bằng Python" --nowordwrap
Or use Ollama API to get accurate metrics:
curl -s http://localhost:11434/api/generate -d '{
"model": "llama3.2",
"prompt": "Explain Docker in 3 sentences",
"stream": false
}' | python3 -c "
import sys, json
d = json.load(sys.stdin)
prompt_tokens = d['prompt_eval_count']
gen_tokens = d['eval_count']
prompt_time = d['prompt_eval_duration'] / 1e9
gen_time = d['eval_duration'] / 1e9
print(f'Prompt: {prompt_tokens} tokens in {prompt_time:.2f}s ({prompt_tokens/prompt_time:.1f} tok/s)')
print(f'Generate: {gen_tokens} tokens in {gen_time:.2f}s ({gen_tokens/gen_time:.1f} tok/s)')
"
7. Model tags you should know
When pulling a model, the default tag is latest (usually instruct + Q4_K_M). But you can specify:
# Chất lượng cao hơn (tốn RAM hơn)
ollama pull llama3.2:8b-instruct-q8_0
# Nhẹ nhất có thể
ollama pull llama3.2:3b-instruct-q4_0
# Model vision
ollama pull gemma3:12b # Tự động có vision
# Chỉ lấy text model
ollama pull gemma3:4b-it-q4_K_M
See all available tags:
# Truy cập: https://ollama.com/library/llama3.2/tags
# Hoặc: https://ollama.com/library/qwen2.5/tags
Summary
| Mac RAM | Recommended model | Size |
|---|---|---|
| 8 GB | llama3.2:3b, phi4-mini | 2-3 GB |
| 16 GB | llama3.2, qwen2.5:7b | 4-5 GB |
| 24-36 GB | qwen2.5:14b, gemma3:12b | 8-10 GB |
| 48 GB+ | qwen2.5:32b | 18 GB |
Quantization: Always start with Q4_K_M (default). Only upgrade to Q5/Q8 when you want better quality and have extra RAM.
Exercises
- Based on your Mac's RAM, choose 2-3 suitable models and download them
- Ask the same question about Vietnamese for each model, noting which model answers best
- Use the benchmark script above to measure tokens/second of each model
- Compare Q4 vs Q8 with the same model: is there a clear difference in quality? How much different speed?
Next article: MLX Framework — Accelerate 3x inference →