Chuyển đến nội dung chính

Lesson 3: Choose the right model - Compare LLM for Mac

Comprehensive comparison table: Llama 3.2 vs Gemma 3 vs Qwen 2.5 vs Mistral vs Phi-4. RAM requirements for each model size. Quantization (Q4, Q5, Q8) affects speed vs quality. Choose model according to use case.

🧠 AI & ML — Lesson 2 Lesson 3: Choose the right model - Compare LLM for Mac

Running AI Local with Ollama on Apple Silicon

Part 1: Platform - Ollama & Apple Silicon

xdev.asia

Introduction

Ollama registry has hundreds of models. So which model to choose? This article is a real-life cheat sheet: comparing the most popular models, helping you choose the right model for each task and Mac configuration.


1. Understand Model Naming Convention

When you see qwen2.5:14b-instruct-q4_K_M, here's how to read:

qwen2.5    :  14b    -  instruct   -  q4_K_M
└── Family    └── Size   └── Variant    └── Quantization
  • Family: Original model name (Llama, Gemma, Qwen, Mistral...)
  • Size: Number of parameters (1B, 3B, 7B, 8B, 14B, 32B, 70B...)
  • Variant: instruct (chat), base (raw), code (coding)
  • Quantization: Model compression level (Q4, Q5, Q8, F16)

2. Quantization — Which Q to choose?

Quantization is a technique to reduce model size by reducing the precision of weights:

QuantizationBits/weightSize ratioQualitySpeed ​​
F1616 bits100% (baseline)BestSlowest
Q8_08 bits~50%Almost equal to F16Faster
Q6_K6 bits~37%Very goodFast
Q5_K_M5 bits~31%GoodFast
Q4_K_M4 bits~25%Pretty goodFastest
Q3_K_M3 bits~19%Significant reductionVery fast
Q2_K2 bits~12%PoorExtremely fast

💡 Recommended: Q4_K_M is the sweet spot for most cases. Reduced size by ~75% but still good quality. Q5_K_M if you want a little higher quality.

Real world example: Llama 3.2 8B

QuantizationSize on diskRAM neededTokens/s (M3 Pro)
F1616 GB~18 GB~12 tok/s
Q8_08.5 GB~10 GB~24 tok/s
Q4_K_M4.9 GB~6.5 GB~32 tok/s

3. Comprehensive model comparison table

Small model group (1B-4B) — MacBook Air 8GB

ModelSize (Q4)RAM minCodeVietnameseOverview
Llama 3.2 3B2.0 GB4 GB★★★☆★★☆☆Jack of all trades
Gemma 3 4B3.3 GB5 GB★★★☆★★★☆Good multilingual
Phi-4 Mini 3.8B2.5 GB4.5 GB★★★★★★☆☆Code/math is very powerful
Qwen 2.5 3B1.9 GB4 GB★★★☆★★★☆Good balance

Medium model group (7B-14B) — MacBook 16-24GB

ModelSize (Q4)RAM minCodeVietnameseOverview
Llama 3.2 8B4.9 GB7 GB★★★★★★★☆Best all round
Gemma 3 12B8.1 GB10 GB★★★★★★★★Multilingual champion
Qwen 2.5 14B9.0 GB11 GB★★★★★★★★★Best Vietnamese
Mistral 7B4.1 GB6 GB★★★★★★★☆Code/reasoning solid
DeepSeek Coder V2 16B10.2 GB13 GB★★★★★★★☆☆Coding beast

Large model group (30B+) — MacBook 32GB+

ModelSize (Q4)RAM minCodeVietnameseOverview
Qwen 2.5 32B18 GB22 GB★★★★★★★★★★Best local model
Llama 3.3 70B40 GB48 GB★★★★★★★★★Needs 64GB+ RAM
DeepSeek V3 (distill 32B)19 GB23 GB★★★★★★★★☆Reasoning king

4. Choose model according to use case

Smart Chatbot / Q&A

# Tiếng Việt tốt nhất
ollama run qwen2.5:14b

# Cân bằng nhất
ollama run llama3.2

# RAM ít (8GB)
ollama run gemma3:4b

Code writing / Code review

# Coding chuyên sâu
ollama run deepseek-coder-v2:16b

# Cân bằng code + chat
ollama run qwen2.5-coder:14b

# RAM ít
ollama run phi4-mini

Summary / Writing

# Tiếng Việt
ollama run qwen2.5:14b

# Tiếng Anh
ollama run llama3.2

Image analysis (Vision)

# Vision tốt nhất
ollama run gemma3:12b   # Có vision built-in

# Nhẹ hơn
ollama run llava:7b

5. Select model according to Mac RAM

8 GB RAM (MacBook Air M1/M2 base)

# Chỉ nên dùng model 3-4B
ollama run llama3.2:3b     # 2.0 GB, chạy tốt
ollama run phi4-mini       # 2.5 GB, code tốt
ollama run gemma3:1b       # 1.0 GB, cực nhẹ

⚠️ With 8GB, close Safari before running model 3B. macOS requires ~4GB for the system.

16 GB RAM

# Sweet spot
ollama run llama3.2        # 8B, 4.9 GB
ollama run qwen2.5:7b      # 7B, 4.7 GB
ollama run mistral          # 7B, 4.1 GB

24-36 GB RAM

# Mở rộng lên 12-14B
ollama run gemma3:12b      # 8.1 GB
ollama run qwen2.5:14b     # 9.0 GB, khuyến nghị
ollama run deepseek-coder-v2:16b  # 10.2 GB

48-64 GB+ RAM

# Model lớn, chất lượng gần cloud
ollama run qwen2.5:32b     # 18 GB
ollama run llama3.3:70b    # 40 GB, cần 48GB+ RAM

6. Realistic benchmark

Run a simple benchmark:

# Đo thời gian generate
time ollama run llama3.2 "Viết function fibonacci bằng Python" --nowordwrap

Or use Ollama API to get accurate metrics:

curl -s http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Explain Docker in 3 sentences",
  "stream": false
}' | python3 -c "
import sys, json
d = json.load(sys.stdin)
prompt_tokens = d['prompt_eval_count']
gen_tokens = d['eval_count']
prompt_time = d['prompt_eval_duration'] / 1e9
gen_time = d['eval_duration'] / 1e9
print(f'Prompt: {prompt_tokens} tokens in {prompt_time:.2f}s ({prompt_tokens/prompt_time:.1f} tok/s)')
print(f'Generate: {gen_tokens} tokens in {gen_time:.2f}s ({gen_tokens/gen_time:.1f} tok/s)')
"

7. Model tags you should know

When pulling a model, the default tag is latest (usually instruct + Q4_K_M). But you can specify:

# Chất lượng cao hơn (tốn RAM hơn)
ollama pull llama3.2:8b-instruct-q8_0

# Nhẹ nhất có thể
ollama pull llama3.2:3b-instruct-q4_0

# Model vision
ollama pull gemma3:12b    # Tự động có vision

# Chỉ lấy text model
ollama pull gemma3:4b-it-q4_K_M

See all available tags:

# Truy cập: https://ollama.com/library/llama3.2/tags
# Hoặc: https://ollama.com/library/qwen2.5/tags

Summary

Mac RAMRecommended modelSize
8 GBllama3.2:3b, phi4-mini2-3 GB
16 GBllama3.2, qwen2.5:7b4-5 GB
24-36 GBqwen2.5:14b, gemma3:12b8-10 GB
48 GB+qwen2.5:32b18 GB

Quantization: Always start with Q4_K_M (default). Only upgrade to Q5/Q8 when you want better quality and have extra RAM.


Exercises

  1. Based on your Mac's RAM, choose 2-3 suitable models and download them
  2. Ask the same question about Vietnamese for each model, noting which model answers best
  3. Use the benchmark script above to measure tokens/second of each model
  4. Compare Q4 vs Q8 with the same model: is there a clear difference in quality? How much different speed?

Next article: MLX Framework — Accelerate 3x inference →