Introduction
You already know Ollama (convenient, has a large API, ecosystem) and MLX (faster on Mac). Natural question: can the two be combined?
Answer: Yes. Ollama supports MLX backend, allowing you to use Ollama's utilities (API, model management) but infer using the MLX engine.
1. MLX backend in Ollama
From version 0.5+, Ollama adds support for MLX on macOS. Instead of using llama.cpp (GGUF format), you can import the MLX model (safetensors format) and Ollama will use the MLX engine.
How it works
┌─────── Backend ────────┐
User ──► Ollama ──►│ llama.cpp (default) │──► Response
API │ MLX (khi dùng MLX model)│
└────────────────────────┘
2. Create Ollama model from MLX weights
Step 1: Download model MLX
# Dùng huggingface-cli
pip3 install huggingface-hub
huggingface-cli download mlx-community/Llama-3.2-3B-Instruct-4bit \
--local-dir ./models/llama-3.2-3b-mlx
Step 2: Create Modelfile
cat > Modelfile.mlx << 'EOF'
FROM ./models/llama-3.2-3b-mlx
PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER num_ctx 4096
SYSTEM """Bạn là trợ lý AI thông minh. Trả lời ngắn gọn, chính xác, bằng tiếng Việt."""
EOF
Step 3: Build model in Ollama
ollama create llama3.2-mlx -f Modelfile.mlx
Step 4: Run
ollama run llama3.2-mlx
Now Ollama will use MLX engine for this model, but you still use Ollama CLI and API as usual.
3. Comparison: same model, two backends
Benchmark script
#!/bin/bash
# benchmark-backends.sh
PROMPT="Write a Python function that implements binary search on a sorted list. Include docstring and type hints."
echo "=== Ollama + llama.cpp (GGUF) ==="
curl -s http://localhost:11434/api/generate -d "{
\"model\": \"llama3.2\",
\"prompt\": \"$PROMPT\",
\"stream\": false
}" | python3 -c "
import sys, json
d = json.load(sys.stdin)
pt = d['prompt_eval_duration']/1e9
gt = d['eval_duration']/1e9
print(f'Prompt: {d[\"prompt_eval_count\"]} tok in {pt:.2f}s = {d[\"prompt_eval_count\"]/pt:.0f} tok/s')
print(f'Generate: {d[\"eval_count\"]} tok in {gt:.2f}s = {d[\"eval_count\"]/gt:.0f} tok/s')
"
echo ""
echo "=== Ollama + MLX ==="
curl -s http://localhost:11434/api/generate -d "{
\"model\": \"llama3.2-mlx\",
\"prompt\": \"$PROMPT\",
\"stream\": false
}" | python3 -c "
import sys, json
d = json.load(sys.stdin)
pt = d['prompt_eval_duration']/1e9
gt = d['eval_duration']/1e9
print(f'Prompt: {d[\"prompt_eval_count\"]} tok in {pt:.2f}s = {d[\"prompt_eval_count\"]/pt:.0f} tok/s')
print(f'Generate: {d[\"eval_count\"]} tok in {gt:.2f}s = {d[\"eval_count\"]/gt:.0f} tok/s')
"
Sample results (M3 Pro 36GB)
| Metrics | llama.cpp | MLX | Speedup |
|---|---|---|---|
| Prompt processing | 285 tok/s | 640 tok/s | 2.2x |
| Token generation | 33 tok/s | 55 tok/s | 1.7x |
| Memory usage | 6.5 GB | 5.8 GB | -10% |
| API response time | 8.2s | 4.8s | 1.7x |
4. Context window tuning
Context window decides how much text model "remembers" in a conversation. Increasing context = consuming more RAM.
Calculate RAM for context
KV Cache memory ≈ 2 × n_layers × n_heads × head_dim × context_length × 2 bytes (FP16)
Llama 3.2 8B with context lengths:
| Context | KV Cache added | Total RAM |
|---|---|---|
| 2048 | ~0.5 GB | ~6 GB |
| 4096 | ~1 GB | ~6.5 GB |
| 8192 | ~2 GB | ~7.5 GB |
| 16384 | ~4 GB | ~9.5 GB |
| 32768 | ~8 GB | ~13.5 GB |
| 131072 | ~32 GB | ~37.5 GB |
Set context in Modelfile
FROM ./models/llama-3.2-3b-mlx
# Context window
PARAMETER num_ctx 8192
# Giảm context nếu ít RAM
# PARAMETER num_ctx 2048
Set context at runtime
# Override context khi chạy
ollama run llama3.2-mlx --num-ctx 16384
💡 Recommended: Start with
4096, gradually increase until the RAM runs out or enough for the use case.
5. Optimize performance
GPU utilization
# Kiểm tra model dùng GPU hay CPU
ollama ps
Ideal output: 100% GPU. If you see CPU, it means the model does not fit in GPU accessible memory.
Keep model in memory
By default, Ollama unloads the model after 5 minutes of idle time. Change:
# Giữ model loaded 30 phút
export OLLAMA_KEEP_ALIVE=30m
# Giữ vĩnh viễn (cho đến khi restart)
export OLLAMA_KEEP_ALIVE=-1
Or in the API:
curl http://localhost:11434/api/generate -d '{
"model": "llama3.2-mlx",
"keep_alive": -1
}'
Run multiple models simultaneously
# Cho phép 3 model cùng lúc
export OLLAMA_MAX_LOADED_MODELS=3
# Chạy song song
export OLLAMA_NUM_PARALLEL=4
⚠️ Each model occupies its own RAM. 3 models × 5GB = 15GB. Make sure there is enough RAM for macOS + other apps.
6. Monitoring performance
Activity Monitor
Open Activity Monitor → GPU tab to see:
- GPU utilization % during inference
- GPU memory usage
Terminal monitoring
# Xem GPU usage real-time
sudo powermetrics --samplers gpu_power -i 1000
# Xem memory pressure
memory_pressure
# Xem Ollama process
ps aux | grep ollama
Ollama logs
# Xem logs chi tiết
cat ~/.ollama/logs/server.log | tail -50
# Follow logs real-time
tail -f ~/.ollama/logs/server.log
7. Recommended Workflow
Based on practical experience, here is the optimal workflow:
Daily use: Ollama (llama.cpp)
# Model mặc định cho chat, hỏi đáp
ollama run qwen2.5:14b
- Stable, large ecosystem (Open WebUI, Continue.dev...)
- Fast enough for interactive chat
When you need speed: Ollama + MLX
# Model MLX cho tác vụ cần nhanh
ollama run qwen2.5-mlx
- Prompt processing is 2x fast → time-to-first-token is very low
- Batch processing multiple requests
Scripting/Pipeline: mlx-lm directly
from mlx_lm import load, generate
model, tokenizer = load("mlx-community/Qwen2.5-14B-Instruct-4bit")
# Custom pipeline, batch processing, fine-tuning...
- Complete control
- No Ollama server overhead
Summary
| Approach | Pros | Disadvantages | When to use |
|---|---|---|---|
| Ollama + llama.cpp | Stable, ecosystem | Slower than MLX | Daily default |
| Ollama + MLX | Faster, use Ollama API | More complicated setup | Need speed + API |
| mlx-lm live | Fastest, flexible | No API server | Scripting, pipeline |
Exercises
- Create Ollama model with MLX backend according to instructions in section 2
- Run a benchmark comparing two backends (llama.cpp vs MLX) for the same model
- Try increasing the context window: 2048 → 4096 → 8192. Record how much RAM usage increases?
- Test
OLLAMA_KEEP_ALIVE=-1— does it affect RAM when idle? - Build a personal workflow: choose model + backend for your 3 daily use cases
Next article: Ollama REST API →