Chuyển đến nội dung chính

Lesson 6: Ollama + MLX backend - Combining the best of two worlds

Configure Ollama to use MLX backend instead of llama.cpp. Detailed benchmarks. Optimize context window. When to use MLX backend, when to use llama.cpp.

🧠 AI & ML — Lesson 2 Lesson 6: Ollama + MLX backend - Good combination best of two worlds

Running AI Local with Ollama on Apple Silicon

Part 2: MLX - 3x acceleration with Apple's native framework

xdev.asia

Introduction

You already know Ollama (convenient, has a large API, ecosystem) and MLX (faster on Mac). Natural question: can the two be combined?

Answer: Yes. Ollama supports MLX backend, allowing you to use Ollama's utilities (API, model management) but infer using the MLX engine.


1. MLX backend in Ollama

From version 0.5+, Ollama adds support for MLX on macOS. Instead of using llama.cpp (GGUF format), you can import the MLX model (safetensors format) and Ollama will use the MLX engine.

How it works

                    ┌─────── Backend ────────┐
User ──► Ollama ──►│ llama.cpp (default)     │──► Response
         API       │ MLX (khi dùng MLX model)│
                    └────────────────────────┘

2. Create Ollama model from MLX weights

Step 1: Download model MLX

# Dùng huggingface-cli
pip3 install huggingface-hub
huggingface-cli download mlx-community/Llama-3.2-3B-Instruct-4bit \
  --local-dir ./models/llama-3.2-3b-mlx

Step 2: Create Modelfile

cat > Modelfile.mlx << 'EOF'
FROM ./models/llama-3.2-3b-mlx

PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER num_ctx 4096

SYSTEM """Bạn là trợ lý AI thông minh. Trả lời ngắn gọn, chính xác, bằng tiếng Việt."""
EOF

Step 3: Build model in Ollama

ollama create llama3.2-mlx -f Modelfile.mlx

Step 4: Run

ollama run llama3.2-mlx

Now Ollama will use MLX engine for this model, but you still use Ollama CLI and API as usual.


3. Comparison: same model, two backends

Benchmark script

#!/bin/bash
# benchmark-backends.sh

PROMPT="Write a Python function that implements binary search on a sorted list. Include docstring and type hints."

echo "=== Ollama + llama.cpp (GGUF) ==="
curl -s http://localhost:11434/api/generate -d "{
  \"model\": \"llama3.2\",
  \"prompt\": \"$PROMPT\",
  \"stream\": false
}" | python3 -c "
import sys, json
d = json.load(sys.stdin)
pt = d['prompt_eval_duration']/1e9
gt = d['eval_duration']/1e9
print(f'Prompt: {d[\"prompt_eval_count\"]} tok in {pt:.2f}s = {d[\"prompt_eval_count\"]/pt:.0f} tok/s')
print(f'Generate: {d[\"eval_count\"]} tok in {gt:.2f}s = {d[\"eval_count\"]/gt:.0f} tok/s')
"

echo ""
echo "=== Ollama + MLX ==="
curl -s http://localhost:11434/api/generate -d "{
  \"model\": \"llama3.2-mlx\",
  \"prompt\": \"$PROMPT\",
  \"stream\": false
}" | python3 -c "
import sys, json
d = json.load(sys.stdin)
pt = d['prompt_eval_duration']/1e9
gt = d['eval_duration']/1e9
print(f'Prompt: {d[\"prompt_eval_count\"]} tok in {pt:.2f}s = {d[\"prompt_eval_count\"]/pt:.0f} tok/s')
print(f'Generate: {d[\"eval_count\"]} tok in {gt:.2f}s = {d[\"eval_count\"]/gt:.0f} tok/s')
"

Sample results (M3 Pro 36GB)

Metricsllama.cppMLXSpeedup
Prompt processing285 tok/s640 tok/s2.2x
Token generation33 tok/s55 tok/s1.7x
Memory usage6.5 GB5.8 GB-10%
API response time8.2s ​​4.8s1.7x

4. Context window tuning

Context window decides how much text model "remembers" in a conversation. Increasing context = consuming more RAM.

Calculate RAM for context

KV Cache memory ≈ 2 × n_layers × n_heads × head_dim × context_length × 2 bytes (FP16)

Llama 3.2 8B with context lengths:

ContextKV Cache addedTotal RAM
2048~0.5 GB~6 GB
4096~1 GB~6.5 GB
8192~2 GB~7.5 GB
16384~4 GB~9.5 GB
32768~8 GB~13.5 GB
131072~32 GB~37.5 GB

Set context in Modelfile

FROM ./models/llama-3.2-3b-mlx

# Context window
PARAMETER num_ctx 8192

# Giảm context nếu ít RAM
# PARAMETER num_ctx 2048

Set context at runtime

# Override context khi chạy
ollama run llama3.2-mlx --num-ctx 16384

💡 Recommended: Start with 4096, gradually increase until the RAM runs out or enough for the use case.


5. Optimize performance

GPU utilization

# Kiểm tra model dùng GPU hay CPU
ollama ps

Ideal output: 100% GPU. If you see CPU, it means the model does not fit in GPU accessible memory.

Keep model in memory

By default, Ollama unloads the model after 5 minutes of idle time. Change:

# Giữ model loaded 30 phút
export OLLAMA_KEEP_ALIVE=30m

# Giữ vĩnh viễn (cho đến khi restart)
export OLLAMA_KEEP_ALIVE=-1

Or in the API:

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2-mlx",
  "keep_alive": -1
}'

Run multiple models simultaneously

# Cho phép 3 model cùng lúc
export OLLAMA_MAX_LOADED_MODELS=3

# Chạy song song
export OLLAMA_NUM_PARALLEL=4

⚠️ Each model occupies its own RAM. 3 models × 5GB = 15GB. Make sure there is enough RAM for macOS + other apps.


6. Monitoring performance

Activity Monitor

Open Activity Monitor → GPU tab to see:

  • GPU utilization % during inference
  • GPU memory usage

Terminal monitoring

# Xem GPU usage real-time
sudo powermetrics --samplers gpu_power -i 1000

# Xem memory pressure
memory_pressure

# Xem Ollama process
ps aux | grep ollama

Ollama logs

# Xem logs chi tiết
cat ~/.ollama/logs/server.log | tail -50

# Follow logs real-time
tail -f ~/.ollama/logs/server.log

7. Recommended Workflow

Based on practical experience, here is the optimal workflow:

Daily use: Ollama (llama.cpp)

# Model mặc định cho chat, hỏi đáp
ollama run qwen2.5:14b
  • Stable, large ecosystem (Open WebUI, Continue.dev...)
  • Fast enough for interactive chat

When you need speed: Ollama + MLX

# Model MLX cho tác vụ cần nhanh
ollama run qwen2.5-mlx
  • Prompt processing is 2x fast → time-to-first-token is very low
  • Batch processing multiple requests

Scripting/Pipeline: mlx-lm directly

from mlx_lm import load, generate

model, tokenizer = load("mlx-community/Qwen2.5-14B-Instruct-4bit")
# Custom pipeline, batch processing, fine-tuning...
  • Complete control
  • No Ollama server overhead

Summary

ApproachProsDisadvantagesWhen to use
Ollama + llama.cppStable, ecosystemSlower than MLXDaily default
Ollama + MLXFaster, use Ollama APIMore complicated setupNeed speed + API
mlx-lm liveFastest, flexibleNo API serverScripting, pipeline

Exercises

  1. Create Ollama model with MLX backend according to instructions in section 2
  2. Run a benchmark comparing two backends (llama.cpp vs MLX) for the same model
  3. Try increasing the context window: 2048 → 4096 → 8192. Record how much RAM usage increases?
  4. Test OLLAMA_KEEP_ALIVE=-1 — does it affect RAM when idle?
  5. Build a personal workflow: choose model + backend for your 3 daily use cases

Next article: Ollama REST API →