Chuyển đến nội dung chính

Lesson 4: MLX Framework - Apple Intelligence under the hood

What is MLX, why did Apple create it? Lazy evaluation architecture, unified computation graph. Compare MLX vs llama.cpp vs Core ML. Actual benchmarks on M1/M2/M3/M4.

🧠 AI & ML — Lesson 0 Lesson 4: MLX Framework - Apple Intelligence under the hood

Running AI Local with Ollama on Apple Silicon

Part 2: MLX - 3x acceleration with Apple's native framework

xdev.asia

Introduction

Ollama uses llama.cpp to run LLM — and that's already very fast. But Apple has a "secret weapon": MLX — a machine learning framework designed specifically for Apple Silicon, taking advantage of all the features of unified memory.

Result? Inference is 2-3x faster than llama.cpp on the same hardware.


1. What is MLX?

MLX is an open source framework from Apple Research, released late 2023. It is designed for Apple Silicon, just like PyTorch is designed for NVIDIA CUDA.

Main features

  • Unified Memory: Model, data and computation share the same memory — zero copy overhead
  • Lazy Evaluation: Only calculate when needed, automatically optimize the computation graph
  • Dynamic Shapes: No need to pre-compile each shape tensor
  • NumPy-like API: Familiar if you know NumPy/PyTorch
  • Multi-device: Automatically takes advantage of GPU, CPU, Neural Engine

MLX vs other frameworks

FeaturesMLXllama.cppPyTorch (MPS)Core ML
TargetApple SiliconCross-platformCross-platformApple only
BackendMetalMetal/CPUMPSANE + GPU
Unified Memory aware✅ Native❌ Ported❌ Ported✅ Native
Ease of use★★★★★★★★☆★★★★★★☆☆
LLM inference speedFastestFastSlowFast (but limited)
Model ecosystemHuggingFace MLXGGUFPyTorchCoreML models
Training support✅❌✅❌

2. Why is MLX faster?

Zero-copy memory access

On llama.cpp run through Metal:

[CPU loads model] → [Copy to GPU buffer] → [GPU compute] → [Copy result back]

On MLX:

[Load model to unified memory] → [GPU compute directly] → [Result already accessible]

There is no step to copy data back and forth between CPU and GPU. On Apple Silicon, both access the same physical memory.

Lazy evaluation graph

import mlx.core as mx

# Không tính ngay!
a = mx.array([1, 2, 3])
b = mx.array([4, 5, 6])
c = a + b        # Chưa tính
d = c * 2         # Chưa tính
result = d.sum()  # Chưa tính

# Chỉ tính khi cần giá trị
mx.eval(result)   # Bây giờ mới tính tất cả, tối ưu tự động

MLX collects all operations into a computation graph and optimizes them before execution. In LLM inference, this makes a big difference because each token generation requires thousands of matrix operations.

Metal shader optimization

MLX uses custom Metal shaders specifically optimized for Apple GPU architecture, especially for quantized matrix multiplication — the most common operation in LLM inference.


3. Install MLX

Prerequisites

# Python 3.9+ (khuyến nghị 3.11+)
python3 --version

# pip
pip3 --version

Install MLX core

pip3 install mlx

Check settings

python3 -c "
import mlx.core as mx
print(f'MLX version: {mx.__version__}')
print(f'Default device: {mx.default_device()}')
a = mx.array([1.0, 2.0, 3.0])
print(f'Test: {a * 2}')
"

Expected output:

MLX version: 0.x.x
Default device: Device(gpu, 0)
Test: array([2, 4, 6], dtype=float32)

💡 Device(gpu, 0) meaning MLX automatically uses Apple GPU. No need for any further configuration.


4. MLX core concepts

Arrays

import mlx.core as mx

# Tạo array (giống NumPy)
a = mx.array([1, 2, 3, 4])
b = mx.zeros((3, 4))
c = mx.random.normal((2, 3))

# Operations
d = mx.matmul(c, b[:, :3].T)  # Matrix multiplication

# Dtype
e = mx.array([1.0, 2.0], dtype=mx.float16)

Device placement

# MLX tự động dùng GPU
# Nhưng bạn có thể chỉ định:
with mx.stream(mx.cpu):
    result_cpu = mx.matmul(a, b)  # Chạy trên CPU

with mx.stream(mx.gpu):
    result_gpu = mx.matmul(a, b)  # Chạy trên GPU

Lazy eval in practice

import mlx.core as mx
import time

# Tạo matrix lớn
a = mx.random.normal((4096, 4096))
b = mx.random.normal((4096, 4096))

# Chưa tính!
c = mx.matmul(a, b)
print(type(c))  # <class 'mlx.core.array'>

# Tính khi cần
start = time.time()
mx.eval(c)
print(f"Matmul 4096x4096: {time.time() - start:.4f}s")

5. Benchmark: MLX vs llama.cpp

Benchmark on MacBook Pro M3 Pro (36GB RAM) with Llama 3.2 8B Q4:

Metricsllama.cpp (Ollama)MLX (mlx-lm)Speedup
Prompt processing280 tok/s650 tok/s2.3x
Token generation32 tok/s58 tok/s1.8x
Time to first token0.45s0.19s2.4x
Memory usage6.5 GB5.8 GB-11%

On M4 Max (128GB):

Metricsllama.cppMLXSpeedup
Prompt (8B Q4)680 tok/s1450 tok/s2.1x
Generation (8B Q4)65 tok/s105 tok/s1.6x
Generation (32B Q4)22 tok/s42 tok/s1.9x

💡 MLX is especially faster at prompt processing (prefill). This creates a feeling of "instant feedback" when chatting.


6. When to use MLX, when to use llama.cpp?

SituationRecommendations
Need maximum speed on MacMLX
Needs OpenAI-compatible APIOllama (llama.cpp)
Need beautiful UI (Open WebUI...)Ollama
Runs on Linux/Windowsllama.cpp
Training/fine-tuning localMLX
Quick prototypingMLX
Production serverOllama (more stable)

Good news: You can use both! Ollama for daily use, MLX for when speed or custom pipeline is needed. Lesson 6 will guide the combination.


7. MLX Ecosystem

MLX doesn't just have a core library. Apple and the community have built a rich ecosystem:

PackagesDescription
mlxCore array framework
mlx-lmLLM inference & fine-tuning
mlx-vlmVision-Language models
mlx-whisperSpeech-to-text
mlx-audioText-to-speech
mlx-imageImage generation (Stable Diffusion)

Hugging Face has a community MLX Community specializing in converting models to MLX format:


Summary

ConceptsRemember
MLXApple's ML framework, designed for Apple Silicon
Zero-copyNo need to copy CPU↔GPU data thanks to unified memory
Lazy evalCollect operations, optimize graph before calculating
SpeedupFaster than llama.cpp ~1.5-2.5x on Mac
Ecosystemmlx-lm, mlx-vlm, mlx-whisper, mlx-audio

Exercises

  1. Install mlx and run basic tests: create array, matmul, check device
  2. Benchmark matmul with increasing size (1024, 2048, 4096, 8192) — plot time chart
  3. Compare mx.array with numpy.array for the same calculation — which is faster?

Next article: Install mlx-lm and run MLX-quantized model →