Chuyển đến nội dung chính

Bài 4: MLX Framework - Apple Intelligence dưới nắp capo

MLX là gì, tại sao Apple tạo ra nó. Kiến trúc lazy evaluation, unified computation graph. So sánh MLX vs llama.cpp vs Core ML. Benchmarks thực tế trên M1/M2/M3/M4.

🧠 AI & ML — Bài 0 Bài 4: MLX Framework - Apple Intelligence dưới nắp capo

Chạy AI Local với Ollama trên Apple Silicon

Phần 2: MLX - Tăng tốc 3x với framework native của Apple

xdev.asia

Giới thiệu

Ollama dùng llama.cpp để chạy LLM — và đó đã là rất nhanh. Nhưng Apple có một "vũ khí bí mật": MLX — framework machine learning được thiết kế riêng cho Apple Silicon, tận dụng mọi đặc tính của unified memory.

Kết quả? Inference nhanh hơn 2-3x so với llama.cpp trên cùng phần cứng.


1. MLX là gì?

MLX là framework mã nguồn mở của Apple Research, ra mắt cuối 2023. Nó được thiết kế cho Apple Silicon, giống như PyTorch được thiết kế cho NVIDIA CUDA.

Đặc điểm chính

  • Unified Memory: Model, data và computation share cùng memory — zero copy overhead
  • Lazy Evaluation: Chỉ tính toán khi cần, tối ưu tự động computation graph
  • Dynamic Shapes: Không cần compile trước cho từng tensor shape
  • NumPy-like API: Quen thuộc nếu bạn biết NumPy/PyTorch
  • Multi-device: Tự động tận dụng GPU, CPU, Neural Engine

MLX vs các framework khác

FeatureMLXllama.cppPyTorch (MPS)Core ML
TargetApple SiliconCross-platformCross-platformApple only
BackendMetalMetal/CPUMPSANE + GPU
Unified Memory aware✅ Native❌ Ported❌ Ported✅ Native
Ease of use★★★★★★★★☆★★★★★★☆☆
LLM inference speedNhanh nhấtNhanhChậmNhanh (but limited)
Model ecosystemHuggingFace MLXGGUFPyTorchCoreML models
Training support✅❌✅❌

2. Tại sao MLX nhanh hơn?

Zero-copy memory access

Trên llama.cpp chạy qua Metal:

[CPU loads model] → [Copy to GPU buffer] → [GPU compute] → [Copy result back]

Trên MLX:

[Load model to unified memory] → [GPU compute directly] → [Result already accessible]

Không có bước copy data qua lại giữa CPU và GPU. Trên Apple Silicon, cả hai đều truy cập cùng physical memory.

Lazy evaluation graph

import mlx.core as mx

# Không tính ngay!
a = mx.array([1, 2, 3])
b = mx.array([4, 5, 6])
c = a + b        # Chưa tính
d = c * 2         # Chưa tính
result = d.sum()  # Chưa tính

# Chỉ tính khi cần giá trị
mx.eval(result)   # Bây giờ mới tính tất cả, tối ưu tự động

MLX gom tất cả operations thành một computation graph và tối ưu trước khi thực thi. Trong LLM inference, điều này tạo ra sự khác biệt lớn vì mỗi token generation cần hàng nghìn matrix operations.

Metal shader optimization

MLX dùng custom Metal shaders được tối ưu riêng cho Apple GPU architecture, đặc biệt cho quantized matrix multiplication — operation phổ biến nhất trong LLM inference.


3. Cài đặt MLX

Prerequisites

# Python 3.9+ (khuyến nghị 3.11+)
python3 --version

# pip
pip3 --version

Cài MLX core

pip3 install mlx

Kiểm tra cài đặt

python3 -c "
import mlx.core as mx
print(f'MLX version: {mx.__version__}')
print(f'Default device: {mx.default_device()}')
a = mx.array([1.0, 2.0, 3.0])
print(f'Test: {a * 2}')
"

Output mong đợi:

MLX version: 0.x.x
Default device: Device(gpu, 0)
Test: array([2, 4, 6], dtype=float32)

💡 Device(gpu, 0) nghĩa là MLX tự động dùng Apple GPU. Không cần config gì thêm.


4. MLX core concepts

Arrays

import mlx.core as mx

# Tạo array (giống NumPy)
a = mx.array([1, 2, 3, 4])
b = mx.zeros((3, 4))
c = mx.random.normal((2, 3))

# Operations
d = mx.matmul(c, b[:, :3].T)  # Matrix multiplication

# Dtype
e = mx.array([1.0, 2.0], dtype=mx.float16)

Device placement

# MLX tự động dùng GPU
# Nhưng bạn có thể chỉ định:
with mx.stream(mx.cpu):
    result_cpu = mx.matmul(a, b)  # Chạy trên CPU

with mx.stream(mx.gpu):
    result_gpu = mx.matmul(a, b)  # Chạy trên GPU

Lazy eval trong thực tế

import mlx.core as mx
import time

# Tạo matrix lớn
a = mx.random.normal((4096, 4096))
b = mx.random.normal((4096, 4096))

# Chưa tính!
c = mx.matmul(a, b)
print(type(c))  # <class 'mlx.core.array'>

# Tính khi cần
start = time.time()
mx.eval(c)
print(f"Matmul 4096x4096: {time.time() - start:.4f}s")

5. Benchmark: MLX vs llama.cpp

Benchmark trên MacBook Pro M3 Pro (36GB RAM) với Llama 3.2 8B Q4:

Metricllama.cpp (Ollama)MLX (mlx-lm)Speedup
Prompt processing280 tok/s650 tok/s2.3x
Token generation32 tok/s58 tok/s1.8x
Time to first token0.45s0.19s2.4x
Memory usage6.5 GB5.8 GB-11%

Trên M4 Max (128GB):

Metricllama.cppMLXSpeedup
Prompt (8B Q4)680 tok/s1450 tok/s2.1x
Generation (8B Q4)65 tok/s105 tok/s1.6x
Generation (32B Q4)22 tok/s42 tok/s1.9x

💡 MLX đặc biệt nhanh hơn ở prompt processing (prefill). Điều này tạo cảm giác "phản hồi tức thì" khi chat.


6. Khi nào dùng MLX, khi nào dùng llama.cpp?

Tình huốngKhuyến nghị
Cần tốc độ tối đa trên MacMLX
Cần OpenAI-compatible APIOllama (llama.cpp)
Cần UI đẹp (Open WebUI...)Ollama
Chạy trên Linux/Windowsllama.cpp
Training/fine-tuning localMLX
Quick prototypingMLX
Production serverOllama (ổn định hơn)

Tin vui: Bạn có thể dùng cả hai! Ollama cho daily use, MLX cho khi cần tốc độ hoặc custom pipeline. Bài 6 sẽ hướng dẫn kết hợp.


7. MLX Ecosystem

MLX không chỉ có core library. Apple và cộng đồng đã xây dựng ecosystem phong phú:

PackageMô tả
mlxCore array framework
mlx-lmLLM inference & fine-tuning
mlx-vlmVision-Language models
mlx-whisperSpeech-to-text
mlx-audioText-to-speech
mlx-imageImage generation (Stable Diffusion)

Hugging Face có cộng đồng MLX Community chuyên convert model sang format MLX:


Tóm tắt

ConceptGhi nhớ
MLXFramework ML của Apple, thiết kế cho Apple Silicon
Zero-copyKhông cần copy data CPU↔GPU nhờ unified memory
Lazy evalGom operations, tối ưu graph trước khi tính
SpeedupNhanh hơn llama.cpp ~1.5-2.5x trên Mac
Ecosystemmlx-lm, mlx-vlm, mlx-whisper, mlx-audio

Bài tập

  1. Cài mlx và chạy test cơ bản: tạo array, matmul, kiểm tra device
  2. Benchmark matmul với kích thước tăng dần (1024, 2048, 4096, 8192) — vẽ biểu đồ thời gian
  3. So sánh mx.array với numpy.array cho cùng phép tính — cái nào nhanh hơn?

Bài tiếp theo: Cài mlx-lm và chạy model MLX-quantized →