Chuyển đến nội dung chính

Lesson 1: Apple Silicon & AI - Why M-chip is the king of local inference

What is Unified Memory Architecture (UMA) and why is it changing the local AI game? Compare M1/M2/M3/M4 with NVIDIA GPU. Memory bandwidth, Neural Engine, GPU cores. Why does LLM 7B-30B run smoothly on MacBook?

🧠 AI & ML — Lesson 0 Lesson 1: Apple Silicon & AI - Why M-chip is the king of local inference

Running AI Local with Ollama on Apple Silicon

Part 1: Platform - Ollama & Apple Silicon

xdev.asia

Introduction

In 2026, running AI locally on a laptop is no longer just "playing for fun" but a truly useful skill. And Apple Silicon is one of the best local inference platforms today.

This article will help you understand why Apple Silicon is strong for AI, before entering the settings in the next article.


1. Unified Memory Architecture (UMA) — Game changer

On most traditional systems, the CPU and GPU have separate memory areas:

[CPU] ←→ [System RAM 32GB]
  ↕ (PCIe bus - bottleneck)
[GPU] ←→ [VRAM 8-24GB]

When running LLM on an NVIDIA GPU, the model must reside entirely in VRAM. If the 13B model requires 8GB but the GPU only has 8GB of VRAM, you will get an Out of Memory Error.

Apple Silicon uses Unified Memory Architecture:

[CPU cores] ←→
[GPU cores] ←→  [Unified Memory 16-192GB]
[Neural Engine] ←→
[Media Engine] ←→

All CPU, GPU, Neural Engine access the same memory pool. Model LLM 13B requires 8GB? Both the CPU and GPU "see" it without needing to copy data back and forth.

💡 Key insight: On Apple Silicon, you can load a much larger model than a discrete graphics card in the same price range, because system RAM = "VRAM".


2. Memory Bandwidth — Read/write speed determines inference

LLM inference depends greatly on memory bandwidth — the speed of reading weights from RAM. This is the main bottleneck when running LLM.

ChipsMemory BandwidthMaximum RAM
M168.25 GB/s16 GB
M1 Pro200 GB/s32 GB
M1 Max400 GB/s64 GB
M1 Ultra800 GB/s128 GB
M2100 GB/s24 GB
M2 Pro200 GB/s32 GB
M2 Max400 GB/s96 GB
M2 Ultra800 GB/s192 GB
M3100 GB/s24 GB
M3 Pro150 GB/s36 GB
M3 Max400 GB/s128 GB
M4120 GB/s32 GB
M4 Pro273 GB/s48 GB
M4 Max546 GB/s128 GB

Comparison with NVIDIA:

GPUVRAMBandwidthPrice
RTX 40608 GB272 GB/s~$300
RTX 4070 Ti12 GB504 GB/s~$750
RTX 409024 GB1008 GB/s~$1600

NVIDIA has higher bandwidth at the high end, but is limited by VRAM. An RTX 4060 has 272 GB/s bandwidth but only 8GB VRAM — not enough for the 13B model.

Meanwhile, MacBook Pro M3 Pro 36GB: bandwidth 150 GB/s, but can load the 30B quantized model comfortably because all 36GB is available.


3. GPU Cores on Apple Silicon

Apple Silicon integrates the GPU directly on the chip (integrated GPU), but this is a very powerful GPU, not a weak "inte HD graphics" type:

ChipsGPU CoresFP16 TFLOPS
M17-82.6
M2103.6
M3104.1
M3 Pro14-187.4
M3 Max30-4014.2
M4104.6
M4 Pro208.7
M4 Max4017.4

These GPU cores run LLM inference extremely effectively, especially when using Apple's MLX framework (will learn in Part 2).


4. Neural Engine — Hidden gem

Each Apple Silicon chip has a Neural Engine — a dedicated processor for machine learning:

  • M1: 16 cores, 11 TOPS
  • M2: 16 cores, 15.8 TOPS
  • M3: 16 cores, 18 TOPS
  • M4: 16 cores, 38 TOPS

Neural Engine is optimized for matrix calculations typical of neural networks. Currently, most inference frameworks (Ollama, llama.cpp) mainly use GPUs, but Core ML and some new frameworks have begun to take advantage of the Neural Engine.


5. Why does LLM 7B-30B run smoothly on Mac?

Let's calculate specifically:

Model Llama 3.2 8B (Q4 quantized):

  • Size: ~4.5 GB
  • Inference requires ~5.5 GB RAM (model + KV cache + overhead)
  • MacBook Air M2 (8GB RAM): works, but tight
  • MacBook Pro M3 (18GB RAM): runs very comfortably

Model Qwen 2.5 32B (Q4 quantized):

  • Size: ~18 GB
  • Inference requires ~22 GB of RAM
  • MacBook Pro 36 GB: runs smoothly
  • MacBook Air 24 GB: can run but should not open many other apps

Actual generation speed on M3 Pro (36GB):

ModelTokens/secondFeeling
Llama 3.2 3B Q4~65 tok/sExtremely fast
Llama 3.2 8B Q4~32 tok/sVery fast
Gemma 3 12B Q4~22 tok/sFast
Qwen 2.5 14B Q4~18 tok/sGood
Qwen 2.5 32B Q4~8 tok/sAcceptable

💡 Humans read ~250 words/minute ≈ ~5 tokens/second. So 8 tok/s is already faster than reading speed.


6. Overall comparison: Mac vs PC for local AI

CriteriaMac Apple SiliconPC + NVIDIA GPU
Max model sizeDepends on RAM (up to 192GB)Depends on VRAM (usually 8-24GB)
Bandwidth100-800 GB/s272-1000 GB/s
Price/GB "VRAM"Much cheaperExpensive (RTX 4090: $1600 for 24GB)
NoiseSilenceFan running loudly
Electricity15-65WGPU alone 100-450W
MobileThin and light laptopDesktop or heavy gaming laptop
EcosystemmacOS onlyLinux/Windows, more tools

Conclusion: Mac Apple Silicon is a great choice for developers who want to run AI locally on a daily basis. Not the fastest, but the best balance between performance, convenience, noise and portability.


7. Test your Mac

Open Terminal and run:

# Xem chip và RAM
system_profiler SPHardwareDataType | grep -E "(Chip|Memory)"

Sample output:

Chip: Apple M3 Pro
Total Number of Cores: 12 (6 performance and 6 efficiency)
Memory: 36 GB

Or use command:

# Xem GPU cores
system_profiler SPDisplaysDataType | grep "Total Number of Cores"

Remember this information — you'll need it to choose the right model in Lesson 3.


Summary

ConceptsRemember
UMACPU + GPU share RAM → load larger model
Memory BandwidthDeciding on inference speed, M3 Max reaches 400 GB/s
GPU CoresIntegrated but powerful, runs LLM inference effectively
Neural EngineDedicated AI processors, being gradually utilized
Model size sweet spot3B-14B for 16GB RAM, 14B-32B for 32GB+ RAM

Exercises

  1. Check your Mac: what chip, how much RAM, how many GPU cores?
  2. Based on the memory bandwidth table, estimate the maximum size model your Mac can run?
  3. Comparison: if you have the same budget to buy RTX 4070 Ti ($750), would it make more sense to buy more RAM for your Mac or buy a separate GPU?

Next post: Install Ollama and run your first LLM in 5 minutes →