Chuyển đến nội dung chính

Lesson 2: Installing Ollama - From zero to running LLM in 5 minutes

Install Ollama on macOS, understand folder structure and model management. Pull and run Llama 3.2, Gemma 3, Mistral, Qwen 2.5. Important Ollama CLI commands: run, pull, list, rm, show, ps.

🧠 AI & ML — Lesson 1 Lesson 2: Installing Ollama - From zero to running LLM in 5 minutes

Running AI Local with Ollama on Apple Silicon

Part 1: Platform - Ollama & Apple Silicon

xdev.asia

Introduction

In the previous article, you understood why Apple Silicon is strong for AI. Now let's turn theory into reality: install Ollama and chat with LLM right on your computer, no internet needed, no API key needed.

Goal of this lesson: After 5 minutes, you will be chatting with an AI model running entirely on your computer.


1. What is Ollama?

Ollama is the easiest tool to run LLM locally today. Think of it as "Docker for LLMs":

  • Pull model from registry like pulling Docker images
  • Run model with a single command
  • Expose API compatible with OpenAI endpoint
  • Manage multiple models at the same time

Ollama uses llama.cpp (an inference engine written in C++) underneath, but wraps it up into a super simple experience.


2. Install Ollama on macOS

Method 1: Download from website (recommended)

# Truy cập https://ollama.com/download và tải bản macOS
# Hoặc dùng curl:
curl -fsSL https://ollama.com/install.sh | sh

Method 2: Use Homebrew

brew install ollama

Confirm installation

ollama --version

Output:

ollama version is 0.6.x

Start the Ollama server

If installed from .dmg, Ollama app will automatically run the server when opened. If installing from CLI:

# Chạy server (giữ terminal này mở)
ollama serve

Server runs on http://localhost:11434.

Check:

curl http://localhost:11434
# Output: Ollama is running

3. Run the first model

Pull and run Llama 3.2

# Pull model (chỉ cần lần đầu, ~2GB cho 3B, ~4.5GB cho 8B)
ollama pull llama3.2

# Chạy và chat
ollama run llama3.2

You will see the prompt:

>>> Send a message (/? for help)

Try asking:

>>> Giải thích Docker trong 3 câu

AI will respond immediately, running entirely on your computer. Press Ctrl+D to escape.

Pull another model

# Gemma 3 - model của Google, mạnh với tiếng Việt
ollama pull gemma3:4b

# Qwen 2.5 - model của Alibaba, đa ngôn ngữ xuất sắc
ollama pull qwen2.5:7b

# Mistral - model của Pháp, code tốt
ollama pull mistral

# Phi-4 - model nhỏ của Microsoft, hiệu quả
ollama pull phi4-mini

4. Important Ollama CLI Commands

View the list of loaded models

ollama list

Sample output:

NAME                ID              SIZE      MODIFIED
llama3.2:latest     a80c4f17acd5    2.0 GB    2 minutes ago
gemma3:4b           2d2a94b1e3fc    3.3 GB    5 minutes ago
qwen2.5:7b          845dbda0ea48    4.7 GB    8 minutes ago

View detailed model information

ollama show llama3.2

Output displays:

  • Architecture (LlamaForCausalLM)
  • Parameters (3.2B)
  • Quantization (Q4_K_M)
  • Context length (128K)
  • System prompt default

See the running model

ollama ps

Output:

NAME              ID            SIZE     PROCESSOR    UNTIL
llama3.2:latest   a80c4f17acd5  3.2 GB   100% GPU     4 minutes from now

💡 100% GPU means the entire model is on GPU memory (Metal on Mac). This is the ideal case.

Delete model

# Xóa một model để giải phóng ổ cứng
ollama rm mistral

Copy model (create a copy with a different name)

ollama cp llama3.2 my-assistant

5. Ollama folder structure

Ollama hosts everything at:

~/.ollama/
├── models/
│   ├── blobs/        # Model weights (file lớn)
│   └── manifests/    # Metadata cho mỗi model
└── logs/             # Logs

Check capacity:

du -sh ~/.ollama/models

⚠️ Note: Heavy model! 3-5 models can take up 20-30GB. If the SSD is small, select the necessary model.

Move the models folder to an external drive

If the main drive is small:

# Dừng Ollama
# Di chuyển thư mục
mv ~/.ollama/models /Volumes/ExternalSSD/ollama-models

# Tạo symlink
ln -s /Volumes/ExternalSSD/ollama-models ~/.ollama/models

# Khởi động lại Ollama

6. Advanced chat with Ollama

System prompt inline

ollama run llama3.2 "Bạn là một chuyên gia Python. Trả lời bằng tiếng Việt." \
  --system "You are a senior Python developer who explains things simply in Vietnamese."

Multi-line input

In chat mode, use """ to enter multiple lines:

>>> """
... Phân tích đoạn code sau:
... def fibonacci(n):
...     if n <= 1: return n
...     return fibonacci(n-1) + fibonacci(n-2)
... """

Set temperature and context

# Temperature thấp = ít sáng tạo, chính xác hơn
ollama run llama3.2 --temperature 0.1

# Context window lớn hơn (tốn RAM hơn)
ollama run llama3.2 --num-ctx 8192

Slash commands in chat

CommandDescription
/set system <prompt>Set system prompt
/show infoView model information
/show modelfileSee Modelfile
/clearDelete chat history
/bye or Ctrl+DExit
/?See help

7. Useful Environment Variables

# Thay đổi host/port
export OLLAMA_HOST=0.0.0.0:11434

# Thay đổi thư mục lưu model
export OLLAMA_MODELS=/path/to/models

# Giới hạn số model load đồng thời
export OLLAMA_MAX_LOADED_MODELS=2

# Bật debug logging
export OLLAMA_DEBUG=1

Added ~/.zshrc to save permanently:

echo 'export OLLAMA_MAX_LOADED_MODELS=2' >> ~/.zshrc
source ~/.zshrc

8. Troubleshooting is common

"Error: model requires more memory than available"

Model is too large for RAM. Solution:

  • Use a smaller model: llama3.2:3b instead llama3.2:8b
  • Close other apps to free up RAM

Unusually slow speed

# Kiểm tra GPU utilization
ollama ps
# Nếu thấy "100% CPU" thay vì "100% GPU" → model quá lớn, không fit GPU memory

Ollama server did not start

# Kiểm tra port xem có bị chiếm không
lsof -i :11434

# Kill process cũ nếu cần
pkill ollama
ollama serve

Summary

CommandDescription
ollama pull <model>Download model
ollama run <model>Run and chat
ollama listView downloaded models
ollama psView running models
ollama show <model>Model information
ollama rm <model>Delete model
ollama serveStart the server

Exercises

  1. Install Ollama and drag 2 models: llama3.2 and gemma3:4b
  2. Chat with each model asking the same question and compare the quality of the answers
  3. Use ollama show See information for both models: quantization, parameter count, context length
  4. Check ollama ps while chatting — how much RAM does the model use? GPU or CPU?
  5. Use du -sh ~/.ollama/models See how much space the model takes up

Next article: Which model to choose for your use case? →