Introduction
Vision models (VLM - Vision Language Models) can "see" and understand images. Thanks to Apple Silicon and Ollama, you run completely locally — no need to send sensitive photos to the cloud.
1. Vision models on Ollama
Popular models
| Model | Size | RAM needed | Features |
|---|---|---|---|
llava:7b | 4.7 GB | 8 GB | First VLM model, stable |
llava:13b | 8.0 GB | 12 GB | More accurate |
llava-llama3 | 5.5 GB | 8 GB | Based on Llama 3 |
gemma3:4b | 3.3 GB | 6 GB | Light, fast, supports vision |
gemma3:12b | 8.1 GB | 12 GB | Good balance |
gemma3:27b | 17 GB | 24 GB | Highest quality |
qwen2.5-vl:7b | 5.3 GB | 10 GB | Good with text/doc |
minicpm-v | 5.5 GB | 8 GB | Compact, strong OCR |
Pull model
# Khuyến nghị cho MacBook 16GB RAM
ollama pull gemma3:12b
# Cho MacBook 8GB RAM
ollama pull gemma3:4b
# Model OCR chuyên dụng
ollama pull minicpm-v
2. Use from CLI
Photo description
# Mô tả ảnh
ollama run gemma3:12b "Describe this image in detail" ./screenshot.png
# Tiếng Việt
ollama run gemma3:12b "Mô tả chi tiết hình ảnh này" ./photo.jpg
Many photos
ollama run gemma3:12b "So sánh 2 bức ảnh này" ./before.png ./after.png
3. Python: image analysis
Basics with Ollama library
import ollama
response = ollama.chat(
model='gemma3:12b',
messages=[{
'role': 'user',
'content': 'Mô tả chi tiết hình ảnh này bằng tiếng Việt',
'images': ['./screenshot.png']
}]
)
print(response['message']['content'])
Use base64
import ollama
import base64
def analyze_image(image_path, prompt="Mô tả hình ảnh này"):
with open(image_path, 'rb') as f:
image_data = base64.b64encode(f.read()).decode('utf-8')
response = ollama.chat(
model='gemma3:12b',
messages=[{
'role': 'user',
'content': prompt,
'images': [image_data]
}]
)
return response['message']['content']
# Sử dụng
result = analyze_image("diagram.png", "Giải thích architecture diagram này")
print(result)
Use API directly
import requests
import base64
def vision_api(image_path, prompt):
with open(image_path, 'rb') as f:
image_b64 = base64.b64encode(f.read()).decode('utf-8')
response = requests.post(
'http://localhost:11434/api/chat',
json={
'model': 'gemma3:12b',
'messages': [{
'role': 'user',
'content': prompt,
'images': [image_b64]
}],
'stream': False
}
)
return response.json()['message']['content']
result = vision_api("error.png", "Đọc error message trong screenshot này")
print(result)
4. Actual use cases
4.1. OCR — Read text from images
import ollama
def ocr_image(image_path):
response = ollama.chat(
model='minicpm-v', # Tốt cho OCR
messages=[{
'role': 'user',
'content': '''Đọc và trích xuất toàn bộ text trong hình ảnh này.
Giữ nguyên format, xuống dòng đúng vị trí.
Chỉ trả về text, không thêm giải thích.''',
'images': [image_path]
}]
)
return response['message']['content']
text = ocr_image("invoice.jpg")
print(text)
4.2. Analyze code from screenshot
def analyze_code_screenshot(image_path):
response = ollama.chat(
model='gemma3:12b',
messages=[{
'role': 'user',
'content': '''Phân tích screenshot code này:
1. Đọc code trong hình
2. Giải thích code làm gì
3. Tìm bugs hoặc issues (nếu có)
4. Đề xuất cải thiện
Trả lời bằng tiếng Việt.''',
'images': [image_path]
}]
)
return response['message']['content']
result = analyze_code_screenshot("code_review.png")
print(result)
4.3. Describe UI/UX
def analyze_ui(image_path):
response = ollama.chat(
model='gemma3:12b',
messages=[{
'role': 'user',
'content': '''Phân tích thiết kế UI này:
1. Layout và bố cục
2. Color scheme
3. Typography
4. UX issues (nếu có)
5. Đề xuất cải thiện
Trả lời theo format markdown.''',
'images': [image_path]
}]
)
return response['message']['content']
4.4. Batch processing of multiple images
import ollama
from pathlib import Path
def batch_analyze(folder_path, prompt="Mô tả hình ảnh"):
image_dir = Path(folder_path)
extensions = {'.jpg', '.jpeg', '.png', '.webp', '.gif'}
results = {}
for img_file in sorted(image_dir.iterdir()):
if img_file.suffix.lower() in extensions:
print(f"📷 Đang phân tích: {img_file.name}...")
response = ollama.chat(
model='gemma3:4b', # Dùng model nhẹ cho batch
messages=[{
'role': 'user',
'content': prompt,
'images': [str(img_file)]
}]
)
results[img_file.name] = response['message']['content']
print(f" ✅ Done\n")
return results
# Phân tích tất cả ảnh trong folder
results = batch_analyze(
"./screenshots",
"Đây là screenshot lỗi. Tóm tắt error message trong 1-2 câu."
)
for filename, analysis in results.items():
print(f"📄 {filename}: {analysis}\n")
5. Compare performance of Vision models
Benchmark on MacBook Pro M3 Pro 18GB RAM:
| Model | Response time | Quality description | OCR accuracy |
|---|---|---|---|
gemma3:4b | ~3s | ⭐⭐⭐ | ⭐⭐⭐ |
gemma3:12b | ~8s | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ |
llava:7b | ~5s | ⭐⭐⭐ | ⭐⭐ |
minicpm-v | ~6s | ⭐⭐⭐ | ⭐⭐⭐⭐⭐ |
qwen2.5-vl:7b | ~7s | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ |
Script benchmark
import ollama
import time
MODELS = ['gemma3:4b', 'gemma3:12b', 'llava:7b']
TEST_IMAGE = './test.png'
PROMPT = 'Describe this image in detail'
for model in MODELS:
print(f"\n--- {model} ---")
start = time.time()
response = ollama.chat(
model=model,
messages=[{
'role': 'user',
'content': PROMPT,
'images': [TEST_IMAGE]
}]
)
elapsed = time.time() - start
content = response['message']['content']
print(f"⏱️ Time: {elapsed:.1f}s")
print(f"📝 Length: {len(content)} chars")
print(f"💬 {content[:200]}...")
6. mlx-vlm — Run Vision models with MLX
If you need higher performance (batch processing, custom pipeline), use mlx-vlm:
pip3 install mlx-vlm
from mlx_vlm import load, generate
# Load model
model, processor = load("mlx-community/Qwen2.5-VL-7B-Instruct-4bit")
# Generate
output = generate(
model,
processor,
"Describe this image",
["./photo.jpg"],
max_tokens=500,
verbose=False,
)
print(output)
Compare Ollama vs mlx-vlm:
| Ollama | mlx-vlm | |
|---|---|---|
| Setup | Easy, pull & run | Need pip install, code |
| API | REST, CLI | Python only |
| Speed | Good | ~20% faster |
| Models | Many options | Less |
| Batch | Each photo | Batch support |
| Use case | General | Production pipeline |
7. Privacy & Security
Vision models local solves many security problems:
- Medical images: X-ray and CT scan analysis does not need to be sent to the cloud
- Confidential documents: OCR contracts, internal invoices
- Screenshot code: Review without revealing the source code
- Camera feed: Process local surveillance videos/photos
# Ví dụ: xử lý tài liệu nội bộ
def process_confidential_doc(image_path):
"""Trích xuất thông tin từ tài liệu mật - hoàn toàn offline."""
response = ollama.chat(
model='minicpm-v',
messages=[{
'role': 'user',
'content': '''Trích xuất thông tin từ tài liệu:
- Ngày tháng
- Tên người/tổ chức
- Số tiền (nếu có)
- Nội dung chính
Format JSON.''',
'images': [image_path]
}]
)
return response['message']['content']
Summary
| Features | Recommended model |
|---|---|
| General photo description | gemma3:12b |
| OCR / text reading | minicpm-v |
| Light, fast | gemma3:4b |
| Analyze documents | qwen2.5-vl:7b |
| Code screenshot | gemma3:12b |
| Batch processing | mlx-vlm + 4bit model |
Exercises
- Pull model
gemma3:4band describe 5 different photos - Build an OCR script to read text from a photo of a book page
- Create a tool to analyze batch error screenshots (5+ images)
- Compare OCR results between
minicpm-vandgemma3:12b - (Bonus) Process images from webcam real-time with OpenCV + Ollama vision
Next article: Optimizing performance — RAM, context, concurrency →