Chuyển đến nội dung chính

Lesson 9: Vision models - Image analysis without cloud

Run LLaVA, Gemma 3 Vision, Qwen VL on Apple Silicon. Image analysis, text reading, OCR, descriptions – all offline. Integrate mlx-vlm for high performance.

🧠 AI & ML — Lesson 2 Lesson 9: Vision models - Image analysis no need for cloud

Running AI Local with Ollama on Apple Silicon

Part 3: API integration & application programming

xdev.asia

Introduction

Vision models (VLM - Vision Language Models) can "see" and understand images. Thanks to Apple Silicon and Ollama, you run completely locally — no need to send sensitive photos to the cloud.


1. Vision models on Ollama

Popular models

ModelSizeRAM neededFeatures
llava:7b4.7 GB8 GBFirst VLM model, stable
llava:13b8.0 GB12 GBMore accurate
llava-llama35.5 GB8 GBBased on Llama 3
gemma3:4b3.3 GB6 GBLight, fast, supports vision
gemma3:12b8.1 GB12 GBGood balance
gemma3:27b17 GB24 GBHighest quality
qwen2.5-vl:7b5.3 GB10 GBGood with text/doc
minicpm-v5.5 GB8 GBCompact, strong OCR

Pull model

# Khuyến nghị cho MacBook 16GB RAM
ollama pull gemma3:12b

# Cho MacBook 8GB RAM
ollama pull gemma3:4b

# Model OCR chuyên dụng
ollama pull minicpm-v

2. Use from CLI

Photo description

# Mô tả ảnh
ollama run gemma3:12b "Describe this image in detail" ./screenshot.png

# Tiếng Việt
ollama run gemma3:12b "Mô tả chi tiết hình ảnh này" ./photo.jpg

Many photos

ollama run gemma3:12b "So sánh 2 bức ảnh này" ./before.png ./after.png

3. Python: image analysis

Basics with Ollama library

import ollama

response = ollama.chat(
    model='gemma3:12b',
    messages=[{
        'role': 'user',
        'content': 'Mô tả chi tiết hình ảnh này bằng tiếng Việt',
        'images': ['./screenshot.png']
    }]
)

print(response['message']['content'])

Use base64

import ollama
import base64

def analyze_image(image_path, prompt="Mô tả hình ảnh này"):
    with open(image_path, 'rb') as f:
        image_data = base64.b64encode(f.read()).decode('utf-8')

    response = ollama.chat(
        model='gemma3:12b',
        messages=[{
            'role': 'user',
            'content': prompt,
            'images': [image_data]
        }]
    )
    return response['message']['content']

# Sử dụng
result = analyze_image("diagram.png", "Giải thích architecture diagram này")
print(result)

Use API directly

import requests
import base64

def vision_api(image_path, prompt):
    with open(image_path, 'rb') as f:
        image_b64 = base64.b64encode(f.read()).decode('utf-8')

    response = requests.post(
        'http://localhost:11434/api/chat',
        json={
            'model': 'gemma3:12b',
            'messages': [{
                'role': 'user',
                'content': prompt,
                'images': [image_b64]
            }],
            'stream': False
        }
    )
    return response.json()['message']['content']

result = vision_api("error.png", "Đọc error message trong screenshot này")
print(result)

4. Actual use cases

4.1. OCR — Read text from images

import ollama

def ocr_image(image_path):
    response = ollama.chat(
        model='minicpm-v',  # Tốt cho OCR
        messages=[{
            'role': 'user',
            'content': '''Đọc và trích xuất toàn bộ text trong hình ảnh này.
Giữ nguyên format, xuống dòng đúng vị trí.
Chỉ trả về text, không thêm giải thích.''',
            'images': [image_path]
        }]
    )
    return response['message']['content']

text = ocr_image("invoice.jpg")
print(text)

4.2. Analyze code from screenshot

def analyze_code_screenshot(image_path):
    response = ollama.chat(
        model='gemma3:12b',
        messages=[{
            'role': 'user',
            'content': '''Phân tích screenshot code này:
1. Đọc code trong hình
2. Giải thích code làm gì
3. Tìm bugs hoặc issues (nếu có)
4. Đề xuất cải thiện

Trả lời bằng tiếng Việt.''',
            'images': [image_path]
        }]
    )
    return response['message']['content']

result = analyze_code_screenshot("code_review.png")
print(result)

4.3. Describe UI/UX

def analyze_ui(image_path):
    response = ollama.chat(
        model='gemma3:12b',
        messages=[{
            'role': 'user',
            'content': '''Phân tích thiết kế UI này:
1. Layout và bố cục
2. Color scheme
3. Typography
4. UX issues (nếu có)
5. Đề xuất cải thiện

Trả lời theo format markdown.''',
            'images': [image_path]
        }]
    )
    return response['message']['content']

4.4. Batch processing of multiple images

import ollama
from pathlib import Path

def batch_analyze(folder_path, prompt="Mô tả hình ảnh"):
    image_dir = Path(folder_path)
    extensions = {'.jpg', '.jpeg', '.png', '.webp', '.gif'}

    results = {}
    for img_file in sorted(image_dir.iterdir()):
        if img_file.suffix.lower() in extensions:
            print(f"📷 Đang phân tích: {img_file.name}...")
            response = ollama.chat(
                model='gemma3:4b',  # Dùng model nhẹ cho batch
                messages=[{
                    'role': 'user',
                    'content': prompt,
                    'images': [str(img_file)]
                }]
            )
            results[img_file.name] = response['message']['content']
            print(f"   ✅ Done\n")

    return results

# Phân tích tất cả ảnh trong folder
results = batch_analyze(
    "./screenshots",
    "Đây là screenshot lỗi. Tóm tắt error message trong 1-2 câu."
)

for filename, analysis in results.items():
    print(f"📄 {filename}: {analysis}\n")

5. Compare performance of Vision models

Benchmark on MacBook Pro M3 Pro 18GB RAM:

ModelResponse timeQuality descriptionOCR accuracy
gemma3:4b~3s⭐⭐⭐⭐⭐⭐
gemma3:12b~8s⭐⭐⭐⭐⭐⭐⭐⭐
llava:7b~5s⭐⭐⭐⭐⭐
minicpm-v~6s⭐⭐⭐⭐⭐⭐⭐⭐
qwen2.5-vl:7b~7s⭐⭐⭐⭐⭐⭐⭐⭐

Script benchmark

import ollama
import time

MODELS = ['gemma3:4b', 'gemma3:12b', 'llava:7b']
TEST_IMAGE = './test.png'
PROMPT = 'Describe this image in detail'

for model in MODELS:
    print(f"\n--- {model} ---")
    start = time.time()
    response = ollama.chat(
        model=model,
        messages=[{
            'role': 'user',
            'content': PROMPT,
            'images': [TEST_IMAGE]
        }]
    )
    elapsed = time.time() - start
    content = response['message']['content']
    print(f"⏱️  Time: {elapsed:.1f}s")
    print(f"📝 Length: {len(content)} chars")
    print(f"💬 {content[:200]}...")

6. mlx-vlm — Run Vision models with MLX

If you need higher performance (batch processing, custom pipeline), use mlx-vlm:

pip3 install mlx-vlm
from mlx_vlm import load, generate

# Load model
model, processor = load("mlx-community/Qwen2.5-VL-7B-Instruct-4bit")

# Generate
output = generate(
    model,
    processor,
    "Describe this image",
    ["./photo.jpg"],
    max_tokens=500,
    verbose=False,
)
print(output)

Compare Ollama vs mlx-vlm:

Ollamamlx-vlm
SetupEasy, pull & runNeed pip install, code
APIREST, CLIPython only
SpeedGood~20% faster
ModelsMany optionsLess
BatchEach photoBatch support
Use caseGeneralProduction pipeline

7. Privacy & Security

Vision models local solves many security problems:

  • Medical images: X-ray and CT scan analysis does not need to be sent to the cloud
  • Confidential documents: OCR contracts, internal invoices
  • Screenshot code: Review without revealing the source code
  • Camera feed: Process local surveillance videos/photos
# Ví dụ: xử lý tài liệu nội bộ
def process_confidential_doc(image_path):
    """Trích xuất thông tin từ tài liệu mật - hoàn toàn offline."""
    response = ollama.chat(
        model='minicpm-v',
        messages=[{
            'role': 'user',
            'content': '''Trích xuất thông tin từ tài liệu:
- Ngày tháng
- Tên người/tổ chức
- Số tiền (nếu có)
- Nội dung chính

Format JSON.''',
            'images': [image_path]
        }]
    )
    return response['message']['content']

Summary

FeaturesRecommended model
General photo descriptiongemma3:12b
OCR / text readingminicpm-v
Light, fastgemma3:4b
Analyze documentsqwen2.5-vl:7b
Code screenshotgemma3:12b
Batch processingmlx-vlm + 4bit model

Exercises

  1. Pull model gemma3:4b and describe 5 different photos
  2. Build an OCR script to read text from a photo of a book page
  3. Create a tool to analyze batch error screenshots (5+ images)
  4. Compare OCR results between minicpm-v and gemma3:12b
  5. (Bonus) Process images from webcam real-time with OpenCV + Ollama vision

Next article: Optimizing performance — RAM, context, concurrency →