簡介
視覺模型(VLM - 視覺語言模型)可以「看到」並理解圖像。借助 Apple Silicon 和 Ollama,您可以完全在本地運行 — 無需將敏感照片發送到雲端。
1. Ollama 上的視覺模型
熱門車型
| 型號 | 尺寸 | 需要內存 | 特點 |
|---|---|---|---|
llava:7b | 4.7GB | 8GB | 第一個 VLM 模型,穩定 |
llava:13b | 8.0GB | 12GB | 更準確 |
llava-llama3 | 5.5 GB | 5.5 GB 8GB | 基於羊駝 3 |
gemma3:4b | 3.3 GB | 3.3 GB 6GB | 輕巧、快速、支援視力 |
gemma3:12b | 8.1GB | 12GB | 良好的平衡性 |
gemma3:27b | 17GB | 24GB | 最高品質 |
qwen2.5-vl:7b | 5.3GB | 10GB | 擅長文字/文件 |
minicpm-v | 5.5 GB | 5.5 GB 8GB | 小巧、強大的 OCR |
拉模型
# Khuyến nghị cho MacBook 16GB RAM
ollama pull gemma3:12b
# Cho MacBook 8GB RAM
ollama pull gemma3:4b
# Model OCR chuyên dụng
ollama pull minicpm-v
2. 從 CLI 使用
圖片說明
# Mô tả ảnh
ollama run gemma3:12b "Describe this image in detail" ./screenshot.png
# Tiếng Việt
ollama run gemma3:12b "Mô tả chi tiết hình ảnh này" ./photo.jpg
很多照片
ollama run gemma3:12b "So sánh 2 bức ảnh này" ./before.png ./after.png
3. Python:影像分析
Ollama 庫的基礎知識
import ollama
response = ollama.chat(
model='gemma3:12b',
messages=[{
'role': 'user',
'content': 'Mô tả chi tiết hình ảnh này bằng tiếng Việt',
'images': ['./screenshot.png']
}]
)
print(response['message']['content'])
使用base64
import ollama
import base64
def analyze_image(image_path, prompt="Mô tả hình ảnh này"):
with open(image_path, 'rb') as f:
image_data = base64.b64encode(f.read()).decode('utf-8')
response = ollama.chat(
model='gemma3:12b',
messages=[{
'role': 'user',
'content': prompt,
'images': [image_data]
}]
)
return response['message']['content']
# Sử dụng
result = analyze_image("diagram.png", "Giải thích architecture diagram này")
print(result)
直接使用API
import requests
import base64
def vision_api(image_path, prompt):
with open(image_path, 'rb') as f:
image_b64 = base64.b64encode(f.read()).decode('utf-8')
response = requests.post(
'http://localhost:11434/api/chat',
json={
'model': 'gemma3:12b',
'messages': [{
'role': 'user',
'content': prompt,
'images': [image_b64]
}],
'stream': False
}
)
return response.json()['message']['content']
result = vision_api("error.png", "Đọc error message trong screenshot này")
print(result)
4. 實際用例
4.1。 OCR — 從圖像中讀取文本
import ollama
def ocr_image(image_path):
response = ollama.chat(
model='minicpm-v', # Tốt cho OCR
messages=[{
'role': 'user',
'content': '''Đọc và trích xuất toàn bộ text trong hình ảnh này.
Giữ nguyên format, xuống dòng đúng vị trí.
Chỉ trả về text, không thêm giải thích.''',
'images': [image_path]
}]
)
return response['message']['content']
text = ocr_image("invoice.jpg")
print(text)
4.2。從截圖分析程式碼
def analyze_code_screenshot(image_path):
response = ollama.chat(
model='gemma3:12b',
messages=[{
'role': 'user',
'content': '''Phân tích screenshot code này:
1. Đọc code trong hình
2. Giải thích code làm gì
3. Tìm bugs hoặc issues (nếu có)
4. Đề xuất cải thiện
Trả lời bằng tiếng Việt.''',
'images': [image_path]
}]
)
return response['message']['content']
result = analyze_code_screenshot("code_review.png")
print(result)
4.3。描述使用者介面/使用者體驗
def analyze_ui(image_path):
response = ollama.chat(
model='gemma3:12b',
messages=[{
'role': 'user',
'content': '''Phân tích thiết kế UI này:
1. Layout và bố cục
2. Color scheme
3. Typography
4. UX issues (nếu có)
5. Đề xuất cải thiện
Trả lời theo format markdown.''',
'images': [image_path]
}]
)
return response['message']['content']
4.4。多幅影像的批次處理
import ollama
from pathlib import Path
def batch_analyze(folder_path, prompt="Mô tả hình ảnh"):
image_dir = Path(folder_path)
extensions = {'.jpg', '.jpeg', '.png', '.webp', '.gif'}
results = {}
for img_file in sorted(image_dir.iterdir()):
if img_file.suffix.lower() in extensions:
print(f"📷 Đang phân tích: {img_file.name}...")
response = ollama.chat(
model='gemma3:4b', # Dùng model nhẹ cho batch
messages=[{
'role': 'user',
'content': prompt,
'images': [str(img_file)]
}]
)
results[img_file.name] = response['message']['content']
print(f" ✅ Done\n")
return results
# Phân tích tất cả ảnh trong folder
results = batch_analyze(
"./screenshots",
"Đây là screenshot lỗi. Tóm tắt error message trong 1-2 câu."
)
for filename, analysis in results.items():
print(f"📄 {filename}: {analysis}\n")
5. 比較 Vision 模型的效能
MacBook Pro M3 Pro 18GB RAM 的基準測試:
| 型號 | 回應時間 | 品質說明 | OCR 準確率 |
|---|---|---|---|
gemma3:4b | 〜3秒 | ⭐⭐⭐ | ⭐⭐⭐ |
gemma3:12b | 〜8秒 | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ |
llava:7b | 〜5秒 | ⭐⭐⭐ | ⭐⭐ |
minicpm-v | 〜6秒 | ⭐⭐⭐ | ⭐⭐⭐⭐⭐ |
qwen2.5-vl:7b | 〜7秒 | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ |
腳本基準測試
import ollama
import time
MODELS = ['gemma3:4b', 'gemma3:12b', 'llava:7b']
TEST_IMAGE = './test.png'
PROMPT = 'Describe this image in detail'
for model in MODELS:
print(f"\n--- {model} ---")
start = time.time()
response = ollama.chat(
model=model,
messages=[{
'role': 'user',
'content': PROMPT,
'images': [TEST_IMAGE]
}]
)
elapsed = time.time() - start
content = response['message']['content']
print(f"⏱️ Time: {elapsed:.1f}s")
print(f"📝 Length: {len(content)} chars")
print(f"💬 {content[:200]}...")
6. mlx-vlm — 使用 MLX 運行 Vision 模型
如果您需要更高的效能(批次、自訂管道),請使用 mlx-vlm:
pip3 install mlx-vlm
from mlx_vlm import load, generate
# Load model
model, processor = load("mlx-community/Qwen2.5-VL-7B-Instruct-4bit")
# Generate
output = generate(
model,
processor,
"Describe this image",
["./photo.jpg"],
max_tokens=500,
verbose=False,
)
print(output)
比較 Ollama 與 mlx-vlm:
| 奧拉瑪 | MLX-VLM | |
|---|---|---|
| 設定 | 輕鬆拉動即可運行 | 需要 pip 安裝,代碼 |
| API | 休息,CLI | 僅限Python |
| 速度 | 好 | 快約 20% |
| 型號 | 多種選擇 | 少 |
| 批次 | 每張照片 | 批量支援 |
| 用例 | 一般 | 生產管線 |
7. 隱私與安全
本地視覺模型解決了許多安全問題:
- 醫學影像:X光和CT掃描分析不需要傳送到雲端
- 機密文件:OCR 合約、內部發票
- 截圖程式碼:在不透露原始程式碼的情況下查看
- 攝影機饋送:處理本地監控影片/照片
# Ví dụ: xử lý tài liệu nội bộ
def process_confidential_doc(image_path):
"""Trích xuất thông tin từ tài liệu mật - hoàn toàn offline."""
response = ollama.chat(
model='minicpm-v',
messages=[{
'role': 'user',
'content': '''Trích xuất thông tin từ tài liệu:
- Ngày tháng
- Tên người/tổ chức
- Số tiền (nếu có)
- Nội dung chính
Format JSON.''',
'images': [image_path]
}]
)
return response['message']['content']
總結
| 特點 | 建議型號 |
|---|---|
| 一般照片說明 | gemma3:12b |
| OCR/文字閱讀 | minicpm-v |
| 輕、快 | gemma3:4b |
| 分析文檔 | qwen2.5-vl:7b |
| 程式碼截圖 | gemma3:12b |
| 批量處理 | mlx-vlm + 4位元模型 |
練習
- 拉動模型
gemma3:4b並描述5張不同的照片 - 建立 OCR 腳本以從書頁照片中讀取文本
- 建立一個分析批次錯誤截圖的工具(5+張圖片)
- 比較 OCR 結果
minicpm-v和gemma3:12b5.(獎勵)使用 OpenCV + Ollama 視覺即時處理來自網路攝影機的影像
下一篇文章:最佳化效能 — RAM、上下文、並發 →