Chuyển đến nội dung chính

Lesson 11: Modelfiles - Custom models & system prompts

Modelfile syntax. Custom system prompt, temperature, top_p, stop tokens. Create specialized models. Manage and share custom models.

🧠 AI & ML — Lesson 1 Lesson 11: Modelfiles - Custom models & system prompts

Running AI Local with Ollama on Apple Silicon

Part 4: Optimization, management & production setup

xdev.asia

Introduction

Modelfile is how you customize models in Ollama — set system prompts, adjust parameters, create your own "character" AI. Like Dockerfile but for AI models.


1. Basic Modelfile

Structure

# Modelfile cơ bản
FROM llama3.2

SYSTEM """Bạn là trợ lý lập trình chuyên về Python.
Luôn trả lời bằng tiếng Việt.
Viết code rõ ràng, có comment."""

PARAMETER temperature 0.7
PARAMETER num_ctx 4096

Create model from Modelfile

# Tạo model
ollama create xdev-py -f ./Modelfile

# Chạy
ollama run xdev-py

# Kiểm tra
ollama list | grep xdev

2. Directives in Modelfile

FROM — Base model

# Từ model có sẵn
FROM llama3.2

# Từ model cụ thể (tag)
FROM llama3.2:3b-instruct-q4_K_M

# Từ custom model khác
FROM xdev-py

SYSTEM — System prompt

SYSTEM """Bạn là xDev AI, trợ lý lập trình.

Quy tắc:
1. Trả lời bằng tiếng Việt
2. Code phải có type hints (Python)
3. Giải thích ngắn gọn, đi thẳng vào vấn đề
4. Dùng markdown formatting
5. Nếu không chắc, nói rõ "tôi không chắc chắn"
"""

PARAMETER — Model parameter

# Temperature: 0 = deterministic, 1 = creative (default: 0.8)
PARAMETER temperature 0.7

# Top_p: nucleus sampling (default: 0.9)
PARAMETER top_p 0.9

# Top_k: limit token candidates (default: 40)
PARAMETER top_k 40

# Context window
PARAMETER num_ctx 4096

# Max tokens to generate
PARAMETER num_predict 1024

# Repeat penalty (default: 1.1)
PARAMETER repeat_penalty 1.1

# Repeat last N tokens to check (default: 64)
PARAMETER repeat_last_n 64

# Stop sequences
PARAMETER stop "<|end|>"
PARAMETER stop "Human:"
PARAMETER stop "---"

# Seed (for reproducibility)
PARAMETER seed 42

# Mirostat sampling (0=disabled, 1=v1, 2=v2)
PARAMETER mirostat 2
PARAMETER mirostat_eta 0.1
PARAMETER mirostat_tau 5.0

TEMPLATE — Chat template

TEMPLATE """{{ if .System }}<|system|>
{{ .System }}<|end|>
{{ end }}{{ if .Prompt }}<|user|>
{{ .Prompt }}<|end|>
{{ end }}<|assistant|>
{{ .Response }}<|end|>
"""

MESSAGE — Pre-seed conversations

MESSAGE user "Xin chào!"
MESSAGE assistant "Chào bạn! Tôi là xDev AI, tôi có thể giúp gì về lập trình?"

LICENSE

LICENSE """
MIT License
Custom model by xDev.asia
"""

3. Realistic custom models example

3.1. Python Expert

# Modelfile.python-expert
FROM llama3.2

SYSTEM """Bạn là Python Expert AI.

Quy tắc:
- Code phải có type hints
- Dùng f-strings thay format()
- Follow PEP 8
- Viết docstring cho functions
- Error handling với specific exceptions
- Suggest test cases khi viết function

Không bao giờ dùng: global variables, bare except, eval, exec.
Trả lời bằng tiếng Việt, code bằng Python."""

PARAMETER temperature 0.3
PARAMETER num_ctx 4096
PARAMETER num_predict 2048
PARAMETER stop "```"
ollama create python-expert -f Modelfile.python-expert
ollama run python-expert "Write function to validate email address"

3.2. Code Reviewer

# Modelfile.code-reviewer
FROM llama3.2

SYSTEM """You are a Senior Code Reviewer.

When reviewing code, you check:
1. 🐛 Bugs & logic errors
2. 🔒 Security vulnerabilities (OWASP Top 10)
3. ⚡ Performance issues
4. 📖 Readability & maintainability
5. 🧪 Testability

Format output:
## Overview
[General comments]

## Issues
- 🔴 Critical: [...]
- 🟡 Warning: [...]
- 🟢 Suggestion: [...]

## Refactored code
[Code fixed]

Reply in Vietnamese."""

PARAMETER temperature 0.2
PARAMETER num_ctx 8192

3.3. DevOps Assistant

# Modelfile.devops
FROM llama3.2

SYSTEM """You are a DevOps Engineer AI.

Expertise:
- Docker, Kubernetes, Helm
- CI/CD (GitHub Actions, GitLab CI, Jenkins)
- Infrastructure as Code (Terraform, Ansible)
- Cloud (AWS, GCP, Azure)
- Monitoring (Prometheus, Grafana)

Rules:
- Always mention security best practices
- Explain "why" not just "how"
- Provides practical YAML/JSON examples
- Warning about common pitfalls
Reply in Vietnamese."""

PARAMETER temperature 0.5
PARAMETER num_ctx 4096

MESSAGE user "Who are you?"
MESSAGE assistant "I am DevOps AI Assistant, specializing in support for Docker, K8s, CI/CD, and Infrastructure. Ask me anything about DevOps!"

3.4. SQL Generator

# Modelfile.sql
FROM llama3.2

SYSTEM """You are a SQL Expert.
- Only write SQL, no lengthy explanations
- Use PostgreSQL syntax
- Include comments in SQL
- Optimize for performance
- Avoid SELECT *

Format: write SQL in code block, with 1-2 lines of explanation.
If the question is vague, ask again before writing the SQL."""

PARAMETER temperature 0.1
PARAMETER num_predict 1024
PARAMETER stop ";"

3.5. Creative Writer (Vietnamese)

# Modelfile.writer
FROM llama3.2

SYSTEM """You are a creative writer writing in Vietnamese.
- Natural, smooth writing style
- Use metaphors and vivid images
- Tone appropriate to requirements (formal, casual, humorous...)
- Know how to write blog posts, poems, short stories, advertisements.

PARAMETER temperature 0.9
PARAMETER top_p 0.95
PARAMETER top_k 60
PARAMETER repeat_penalty 1.2
PARAMETER num_predict 4096

4. Quản lý custom models

Liệt kê models

ollama list

Output:

NAME ID SIZE MODIFIED
python-expert abc123def 2.0 GB 5 minutes ago
code-reviewer def456ghi 2.0 GB 10 minutes ago
llama3.2:latest xyz789abc 2.0 GB 2 hours ago

Xem Modelfile của model

ollama show python-expert --modelfile

Copy model

ollama cp python-expert python-expert-v2

Xóa model

ollama rm python-expert-v2

Export/Import workflow

# Export: save Modelfile
ollama show python-expert --modelfile > Modelfile.python-expert

# Import on another machine
ollama create python-expert -f Modelfile.python-expert

5. Tham số model — Deep dive

Temperature

Kiểm soát mức độ "sáng tạo":

temperature = 0.0 → Always choose the token with the highest probability (deterministic)
temperature = 0.5 → Balance between accuracy and diversity
temperature = 1.0 → Creative, sometimes inaccurate
temperature = 2.0 → Very random (not recommended)
Use caseTemperature
Code generation0.1 - 0.3
Q&A, factual0.3 - 0.5
Chat thông thường0.5 - 0.8
Creative writing0.8 - 1.0
Brainstorming0.9 - 1.2

Top_p vs Top_k

top_k = 40: Choose from 40 tokens with the highest probability
top_p = 0.9: Choose from tokens that account for 90% cumulative probability

Thường dùng một trong hai, không cả hai.

Stop sequences

# Stop generating when a keyword is encountered
PARAMETER stop "<|end|>"
PARAMETER stop "Human:"
PARAMETER stop "User:"

# Useful when the model often "loops" or creates its own dialogue

6. Test và iterate

Script test model

#!/usr/bin/env python3
"""Test custom model with multiple prompts."""

import ollama

MODEL = "python-expert"
TEST_PROMPTS = [
    "Write a function that reads CSV file and returns a list of dicts",
    "Explain the @property decorator",
    "Find bugs in this code: def add(a, b): return a * b",
    "Write unit test for function validate_email()",
]

for i, prompt in enumerate(TEST_PROMPTS, 1):
    print(f"\n{'='*60}")
    print(f"Test {i}: {prompt}")
    print('='*60)

    response = ollama.chat(
        model=MODEL,
        messages=[{'role': 'user', 'content': prompt}]
    )
    print(response['message']['content'])

So sánh models

import ollama

MODELS = ['llama3.2', 'python-expert', 'code-reviewer']
PROMPT = "Review code: def calc(x): return x*2+1"

for models in MODELS:
    print(f"\n--- {model} ---")
    response = ollama.chat(
        model=model,
        messages=[{'role': 'user', 'content': PROMPT}]
    )
    print(response['message']['content'][:500])

7. Tổ chức Modelfiles trong project

ai-models/
├── README.md
├── Modelfile.python-expert
├── Modelfile.code-reviewer
├── Modelfile.devops
├── Modelfile.sql
├── Modelfile.writer
├── setup.sh # Script to create all models
└── test.py # Script test models

setup.sh:

#!/bin/bash
echo "🚀 Creating custom Ollama models..."

for f in Modelfile.*; due
    name="${f#Modelfile.}"
    echo "📦 Creating: $name"
    ollama create "$name" -f "$f"
done

echo "✅ Done! Models:"
ollama list

Tóm tắt

DirectiveMục đíchVí dụ
FROMBase modelFROM llama3.2
SYSTEMSystem promptPersonality, rules
PARAMETERTuning paramstemperature, num_ctx
TEMPLATEChat formatCustom template
MESSAGESeed conversationPre-defined Q&A
LICENSELicense infoMIT, custom

Exercises

  1. Create Modelfile for "Vietnamese Technical Writer" — writing technical blog posts
  2. Create a "JSON Generator" model — only output JSON, no extra text
  3. Compare temperature 0.1 vs 0.5 vs 0.9 with the same prompt and comments
  4. Organize folder ai-models/ with 3+ Modelfiles and setup.sh script
  5. (Bonus) Create "Interview Bot" model — ask interview questions, score

Next article: Complete Workflow — Personal AI setup →