Chuyển đến nội dung chính

Lesson 7: Ollama REST API - OpenAI-compatible endpoint

Ollama exposes OpenAI compatible REST APIs: /api/chat, /api/generate, /api/embeddings. Use curl and Python requests. Streaming responses. Integration with OpenAI SDK.

🧠 AI & ML — Lesson 0 Lesson 7: Ollama REST API - OpenAI-compatible endpoint. endpoint

Running AI Local with Ollama on Apple Silicon

Part 3: API integration & application programming

xdev.asia

Introduction

Ollama is more than just a CLI tool. It exposes a REST API that runs on top http://localhost:11434, allowing any application to call LLM. In particular, Ollama supports OpenAI-compatible endpoint — meaning code written for the OpenAI API can use the Ollama model almost unchanged.


1. Basic API

Check the server

curl http://localhost:11434
# Output: Ollama is running

/api/generate — Generate text

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Viết hàm fibonacci bằng Python",
  "stream": false
}'

Response:

{
  "model": "llama3.2",
  "response": "```python\ndef fibonacci(n):\n ...",
  "done": true,
  "total_duration": 3241920000,
  "prompt_eval_count": 12,
  "prompt_eval_duration": 180000000,
  "eval_count": 156,
  "eval_duration": 2890000000
}

/api/chat — Chat với history

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.2",
  "messages": [
    {"role": "system", "content": "You are a Python expert, answer in Vietnamese."},
    {"role": "user", "content": "What is dictionary comprehension?"}
  ],
  "stream": false
}'

/api/embeddings — Vector embeddings

curl http://localhost:11434/api/embeddings -d '{
  "model": "nomic-embed-text",
  "prompt": "Docker is a containerization technology"
}'

Response chứa vector embedding (mảng số thực), dùng cho RAG, semantic search.


2. Streaming responses

Mặc định stream: true — Ollama trả về từng chunk:

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Explain microservices"
}'

Mỗi dòng là một JSON object:

{"model":"llama3.2","response":"Micro","done":false}
{"model":"llama3.2","response":"services","done":false}
{"model":"llama3.2","response":" is","done":false}
...
{"model":"llama3.2","response":"","done":true,"total_duration":...}

3. OpenAI-compatible endpoint

Ollama expose endpoint tương thích OpenAI tại /v1/:

# Chat completions (like OpenAI)
curl http://localhost:11434/v1/chat/completions -d '{
  "model": "llama3.2",
  "messages": [
    {"role": "user", "content": "Hello, who are you?"}
  ]
}'

Dùng OpenAI Python SDK

pip3 install openai
from openai import OpenAI

# Points to Ollama instead of OpenAI
client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama", # Ollama does not need a key, but the SDK requires it
)

response = client.chat.completions.create(
    model="llama3.2",
    messages=[
        {"role": "system", "content": "Reply in Vietnamese."},
        {"role": "user", "content": "What is Kubernetes?"},
    ],
    temperature=0.7,
    max_tokens=500,
)

print(response.choices[0].message.content)

Streaming với OpenAI SDK

from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

stream = client.chat.completions.create(
    model="llama3.2",
    messages=[{"role": "user", "content": "Write quicksort function in Python"}],
    stream=True,
)

for chunks in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
print()

💡 Tại sao điều này quan trọng? Bất kỳ thư viện/framework nào hỗ trợ OpenAI API (LangChain, LlamaIndex, Vercel AI SDK, Cursor, Continue.dev...) đều có thể dùng Ollama chỉ bằng cách đổi base_url.


4. Python requests (không cần SDK)

import requests
import json

def chat(messages, model="llama3.2", stream=False):
    response = requests.post(
        "http://localhost:11434/api/chat",
        json={"model": model, "messages": messages, "stream": stream}
    )
    if streams:
        for line in response.iter_lines():
            if line:
                data = json.loads(line)
                if not data["done"]:
                    yield data["message"]["content"]
    else:
        return response.json()["message"]["content"]

# Non-streaming
result = chat([
    {"role": "user", "content": "What is Docker?"}
])
print(result)

#Streaming
print("\n--- Streaming ---")
for token in chat(
    [{"role": "user", "content": "Write a binary search function"}],
    stream=True
):
    print(token, end="", flush=True)

5. JavaScript/TypeScript (Node.js)

Fetch API

async function chat(messages) {
  const response = await fetch('http://localhost:11434/api/chat', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({
      model: 'llama3.2',
      messages,
      stream: false,
    }),
  });
  const data = await response.json();
  return data.message.content;
}

// Use
const result = await chat([
  { role: 'user', content: 'Explain Promise in JavaScript' },
]);
console.log(result);

Streaming với ReadableStream

async function streamChat(messages) {
  const response = await fetch('http://localhost:11434/api/chat', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({
      model: 'llama3.2',
      messages,
      stream: true,
    }),
  });

  const reader = response.body.getReader();
  const decoder = new TextDecoder();

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;

    const lines = decoder.decode(value).split('\n').filter(Boolean);
    for (const line of lines) {
      const data = JSON.parse(line);
      if (!data.done) {
        process.stdout.write(data.message.content);
      }
    }
  }
  console.log();
}

await streamChat([
  { role: 'user', content: 'Write an async function fetch data in TypeScript' },
]);

Dùng OpenAI SDK cho JavaScript

npm install openai
import OpenAI from 'openai';

const client = new OpenAI({
  baseURL: 'http://localhost:11434/v1',
  apiKey: 'ollama',
});

const completion = await client.chat.completions.create({
  model: 'llama3.2',
  messages: [{ role: 'user', content: 'Hello!' }],
});

console.log(completion.choices[0].message.content);

6. API Parameters quan trọng

ParameterMô tảMặc định
temperatureĐộ sáng tạo (0.0-2.0)0.8
top_pNucleus sampling0.9
top_kTop-k sampling40
num_predictMax tokens generate128
num_ctxContext window size2048
stopStop sequences[]
seedRandom seed (cho reproducible)random
keep_aliveThời gian giữ model loaded"5m"
curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Write a poem about programming",
  "options": {
    "temperature": 1.2,
    "top_p": 0.95,
    "num_predict": 300,
    "seed": 42
  },
  "stream": false
}'

7. Model management API

# List models
curl http://localhost:11434/api/tags

# Show model info
curl http://localhost:11434/api/show -d '{"name": "llama3.2"}'

# Pull model
curl http://localhost:11434/api/pull -d '{"name": "gemma3:4b"}'

# Delete model
curl http://localhost:11434/api/delete -d '{"name": "mistral"}'

# Running models
curl http://localhost:11434/api/ps

8. Expose API ra mạng local

Mặc định Ollama chỉ listen trên localhost. Để expose cho các máy khác trong mạng:

# Listen on all interfaces
export OLLAMA_HOST=0.0.0.0:11434
ollama serve

Giờ các máy khác trong mạng có thể truy cập:

# From another device
curl http://192.168.1.100:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Hello!",
  "stream": false
}'

⚠️ Bảo mật: Chỉ expose trong mạng nội bộ (LAN). Không expose ra internet public nếu không có authentication.


Tóm tắt

EndpointMethodMô tả
/api/generatePOSTGenerate text
/api/chatPOSTChat với message history
/api/embeddingsPOSTVector embeddings
/v1/chat/completionsPOSTOpenAI-compatible
/api/tagsGETList models
/api/showPOSTModel info
/api/pullPOSTPull model
/api/psGETRunning models

Bài tập

  1. Dùng curl gọi /api/chat với system prompt tiếng Việt
  2. Viết Python script tạo chatbot terminal dùng requests + streaming
  3. Dùng OpenAI SDK Python/JS trỏ tới Ollama, chạy chat
  4. Đo thời gian response với các temperature different (0.1, 0.7, 1.5)
  5. Try exposing Ollama to LAN and calling from your phone (via browser)

Next article: Build a local chatbot with Python →