Originally published on
[DevToolHub].
Every Ollama install runs a local HTTP server on port 11434
, and that server is the real interface to the models. The ollama run
command is a thin client on top of it. Once you know the two main endpoints, the streaming format, and the options object, you can wire a local model into any application.
There is also an OpenAI-compatible route, so existing code that talks to OpenAI can point at Ollama with a base-URL change.
The Ollama API is a plain REST API served at http://localhost:11434
. You send JSON with POST
, and by default you get a stream of newline-delimited JSON objects back. No API key is required for local access.
| Endpoint | Method | Purpose |
|---|---|---|
/api/generate |
||
| POST | Single-prompt text completion | |
/api/chat |
||
| POST | Multi-turn chat with message history and tools | |
/api/embed |
||
| POST | Generate embeddings | |
/api/tags |
||
| GET | List installed models | |
/api/ps |
||
| GET | List models currently loaded in memory | |
/api/pull |
||
| POST | Download a model |
Check the server with curl http://localhost:11434/api/version
.
Use /api/generate
for a single prompt with no history:
curl http://localhost:11434/api/generate -d '{
"model": "llama3.1",
"prompt": "Summarize in one sentence: Ollama serves a local HTTP API on port 11434.",
"stream": false
}'
Use /api/chat
for turn-by-turn context or tool calling. Pass a messages
array with role
values of system
, user
, assistant
, or tool
, and send the whole history on each request:
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1",
"messages": [
{"role": "system", "content": "You answer in one short sentence."},
{"role": "user", "content": "What is the KV cache?"}
],
"stream": false
}'
For application code, /api/chat
is the better default even for single questions.
By default stream
is true
, and Ollama returns one JSON object per chunk. The final chunk has "done": true
plus timing data:
{"model":"llama3.1","response":"","done":true,
"total_duration":4883583458,"prompt_eval_count":26,
"eval_count":298,"eval_duration":3789981000}
All durations are in nanoseconds. Tokens per second is eval_count / eval_duration * 1e9
. Set "stream": false
for a single response object.
The options
object tunes sampling and context per request:
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1",
"messages": [{"role": "user", "content": "Name three container runtimes."}],
"stream": false,
"options": {"temperature": 0.2, "num_ctx": 8192, "num_predict": 200, "seed": 42, "stop": ["\n\n"]}
}'
temperature
— lower is more deterministicnum_ctx
— context window in tokens for this requestnum_predict
— cap on generated tokensseed
— with temperature: 0
, gives repeatable outputstop
— strings that end generationSetting num_ctx
above what your hardware holds forces a partial CPU offload. Confirm with ollama ps
.
Set format
to "json"
for any valid JSON, or pass a JSON schema object to force a shape:
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1",
"messages": [{"role": "user", "content": "List two Linux distros with release years. Respond in JSON."}],
"stream": false,
"format": {
"type": "object",
"properties": {
"distros": {"type": "array", "items": {
"type": "object",
"properties": {"name": {"type": "string"}, "year": {"type": "integer"}},
"required": ["name", "year"]
}}
},
"required": ["distros"]
}
}'
Keep the word "JSON" in the prompt and use a low temperature.
/api/chat
supports function calling through a tools
array. The model replies with a tool_calls
entry instead of text when it decides to use one:
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1",
"messages": [{"role": "user", "content": "What is the weather in Toronto?"}],
"stream": false,
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}
}
}]
}'
Run the function, then send the result back as a message with "role": "tool"
. Tool support depends on the model.
Install with pip install ollama
.
from ollama import chat
response = chat(model='llama3.1', messages=[
{'role': 'user', 'content': 'Why is the sky blue?'},
])
print(response.message.content)
Streaming:
from ollama import chat
stream = chat(model='llama3.1',
messages=[{'role': 'user', 'content': 'Explain the KV cache in two sentences.'}],
stream=True)
for chunk in stream:
print(chunk['message']['content'], end='', flush=True)
Remote host:
from ollama import Client
client = Client(host='http://192.168.1.50:11434')
There is an AsyncClient
with the same methods, plus embed()
, list()
, ps()
, and pull()
.
Ollama serves an OpenAI-style API at http://localhost:11434/v1
with /v1/chat/completions
, /v1/completions
, /v1/embeddings
, and /v1/models
:
from openai import OpenAI
client = OpenAI(base_url='http://localhost:11434/v1', api_key='ollama')
response = client.chat.completions.create(
model='llama3.1',
messages=[{'role': 'user', 'content': 'Hello'}],
)
print(response.choices[0].message.content)
The api_key
is required by the SDK but ignored by Ollama. Use /v1
for compatibility and /api
for full features.
The Ollama API has no built-in authentication. Anyone who can reach port 11434
can use, pull, or delete your models. Keep it bound to 127.0.0.1
and put a reverse proxy in front with auth and TLS:
server {
listen 443 ssl;
server_name ollama.example.com;
location / {
proxy_pass http://127.0.0.1:11434;
proxy_set_header Host localhost:11434;
auth_basic "Ollama";
auth_basic_user_file /etc/nginx/.htpasswd;
}
}
Setting OLLAMA_HOST=0.0.0.0
without a proxy puts an unauthenticated model server on the open network. Only do that inside a private network or behind an IP-restricted firewall.
Q: What port does the Ollama API use?
A: Port 11434
on 127.0.0.1
by default. Change it with OLLAMA_HOST
, for example OLLAMA_HOST=0.0.0.0:11434
.
Q: Does the Ollama API need an API key?
A: No, not for local use. Ollama's hosted cloud models use a key; self-hosted remote access should sit behind a reverse proxy that adds authentication.
Q: What is the difference between /api/generate and /api/chat?
A: /api/generate
takes a single prompt
string. /api/chat
takes a messages
array with roles and supports tool calling. Use /api/chat
for application code.
Q: How do I get JSON output from the Ollama API?
A: Set format
to "json"
or to a JSON schema object. Keep the word "JSON" in your prompt and use a low temperature.
Q: Can I use the OpenAI Python SDK with Ollama?
A: Yes. Point base_url
at http://localhost:11434/v1
and pass any non-empty api_key
.