cd /news/developer-tools/ollama-api-a-practical-guide-with-ex… · home topics developer-tools article
[ARTICLE · art-121459] src=dev.to ↗ pub= topic=developer-tools verified=true sentiment=· neutral

Ollama API: A Practical Guide with Examples

Ollama's local HTTP server on port 11434 provides a REST API for interacting with models, with endpoints for generation, chat, embeddings, and model management. The API supports streaming responses, configurable options, JSON formatting, and tool calling, and offers an OpenAI-compatible route for easy integration.

read5 min views1 publishedSep 4, 2026

Originally published on

[DevToolHub].

Every Ollama install runs a local HTTP server on port 11434

, and that server is the real interface to the models. The ollama run

command is a thin client on top of it. Once you know the two main endpoints, the streaming format, and the options object, you can wire a local model into any application.

There is also an OpenAI-compatible route, so existing code that talks to OpenAI can point at Ollama with a base-URL change.

The Ollama API is a plain REST API served at http://localhost:11434

. You send JSON with POST

, and by default you get a stream of newline-delimited JSON objects back. No API key is required for local access.

Endpoint Method Purpose
/api/generate
POST Single-prompt text completion
/api/chat
POST Multi-turn chat with message history and tools
/api/embed
POST Generate embeddings
/api/tags
GET List installed models
/api/ps
GET List models currently loaded in memory
/api/pull
POST Download a model

Check the server with curl http://localhost:11434/api/version

.

Use /api/generate

for a single prompt with no history:

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.1",
  "prompt": "Summarize in one sentence: Ollama serves a local HTTP API on port 11434.",
  "stream": false
}'

Use /api/chat

for turn-by-turn context or tool calling. Pass a messages

array with role

values of system

, user

, assistant

, or tool

, and send the whole history on each request:

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1",
  "messages": [
    {"role": "system", "content": "You answer in one short sentence."},
    {"role": "user", "content": "What is the KV cache?"}
  ],
  "stream": false
}'

For application code, /api/chat

is the better default even for single questions.

By default stream

is true

, and Ollama returns one JSON object per chunk. The final chunk has "done": true

plus timing data:

{"model":"llama3.1","response":"","done":true,
 "total_duration":4883583458,"prompt_eval_count":26,
 "eval_count":298,"eval_duration":3789981000}

All durations are in nanoseconds. Tokens per second is eval_count / eval_duration * 1e9

. Set "stream": false

for a single response object.

The options

object tunes sampling and context per request:

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1",
  "messages": [{"role": "user", "content": "Name three container runtimes."}],
  "stream": false,
  "options": {"temperature": 0.2, "num_ctx": 8192, "num_predict": 200, "seed": 42, "stop": ["\n\n"]}
}'

temperature

— lower is more deterministicnum_ctx

— context window in tokens for this requestnum_predict

— cap on generated tokensseed

— with temperature: 0

, gives repeatable outputstop

— strings that end generationSetting num_ctx

above what your hardware holds forces a partial CPU offload. Confirm with ollama ps

.

Set format

to "json"

for any valid JSON, or pass a JSON schema object to force a shape:

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1",
  "messages": [{"role": "user", "content": "List two Linux distros with release years. Respond in JSON."}],
  "stream": false,
  "format": {
    "type": "object",
    "properties": {
      "distros": {"type": "array", "items": {
        "type": "object",
        "properties": {"name": {"type": "string"}, "year": {"type": "integer"}},
        "required": ["name", "year"]
      }}
    },
    "required": ["distros"]
  }
}'

Keep the word "JSON" in the prompt and use a low temperature.

/api/chat

supports function calling through a tools

array. The model replies with a tool_calls

entry instead of text when it decides to use one:

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1",
  "messages": [{"role": "user", "content": "What is the weather in Toronto?"}],
  "stream": false,
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Get the current weather for a city",
      "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}
    }
  }]
}'

Run the function, then send the result back as a message with "role": "tool"

. Tool support depends on the model.

Install with pip install ollama

.

from ollama import chat

response = chat(model='llama3.1', messages=[
    {'role': 'user', 'content': 'Why is the sky blue?'},
])
print(response.message.content)

Streaming:

from ollama import chat

stream = chat(model='llama3.1',
    messages=[{'role': 'user', 'content': 'Explain the KV cache in two sentences.'}],
    stream=True)
for chunk in stream:
    print(chunk['message']['content'], end='', flush=True)

Remote host:

from ollama import Client
client = Client(host='http://192.168.1.50:11434')

There is an AsyncClient

with the same methods, plus embed()

, list()

, ps()

, and pull()

.

Ollama serves an OpenAI-style API at http://localhost:11434/v1

with /v1/chat/completions

, /v1/completions

, /v1/embeddings

, and /v1/models

:

from openai import OpenAI

client = OpenAI(base_url='http://localhost:11434/v1', api_key='ollama')
response = client.chat.completions.create(
    model='llama3.1',
    messages=[{'role': 'user', 'content': 'Hello'}],
)
print(response.choices[0].message.content)

The api_key

is required by the SDK but ignored by Ollama. Use /v1

for compatibility and /api

for full features.

The Ollama API has no built-in authentication. Anyone who can reach port 11434

can use, pull, or delete your models. Keep it bound to 127.0.0.1

and put a reverse proxy in front with auth and TLS:

server {
    listen 443 ssl;
    server_name ollama.example.com;
    location / {
        proxy_pass http://127.0.0.1:11434;
        proxy_set_header Host localhost:11434;
        auth_basic "Ollama";
        auth_basic_user_file /etc/nginx/.htpasswd;
    }
}

Setting OLLAMA_HOST=0.0.0.0

without a proxy puts an unauthenticated model server on the open network. Only do that inside a private network or behind an IP-restricted firewall.

Q: What port does the Ollama API use?

A: Port 11434

on 127.0.0.1

by default. Change it with OLLAMA_HOST

, for example OLLAMA_HOST=0.0.0.0:11434

.

Q: Does the Ollama API need an API key?

A: No, not for local use. Ollama's hosted cloud models use a key; self-hosted remote access should sit behind a reverse proxy that adds authentication.

Q: What is the difference between /api/generate and /api/chat?

A: /api/generate

takes a single prompt

string. /api/chat

takes a messages

array with roles and supports tool calling. Use /api/chat

for application code.

Q: How do I get JSON output from the Ollama API?

A: Set format

to "json"

or to a JSON schema object. Keep the word "JSON" in your prompt and use a low temperature.

Q: Can I use the OpenAI Python SDK with Ollama?

A: Yes. Point base_url

at http://localhost:11434/v1

and pass any non-empty api_key

.

── more in #developer-tools 4 stories · sorted by recency
── more on @ollama 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ollama-api-a-practic…] indexed:0 read:5min 2026-09-04 ·