Ollama API: A Practical Guide with Examples Ollama's local HTTP server on port 11434 provides a REST API for interacting with models, with endpoints for generation, chat, embeddings, and model management. The API supports streaming responses, configurable options, JSON formatting, and tool calling, and offers an OpenAI-compatible route for easy integration. Originally published on DevToolHub . Every Ollama install runs a local HTTP server on port 11434 , and that server is the real interface to the models. The ollama run command is a thin client on top of it. Once you know the two main endpoints, the streaming format, and the options object, you can wire a local model into any application. There is also an OpenAI-compatible route, so existing code that talks to OpenAI can point at Ollama with a base-URL change. The Ollama API is a plain REST API served at http://localhost:11434 . You send JSON with POST , and by default you get a stream of newline-delimited JSON objects back. No API key is required for local access. | Endpoint | Method | Purpose | |---|---|---| /api/generate | POST | Single-prompt text completion | /api/chat | POST | Multi-turn chat with message history and tools | /api/embed | POST | Generate embeddings | /api/tags | GET | List installed models | /api/ps | GET | List models currently loaded in memory | /api/pull | POST | Download a model | Check the server with curl http://localhost:11434/api/version . Use /api/generate for a single prompt with no history: curl http://localhost:11434/api/generate -d '{ "model": "llama3.1", "prompt": "Summarize in one sentence: Ollama serves a local HTTP API on port 11434.", "stream": false }' Use /api/chat for turn-by-turn context or tool calling. Pass a messages array with role values of system , user , assistant , or tool , and send the whole history on each request: curl http://localhost:11434/api/chat -d '{ "model": "llama3.1", "messages": {"role": "system", "content": "You answer in one short sentence."}, {"role": "user", "content": "What is the KV cache?"} , "stream": false }' For application code, /api/chat is the better default even for single questions. By default stream is true , and Ollama returns one JSON object per chunk. The final chunk has "done": true plus timing data: {"model":"llama3.1","response":"","done":true, "total duration":4883583458,"prompt eval count":26, "eval count":298,"eval duration":3789981000} All durations are in nanoseconds. Tokens per second is eval count / eval duration 1e9 . Set "stream": false for a single response object. The options object tunes sampling and context per request: curl http://localhost:11434/api/chat -d '{ "model": "llama3.1", "messages": {"role": "user", "content": "Name three container runtimes."} , "stream": false, "options": {"temperature": 0.2, "num ctx": 8192, "num predict": 200, "seed": 42, "stop": "\n\n" } }' temperature — lower is more deterministic num ctx — context window in tokens for this request num predict — cap on generated tokens seed — with temperature: 0 , gives repeatable output stop — strings that end generationSetting num ctx above what your hardware holds forces a partial CPU offload. Confirm with ollama ps . Set format to "json" for any valid JSON, or pass a JSON schema object to force a shape: curl http://localhost:11434/api/chat -d '{ "model": "llama3.1", "messages": {"role": "user", "content": "List two Linux distros with release years. Respond in JSON."} , "stream": false, "format": { "type": "object", "properties": { "distros": {"type": "array", "items": { "type": "object", "properties": {"name": {"type": "string"}, "year": {"type": "integer"}}, "required": "name", "year" }} }, "required": "distros" } }' Keep the word "JSON" in the prompt and use a low temperature. /api/chat supports function calling through a tools array. The model replies with a tool calls entry instead of text when it decides to use one: curl http://localhost:11434/api/chat -d '{ "model": "llama3.1", "messages": {"role": "user", "content": "What is the weather in Toronto?"} , "stream": false, "tools": { "type": "function", "function": { "name": "get weather", "description": "Get the current weather for a city", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": "city" } } } }' Run the function, then send the result back as a message with "role": "tool" . Tool support depends on the model. Install with pip install ollama . python from ollama import chat response = chat model='llama3.1', messages= {'role': 'user', 'content': 'Why is the sky blue?'}, print response.message.content Streaming: python from ollama import chat stream = chat model='llama3.1', messages= {'role': 'user', 'content': 'Explain the KV cache in two sentences.'} , stream=True for chunk in stream: print chunk 'message' 'content' , end='', flush=True Remote host: python from ollama import Client client = Client host='http://192.168.1.50:11434' There is an AsyncClient with the same methods, plus embed , list , ps , and pull . Ollama serves an OpenAI-style API at http://localhost:11434/v1 with /v1/chat/completions , /v1/completions , /v1/embeddings , and /v1/models : python from openai import OpenAI client = OpenAI base url='http://localhost:11434/v1', api key='ollama' response = client.chat.completions.create model='llama3.1', messages= {'role': 'user', 'content': 'Hello'} , print response.choices 0 .message.content The api key is required by the SDK but ignored by Ollama. Use /v1 for compatibility and /api for full features. The Ollama API has no built-in authentication. Anyone who can reach port 11434 can use, pull, or delete your models. Keep it bound to 127.0.0.1 and put a reverse proxy in front with auth and TLS: server { listen 443 ssl; server name ollama.example.com; location / { proxy pass http://127.0.0.1:11434; proxy set header Host localhost:11434; auth basic "Ollama"; auth basic user file /etc/nginx/.htpasswd; } } Setting OLLAMA HOST=0.0.0.0 without a proxy puts an unauthenticated model server on the open network. Only do that inside a private network or behind an IP-restricted firewall. Q: What port does the Ollama API use? A: Port 11434 on 127.0.0.1 by default. Change it with OLLAMA HOST , for example OLLAMA HOST=0.0.0.0:11434 . Q: Does the Ollama API need an API key? A: No, not for local use. Ollama's hosted cloud models use a key; self-hosted remote access should sit behind a reverse proxy that adds authentication. Q: What is the difference between /api/generate and /api/chat? A: /api/generate takes a single prompt string. /api/chat takes a messages array with roles and supports tool calling. Use /api/chat for application code. Q: How do I get JSON output from the Ollama API? A: Set format to "json" or to a JSON schema object. Keep the word "JSON" in your prompt and use a low temperature. Q: Can I use the OpenAI Python SDK with Ollama? A: Yes. Point base url at http://localhost:11434/v1 and pass any non-empty api key .