{"slug": "llmman-the-unified-interface-for-language-models", "title": "llmman, the unified interface for language models", "summary": "Eric Curtin's llmman, a Rust command-line tool billed as \"Run any agent on any model,\" manages models as standard OCI Image Layouts pulled from Docker Hub, Hugging Face or private registries and serves them behind a single API on port 17434 that speaks OpenAI, Ollama and Anthropic formats. The tool can also launch code agents including Claude Code, Codex, OpenCode and Gemini against a local model or remote endpoint, and works with llama.cpp as its inference runtime, which llmman downloads and caches automatically. Installation is available via a curl script or Homebrew, and models are fetched with commands such as \"llmman pull hf.co/unsloth/Qwen3.5-0.8B-GGUF.", "body_md": "## [llmman in two minutes](#llmman-in-two-minutes)\n\n**[llmman](https://github.com/llmmanorg/llmman)** is a command-line tool written in Rust with the slogan **\"Run any agent on any model\"**.\n\n**llmman** is a project initiated by **[Eric Curtin](https://github.com/ericcurtin)** 👏\n\nIts 3 main functions are as follows:\n\n1. **It manages models like OCI images.** A model is retrieved from Docker Hub, Hugging Face, ... or your private registry, and is stored locally in a standard*OCI Image Layout* . So, no proprietary format, we stay standard.\n2. The part that interests me the most: **It serves these models behind a single API** (`llmman serve` , port`17434` ) that knows how to speak**OpenAI, Ollama and Anthropic** .\n3. **It can launch code agents** (Claude Code, Codex, OpenCode, Gemini, ...) pointed at a local model or a remote endpoint, in a single command.\n\nAnd **llmman** knows how to work with **[`llamacpp`](https://github.com/ggml-org/llama.cpp)** as an inference engine (otherwise known as: as a runtime).\n\nThis is convenient, because my preferred engines are **llamacpp** and **[Docker Model Runner](https://docs.docker.com/ai/model-runner/)** (which itself uses `llamacpp`).\n\nSo, today let's see how to use **llmman** with **llamacpp** to serve my favorite model.\n\n**llmman** comes with many features, I encourage you to consult the excellent documentation.\n\n## [Installation](#installation)\n\nI tested this installation on both **macOS** and **Linux**.\n\nI used the following options:\n\n```\n# Official script\ncurl -fsSL https://llmmanorg.github.io/install.sh | sh\n\n# or via Homebrew\nbrew install llmmanorg/tap/llmman\n```\n\nFor **Linux**, I used the `curl` option, and for **macOS**, I used the `curl` option on one Mac and the `brew` option on another Mac.\n\n### [Choosing the `llamacpp` runtime](#choosing-the-llamacpp-runtime)\n\n`llamacpp` runtime\nThere are several solutions here as well. I chose to let **llmman** handle it, which will download and cache the official **llamacpp** binary (once), and then use it at every startup. (It is also possible to specify the path to your existing **llamacpp** installation).\n\n### [Starting up](#starting-up)\n\nIt's simple, just run this command:\n\n```\nllmman serve\n```\n\nTo verify that everything is working correctly, you can use the following commands:\n\n```\ncurl -s http://127.0.0.1:17434/api/version\ncurl -s http://127.0.0.1:17434/v1/models   # {\"data\":[],\"object\":\"list\"} if you haven't installed any model\n```\n\n### [We need a model](#we-need-a-model)\n\nTo retrieve models, it's also simple, and afterwards **llmman** will allow you to easily switch from one to another. For example, if you want to install `Qwen3.5-0.8B-GGUF`, the version provided by **[Unsloth](https://unsloth.ai/)** on the **[Hugging Face](https://huggingface.co/)** platform, type the following command:\n\n```\nllmman pull hf.co/unsloth/Qwen3.5-0.8B-GGUF\n```\n\nWait a little while for the download, and once it's finished, you can verify that everything is fine and that your model is indeed there with these commands:\n\n```\nllmman list\nllmman show hf.co/unsloth/Qwen3.5-0.8B-GGUF\n```\n\nThere, you are now ready to query your model.\n\n## [Testing the three API versions (Ollama, Anthropic, OpenAI)](#testing-the-three-api-versions-ollama-anthropic-openai)\n\nYou can query the various APIs using `curl`:\n\n```\n# OpenAI\ncurl -s http://127.0.0.1:17434/v1/chat/completions \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"hf.co/unsloth/Qwen3.5-0.8B-GGUF\",\"messages\":[{\"role\":\"user\",\"content\":\"Explain why the hawaiian pizza is the best pizza in the world\"}],\"max_tokens\":300,\"chat_template_kwargs\":{\"enable_thinking\":false}}'\n\n# Anthropic\ncurl -s http://127.0.0.1:17434/v1/messages \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"hf.co/unsloth/Qwen3.5-0.8B-GGUF\",\"max_tokens\":200,\"messages\":[{\"role\":\"user\",\"content\":\"Explain why the hawaiian pizza is the best pizza in the world\"}],\"thinking\":{\"type\":\"disabled\"}}'\n\n# Ollama\ncurl -s http://127.0.0.1:17434/api/chat -d '{\\n  \"model\": \"hf.co/unsloth/Qwen3.5-0.8B-GGUF\",\\n  \"messages\": [{\"role\": \"user\", \"content\": \"Explain why the hawaiian pizza is the best pizza in the world\"}],\\n  \"stream\": false\\n}'\n```\n\nAs you can see, it's easy to implement **llmman** and then use it in generative AI applications or with your favorite agents.\n\n## [Go Example with the OpenAI Go SDK](#go-example-with-the-openai-go-sdk)\n\nYou can therefore use the API exposed by **llmman** with the usual frameworks, as here with Go and the OpenAI API:\n\n```\npackage main\n\nimport (\n    \"context\"\n    \"fmt\"\n\n    openai \"github.com/openai/openai-go/v3\"\n    \"github.com/openai/openai-go/v3/option\"\n)\n\nfunc main() {\n    ctx := context.Background()\n\n    model := \"hf.co/unsloth/Qwen3.5-0.8B-GGUF\"\n    baseURL := \"http://127.0.0.1:17434/v1\"\n\n    client := openai.NewClient(\n        option.WithBaseURL(baseURL),\n        option.WithAPIKey(\"not-needed\"),\n    )\n\n    completion, err := client.Chat.Completions.New(\n        ctx,\n        openai.ChatCompletionNewParams{\n            Model: model,\n            Messages: []openai.ChatCompletionMessageParamUnion{\n                openai.SystemMessage(\"You are a pizza expert.\"),\n                openai.UserMessage(\"Tell me more about Hawaiian pizza.\"),\n            },\n            Temperature: openai.Float(0.0),\n            TopP:        openai.Float(0.9),\n            MaxTokens:   openai.Int(2048),\n        },\n        option.WithJSONSet(\n            \"chat_template_kwargs\",\n            map[string]any{\n                \"enable_thinking\": false,\n            },\n        ),\n    )\n    if err != nil {\n        fmt.Println(\"[error:\", err, \"]\")\n        return\n    }\n\n    if len(completion.Choices) == 0 {\n        fmt.Println(\"[error: empty completion]\")\n        return\n    }\n\n    fmt.Println(completion.Choices[0].Message.Content)\n}\n```\n\n## [One last thing before you go](#one-last-thing-before-you-go)\n\n**llmman** provides a particularly well-made web UI. To access it, it's simple: open this URL in your browser: [http://127.0.0.1:17434/](http://127.0.0.1:17434/) and you will be able to search for and download models, and \"chat with your models\":\n\nThat's all for today. I'll leave you to play with **llmman**. Feel free to ask questions 🙂\n\nWritten by\n\nKeep reading\n\n### Docker Agent + llama.cpp: a local code agent in 5 minutes 🤖💙🦙\n\nServe a GGUF model with llama.cpp, point Docker Agent at it with a short agent.yaml, and get a fully local code agent running in minutes.\n\nSep 15, 2026\n\n### Zed Editor + Docker Agent + ACP: coding with a local agent plugged into llmman\n\nPlug Zed's agent panel into Docker Agent over ACP, then point it at llmman to chat with a fully local, OpenAI-compatible coding agent.\n\nOct 3, 2026\n\n### Docker Model Runner, standalone edition (`dmr`)\n\nInstall and use dmr, the standalone Docker Model Runner: one binary with the inference daemon and the full model CLI, no Docker Desktop needed.\n\nJul 31, 2026", "url": "https://wpnews.pro/news/llmman-the-unified-interface-for-language-models", "canonical_source": "https://k33g.org/p/20261001-llmman", "published_at": "2026-10-01 00:00:00+00:00", "updated_at": "2026-10-07 18:16:44.703848+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-agents", "developer-tools", "ai-infrastructure"], "entities": ["llmman", "Eric Curtin", "llama.cpp", "Docker Model Runner", "Hugging Face", "Unsloth", "Qwen3.5-0.8B-GGUF", "Claude Code"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/llmman-the-unified-interface-for-language-models", "markdown": "https://wpnews.pro/news/llmman-the-unified-interface-for-language-models.md", "text": "https://wpnews.pro/news/llmman-the-unified-interface-for-language-models.txt", "jsonld": "https://wpnews.pro/news/llmman-the-unified-interface-for-language-models.jsonld"}}