{"slug": "docker-agent-llama-cpp-a-local-code-agent-in-5-minutes", "title": "Docker Agent + llama.cpp: a local code agent in 5 minutes 🤖💙🦙", "summary": "Docker Agent can be connected to a local llama.cpp server to run a code agent entirely on-device, with llama.cpp installed via a single curl command and started in serve mode on http://localhost:8080. The walkthrough serves the JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M model through llama.cpp's OpenAI-compatible /v1/chat/completions endpoint and configures Docker Agent via an agent.yaml file using the openai_ch API type, with Docker Agent available as a Docker Desktop plugin or a standalone brew install.", "body_md": "# Docker Agent + llama.cpp: a local code agent in 5 minutes 🤖💙🦙\n\nMy favourite code agent is still **[Docker Agent](https://docker.github.io/docker-agent/)**, especially when I work with \"local\" LLMs, be they \"big\" ([gemma-4-26B-A4B-it](https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF)) when I'm on my work laptop, or more modest ([Mellum2-12B-A2.5B-Instruct](https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M)) when I'm on my personal Mac Book Air.\n\nA big advantage of **[Docker Agent](https://docker.github.io/docker-agent/)** is being able to connect to various model providers, **remote** ones (like Anthropic, OpenAI, [OVH's AI endpoints](https://blog.ovhcloud.com/en/posts/discovering-docker-agent-with-ai-endpoints/), [MistralAI](https://mistral.ai/), ...) or **local** ones (like [Ollama](https://ollama.com/), [Docker Model Runner](https://docs.docker.com/ai/model-runner/), [llama.cpp](https://llama.app/), ...)\n\nToday, the one I'm interested in is **[llama.cpp](https://llama.app/)**, because when a new model in GGUF format shows up, **[llama.cpp](https://llama.app/)** is generally the first to be updated to handle the model's specifics.\n\n[Kronk](https://www.kronkai.com/) is very reactive too, I had written an [article](https://k33g.org/p/20260510-kronk-docker-agent-sbx) about it that I'll have to refresh\n\nBut let's get back to **[llama.cpp](https://llama.app/)**, which we'll have to install and start.\n\n## [Prerequisites](#prerequisites)\n\n### [Llama.cpp](#llamacpp)\n\nInstalling it is very simple, a single `curl` command is enough (otherwise there are other options, I'll let you refer to the llama.cpp website).\n\n```\ncurl -LsSf https://llama.app/install.sh | sh\n```\n\nThen all you have to do is start `llama` in `serve` mode (so, in API mode) so that it can be used by a code agent:\n\n```\nllama serve -hf JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M:Q4_K_M\n```\n\nIf the model isn't present on your machine, **[llama.cpp](https://llama.app/)** will download it (from [Hugging Face](https://huggingface.co/), hence the `-hf` flag):\n\nAnd then the model is served on [http://localhost:8080](http://localhost:8080)\n\nYou can run a few checks to verify:\n\n```\n# Server health\ncurl http://localhost:8080/health\n# {\"status\":\"ok\"}\n\n# Exact name of the exposed model\ncurl http://localhost:8080/v1/models\n\n# First chat completion\ncurl http://localhost:8080/v1/chat/completions \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"model\": \"JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M:Q4_K_M\",\n    \"messages\": [{\"role\": \"user\", \"content\": \"Say hello in one short sentence.\"}],\n    \"max_tokens\": 800\n  }'\n\n# {\"choices\":[{\"finish_reason\":\"stop\",\"index\":0,\"message\":{\"role\":\"assistant\",\"content\":\"Hello!\"}}], ...\n```\n\n### [Docker Agent](#docker-agent)\n\nThere are several ways to get **[Docker Agent](https://docker.github.io/docker-agent/)**. The simplest one is to have a recent version of **[Docker Desktop](https://docs.docker.com/desktop/)**, and in that case all you need to type (to check) is:\n\n```\ndocker agent version\n```\n\nIn that case, [Docker Agent](https://docker.github.io/docker-agent/) is a [Docker Desktop](https://docs.docker.com/desktop/) plugin.\n\nBut you can install [Docker Agent](https://docker.github.io/docker-agent/) in a **\"standalone\"** version (you don't need `docker` or Docker Desktop to make it work), on Mac or Linux:\n\n```\nbrew install docker-agent\n```\n\nYou can also download the latest release from [https://github.com/docker/docker-agent/releases](https://github.com/docker/docker-agent/releases), and you'll find a **Windows** version there too.\n\nAnd this time you'll use the `docker-agent` command instead of `docker agent`:\n\n```\ndocker-agent version\n```\n\nSo all that's left is to write the configuration of our agent.\n\n## [Creating a configuration for Docker Agent](#creating-a-configuration-for-docker-agent)\n\nIn a folder of your choice, create an `agent.yaml` file with the content below:\n\n```\nproviders:\n  llamacpp:\n    api_type: openai_chatcompletions\n    base_url: http://localhost:8080/v1\n    # If you work from a container\n    #base_url: http://host.docker.internal:8080/v1\n\nmodels:\n  mellum2:\n    provider: llamacpp\n    model: JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M:Q4_K_M\n    #max_tokens: 8192\n    temperature: 0.7\n    provider_opts:\n      context_size: 262144\n\nagents:\n  root:\n    model: mellum2\n    description: A helpful AI assistant running on a local llama.cpp server\n    instruction: |\n      You name is Bob 🤓, you are a knowledgeable code assistant that helps users with various tasks.\n      Be helpful, accurate, and concise in your responses.\n      You have access to the local filesystem and shell: use these tools\n    welcome_message: |\n      🤖 Local Assistant propulsed by **llama.cpp** 🦙\n      \n    toolsets:\n      - type: filesystem\n      - type: shell\n```\n\nSo we have defined:\n\n- An OpenAI API compatible \"LLM provider\": `llamacpp`\n- A model (LLM): `mellum2`\n- Then a main agent: `root` with its system instructions, its model and a set of tools (`toolsets` ) to interact with the host system (and this is where I have to tell you that it's better to run a code agent in a sandbox, a VM or a container, and to go have a look at**[Docker SBX](https://docs.docker.com/ai/sandboxes/)** , which offers a container inside a micro VM).\n\nAll that's left is to launch our new agent.\n\n## [Starting Docker Agent](#starting-docker-agent)\n\n```\ndocker-agent run agent.yaml\n```\n\nor `docker agent run agent.yaml` depending your installation\n\nYou'll land on this TUI:\n\nAnd you can start interacting with your new code agent:\n\nThat's all for today (feel free to comment or ask questions). In an upcoming blog post, we'll see how to use **[Docker Agent](https://docker.github.io/docker-agent/)** in ACP (Agent Client Protocol) mode with **[Zed Editor](https://zed.dev/)**.\n\nWritten by\n\nKeep reading\n\n### Zed Editor + Docker Agent + ACP: coding with a local agent plugged into llmman\n\nPlug Zed's agent panel into Docker Agent over ACP, then point it at llmman to chat with a fully local, OpenAI-compatible coding agent.\n\nOct 3, 2026\n\n### A mini code agent with Docker Agent + Docker Model Runner - Part 1\n\nBuild a mini local code agent with Docker Agent and Docker Model Runner: a single shell tool on the small Mellum2 model, the agent loop, and running it safely in an sbx sandbox.\n\nAug 1, 2026\n\n### Zed Editor + Docker Agent + ACP, but in a sandbox: running the agent with sbx\n\nRun Docker Agent inside a Docker sandbox (sbx) and connect Zed to it over ACP, keeping llmman local while isolating the agent from your machine.\n\nOct 4, 2026\nFrom other blogs", "url": "https://wpnews.pro/news/docker-agent-llama-cpp-a-local-code-agent-in-5-minutes", "canonical_source": "https://k33g.org/p/20260915-docker-agent-llamacpp", "published_at": "2026-09-15 00:00:00+00:00", "updated_at": "2026-10-07 18:16:52.263958+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "developer-tools", "ai-products"], "entities": ["Docker Agent", "llama.cpp", "Docker Desktop", "JetBrains", "Mellum2-12B-A2.5B-Instruct", "gemma-4-26B-A4B-it", "Hugging Face", "Ollama"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/docker-agent-llama-cpp-a-local-code-agent-in-5-minutes", "markdown": "https://wpnews.pro/news/docker-agent-llama-cpp-a-local-code-agent-in-5-minutes.md", "text": "https://wpnews.pro/news/docker-agent-llama-cpp-a-local-code-agent-in-5-minutes.txt", "jsonld": "https://wpnews.pro/news/docker-agent-llama-cpp-a-local-code-agent-in-5-minutes.jsonld"}}