Docker Agent + llama.cpp: a local code agent in 5 minutes 🤖💙🦙 Docker Agent can be connected to a local llama.cpp server to run a code agent entirely on-device, with llama.cpp installed via a single curl command and started in serve mode on http://localhost:8080. The walkthrough serves the JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M model through llama.cpp's OpenAI-compatible /v1/chat/completions endpoint and configures Docker Agent via an agent.yaml file using the openai_ch API type, with Docker Agent available as a Docker Desktop plugin or a standalone brew install. Docker Agent + llama.cpp: a local code agent in 5 minutes 🤖💙🦙 My favourite code agent is still Docker Agent https://docker.github.io/docker-agent/ , especially when I work with "local" LLMs, be they "big" gemma-4-26B-A4B-it https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF when I'm on my work laptop, or more modest Mellum2-12B-A2.5B-Instruct https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4 K M when I'm on my personal Mac Book Air. A big advantage of Docker Agent https://docker.github.io/docker-agent/ is being able to connect to various model providers, remote ones like Anthropic, OpenAI, OVH's AI endpoints https://blog.ovhcloud.com/en/posts/discovering-docker-agent-with-ai-endpoints/ , MistralAI https://mistral.ai/ , ... or local ones like Ollama https://ollama.com/ , Docker Model Runner https://docs.docker.com/ai/model-runner/ , llama.cpp https://llama.app/ , ... Today, the one I'm interested in is llama.cpp https://llama.app/ , because when a new model in GGUF format shows up, llama.cpp https://llama.app/ is generally the first to be updated to handle the model's specifics. Kronk https://www.kronkai.com/ is very reactive too, I had written an article https://k33g.org/p/20260510-kronk-docker-agent-sbx about it that I'll have to refresh But let's get back to llama.cpp https://llama.app/ , which we'll have to install and start. Prerequisites prerequisites Llama.cpp llamacpp Installing it is very simple, a single curl command is enough otherwise there are other options, I'll let you refer to the llama.cpp website . curl -LsSf https://llama.app/install.sh | sh Then all you have to do is start llama in serve mode so, in API mode so that it can be used by a code agent: llama serve -hf JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4 K M:Q4 K M If the model isn't present on your machine, llama.cpp https://llama.app/ will download it from Hugging Face https://huggingface.co/ , hence the -hf flag : And then the model is served on http://localhost:8080 http://localhost:8080 You can run a few checks to verify: Server health curl http://localhost:8080/health {"status":"ok"} Exact name of the exposed model curl http://localhost:8080/v1/models First chat completion curl http://localhost:8080/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4 K M:Q4 K M", "messages": {"role": "user", "content": "Say hello in one short sentence."} , "max tokens": 800 }' {"choices": {"finish reason":"stop","index":0,"message":{"role":"assistant","content":"Hello "}} , ... Docker Agent docker-agent There are several ways to get Docker Agent https://docker.github.io/docker-agent/ . The simplest one is to have a recent version of Docker Desktop https://docs.docker.com/desktop/ , and in that case all you need to type to check is: docker agent version In that case, Docker Agent https://docker.github.io/docker-agent/ is a Docker Desktop https://docs.docker.com/desktop/ plugin. But you can install Docker Agent https://docker.github.io/docker-agent/ in a "standalone" version you don't need docker or Docker Desktop to make it work , on Mac or Linux: brew install docker-agent You can also download the latest release from https://github.com/docker/docker-agent/releases https://github.com/docker/docker-agent/releases , and you'll find a Windows version there too. And this time you'll use the docker-agent command instead of docker agent : docker-agent version So all that's left is to write the configuration of our agent. Creating a configuration for Docker Agent creating-a-configuration-for-docker-agent In a folder of your choice, create an agent.yaml file with the content below: providers: llamacpp: api type: openai chatcompletions base url: http://localhost:8080/v1 If you work from a container base url: http://host.docker.internal:8080/v1 models: mellum2: provider: llamacpp model: JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4 K M:Q4 K M max tokens: 8192 temperature: 0.7 provider opts: context size: 262144 agents: root: model: mellum2 description: A helpful AI assistant running on a local llama.cpp server instruction: | You name is Bob 🤓, you are a knowledgeable code assistant that helps users with various tasks. Be helpful, accurate, and concise in your responses. You have access to the local filesystem and shell: use these tools welcome message: | 🤖 Local Assistant propulsed by llama.cpp 🦙 toolsets: - type: filesystem - type: shell So we have defined: - An OpenAI API compatible "LLM provider": llamacpp - A model LLM : mellum2 - Then a main agent: root with its system instructions, its model and a set of tools toolsets to interact with the host system and this is where I have to tell you that it's better to run a code agent in a sandbox, a VM or a container, and to go have a look at Docker SBX https://docs.docker.com/ai/sandboxes/ , which offers a container inside a micro VM . All that's left is to launch our new agent. Starting Docker Agent starting-docker-agent docker-agent run agent.yaml or docker agent run agent.yaml depending your installation You'll land on this TUI: And you can start interacting with your new code agent: That's all for today feel free to comment or ask questions . In an upcoming blog post, we'll see how to use Docker Agent https://docker.github.io/docker-agent/ in ACP Agent Client Protocol mode with Zed Editor https://zed.dev/ . Written by Keep reading Zed Editor + Docker Agent + ACP: coding with a local agent plugged into llmman Plug Zed's agent panel into Docker Agent over ACP, then point it at llmman to chat with a fully local, OpenAI-compatible coding agent. Oct 3, 2026 A mini code agent with Docker Agent + Docker Model Runner - Part 1 Build a mini local code agent with Docker Agent and Docker Model Runner: a single shell tool on the small Mellum2 model, the agent loop, and running it safely in an sbx sandbox. Aug 1, 2026 Zed Editor + Docker Agent + ACP, but in a sandbox: running the agent with sbx Run Docker Agent inside a Docker sandbox sbx and connect Zed to it over ACP, keeping llmman local while isolating the agent from your machine. Oct 4, 2026 From other blogs