cd /news/ai-agents/docker-agent-llama-cpp-a-local-code-… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-147010] src=k33g.org β†— pub= topic=ai-agents verified=true sentiment=↑ positive

Docker Agent + llama.cpp: a local code agent in 5 minutes πŸ€–πŸ’™πŸ¦™

Docker Agent can be connected to a local llama.cpp server to run a code agent entirely on-device, with llama.cpp installed via a single curl command and started in serve mode on http://localhost:8080. The walkthrough serves the JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M model through llama.cpp's OpenAI-compatible /v1/chat/completions endpoint and configures Docker Agent via an agent.yaml file using the openai_ch API type, with Docker Agent available as a Docker Desktop plugin or a standalone brew install.

read4 min views5 publishedSep 15, 2026
Docker Agent + llama.cpp: a local code agent in 5 minutes πŸ€–πŸ’™πŸ¦™
Image: K33G (auto-discovered)

My favourite code agent is still Docker Agent, especially when I work with "local" LLMs, be they "big" (gemma-4-26B-A4B-it) when I'm on my work laptop, or more modest (Mellum2-12B-A2.5B-Instruct) when I'm on my personal Mac Book Air.

A big advantage of Docker Agent is being able to connect to various model providers, remote ones (like Anthropic, OpenAI, OVH's AI endpoints, MistralAI, ...) or local ones (like Ollama, Docker Model Runner, llama.cpp, ...)

Today, the one I'm interested in is llama.cpp, because when a new model in GGUF format shows up, llama.cpp is generally the first to be updated to handle the model's specifics.

Kronk is very reactive too, I had written an article about it that I'll have to refresh

But let's get back to llama.cpp, which we'll have to install and start.

Prerequisites #

Llama.cpp

Installing it is very simple, a single curl command is enough (otherwise there are other options, I'll let you refer to the llama.cpp website).

curl -LsSf https://llama.app/install.sh | sh

Then all you have to do is start llama in serve mode (so, in API mode) so that it can be used by a code agent:

llama serve -hf JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M:Q4_K_M

If the model isn't present on your machine, llama.cpp will download it (from Hugging Face, hence the -hf flag):

And then the model is served on http://localhost:8080

You can run a few checks to verify:

curl http://localhost:8080/health

curl http://localhost:8080/v1/models

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M:Q4_K_M",
    "messages": [{"role": "user", "content": "Say hello in one short sentence."}],
    "max_tokens": 800
  }'

Docker Agent

There are several ways to get Docker Agent. The simplest one is to have a recent version of Docker Desktop, and in that case all you need to type (to check) is:

docker agent version

In that case, Docker Agent is a Docker Desktop plugin.

But you can install Docker Agent in a "standalone" version (you don't need docker or Docker Desktop to make it work), on Mac or Linux:

brew install docker-agent

You can also download the latest release from https://github.com/docker/docker-agent/releases, and you'll find a Windows version there too.

And this time you'll use the docker-agent command instead of docker agent:

docker-agent version

So all that's left is to write the configuration of our agent.

Creating a configuration for Docker Agent #

In a folder of your choice, create an agent.yaml file with the content below:

providers:
  llamacpp:
    api_type: openai_chatcompletions
    base_url: http://localhost:8080/v1
    #base_url: http://host.docker.internal:8080/v1

models:
  mellum2:
    provider: llamacpp
    model: JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M:Q4_K_M
    #max_tokens: 8192
    temperature: 0.7
    provider_opts:
      context_size: 262144

agents:
  root:
    model: mellum2
    description: A helpful AI assistant running on a local llama.cpp server
    instruction: |
      You name is Bob πŸ€“, you are a knowledgeable code assistant that helps users with various tasks.
      Be helpful, accurate, and concise in your responses.
      You have access to the local filesystem and shell: use these tools
    welcome_message: |
      πŸ€– Local Assistant propulsed by **llama.cpp** πŸ¦™
      
    toolsets:
      - type: filesystem
      - type: shell

So we have defined:

  • An OpenAI API compatible "LLM provider": llamacpp
  • A model (LLM): mellum2
  • Then a main agent: root with its system instructions, its model and a set of tools (toolsets ) to interact with the host system (and this is where I have to tell you that it's better to run a code agent in a sandbox, a VM or a container, and to go have a look at**Docker SBX** , which offers a container inside a micro VM).

All that's left is to launch our new agent.

Starting Docker Agent #

docker-agent run agent.yaml

or docker agent run agent.yaml depending your installation

You'll land on this TUI:

And you can start interacting with your new code agent:

That's all for today (feel free to comment or ask questions). In an upcoming blog post, we'll see how to use Docker Agent in ACP (Agent Client Protocol) mode with Zed Editor.

Written by

Keep reading

Zed Editor + Docker Agent + ACP: coding with a local agent plugged into llmman

Plug Zed's agent panel into Docker Agent over ACP, then point it at llmman to chat with a fully local, OpenAI-compatible coding agent.

Oct 3, 2026

A mini code agent with Docker Agent + Docker Model Runner - Part 1

Build a mini local code agent with Docker Agent and Docker Model Runner: a single shell tool on the small Mellum2 model, the agent loop, and running it safely in an sbx sandbox.

Aug 1, 2026

Zed Editor + Docker Agent + ACP, but in a sandbox: running the agent with sbx

Run Docker Agent inside a Docker sandbox (sbx) and connect Zed to it over ACP, keeping llmman local while isolating the agent from your machine.

Oct 4, 2026 From other blogs

── more in #ai-agents 4 stories Β· sorted by recency
── more on @docker agent 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/docker-agent-llama-c…] indexed:0 read:4min 2026-09-15 Β· β€”