# Docker Agent + llama.cpp: a local code agent in 5 minutes 🤖💙🦙

> Source: <https://k33g.org/p/20260915-docker-agent-llamacpp>
> Published: 2026-09-15 00:00:00+00:00

# Docker Agent + llama.cpp: a local code agent in 5 minutes 🤖💙🦙

My favourite code agent is still **[Docker Agent](https://docker.github.io/docker-agent/)**, especially when I work with "local" LLMs, be they "big" ([gemma-4-26B-A4B-it](https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF)) when I'm on my work laptop, or more modest ([Mellum2-12B-A2.5B-Instruct](https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M)) when I'm on my personal Mac Book Air.

A big advantage of **[Docker Agent](https://docker.github.io/docker-agent/)** is being able to connect to various model providers, **remote** ones (like Anthropic, OpenAI, [OVH's AI endpoints](https://blog.ovhcloud.com/en/posts/discovering-docker-agent-with-ai-endpoints/), [MistralAI](https://mistral.ai/), ...) or **local** ones (like [Ollama](https://ollama.com/), [Docker Model Runner](https://docs.docker.com/ai/model-runner/), [llama.cpp](https://llama.app/), ...)

Today, the one I'm interested in is **[llama.cpp](https://llama.app/)**, because when a new model in GGUF format shows up, **[llama.cpp](https://llama.app/)** is generally the first to be updated to handle the model's specifics.

[Kronk](https://www.kronkai.com/) is very reactive too, I had written an [article](https://k33g.org/p/20260510-kronk-docker-agent-sbx) about it that I'll have to refresh

But let's get back to **[llama.cpp](https://llama.app/)**, which we'll have to install and start.

## [Prerequisites](#prerequisites)

### [Llama.cpp](#llamacpp)

Installing it is very simple, a single `curl` command is enough (otherwise there are other options, I'll let you refer to the llama.cpp website).

```
curl -LsSf https://llama.app/install.sh | sh
```

Then all you have to do is start `llama` in `serve` mode (so, in API mode) so that it can be used by a code agent:

```
llama serve -hf JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M:Q4_K_M
```

If the model isn't present on your machine, **[llama.cpp](https://llama.app/)** will download it (from [Hugging Face](https://huggingface.co/), hence the `-hf` flag):

And then the model is served on [http://localhost:8080](http://localhost:8080)

You can run a few checks to verify:

```
# Server health
curl http://localhost:8080/health
# {"status":"ok"}

# Exact name of the exposed model
curl http://localhost:8080/v1/models

# First chat completion
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M:Q4_K_M",
    "messages": [{"role": "user", "content": "Say hello in one short sentence."}],
    "max_tokens": 800
  }'

# {"choices":[{"finish_reason":"stop","index":0,"message":{"role":"assistant","content":"Hello!"}}], ...
```

### [Docker Agent](#docker-agent)

There are several ways to get **[Docker Agent](https://docker.github.io/docker-agent/)**. The simplest one is to have a recent version of **[Docker Desktop](https://docs.docker.com/desktop/)**, and in that case all you need to type (to check) is:

```
docker agent version
```

In that case, [Docker Agent](https://docker.github.io/docker-agent/) is a [Docker Desktop](https://docs.docker.com/desktop/) plugin.

But you can install [Docker Agent](https://docker.github.io/docker-agent/) in a **"standalone"** version (you don't need `docker` or Docker Desktop to make it work), on Mac or Linux:

```
brew install docker-agent
```

You can also download the latest release from [https://github.com/docker/docker-agent/releases](https://github.com/docker/docker-agent/releases), and you'll find a **Windows** version there too.

And this time you'll use the `docker-agent` command instead of `docker agent`:

```
docker-agent version
```

So all that's left is to write the configuration of our agent.

## [Creating a configuration for Docker Agent](#creating-a-configuration-for-docker-agent)

In a folder of your choice, create an `agent.yaml` file with the content below:

```
providers:
  llamacpp:
    api_type: openai_chatcompletions
    base_url: http://localhost:8080/v1
    # If you work from a container
    #base_url: http://host.docker.internal:8080/v1

models:
  mellum2:
    provider: llamacpp
    model: JetBrains/Mellum2-12B-A2.5B-Instruct-GGUF-Q4_K_M:Q4_K_M
    #max_tokens: 8192
    temperature: 0.7
    provider_opts:
      context_size: 262144

agents:
  root:
    model: mellum2
    description: A helpful AI assistant running on a local llama.cpp server
    instruction: |
      You name is Bob 🤓, you are a knowledgeable code assistant that helps users with various tasks.
      Be helpful, accurate, and concise in your responses.
      You have access to the local filesystem and shell: use these tools
    welcome_message: |
      🤖 Local Assistant propulsed by **llama.cpp** 🦙
      
    toolsets:
      - type: filesystem
      - type: shell
```

So we have defined:

- An OpenAI API compatible "LLM provider": `llamacpp`
- A model (LLM): `mellum2`
- Then a main agent: `root` with its system instructions, its model and a set of tools (`toolsets` ) to interact with the host system (and this is where I have to tell you that it's better to run a code agent in a sandbox, a VM or a container, and to go have a look at**[Docker SBX](https://docs.docker.com/ai/sandboxes/)** , which offers a container inside a micro VM).

All that's left is to launch our new agent.

## [Starting Docker Agent](#starting-docker-agent)

```
docker-agent run agent.yaml
```

or `docker agent run agent.yaml` depending your installation

You'll land on this TUI:

And you can start interacting with your new code agent:

That's all for today (feel free to comment or ask questions). In an upcoming blog post, we'll see how to use **[Docker Agent](https://docker.github.io/docker-agent/)** in ACP (Agent Client Protocol) mode with **[Zed Editor](https://zed.dev/)**.

Written by

Keep reading

### Zed Editor + Docker Agent + ACP: coding with a local agent plugged into llmman

Plug Zed's agent panel into Docker Agent over ACP, then point it at llmman to chat with a fully local, OpenAI-compatible coding agent.

Oct 3, 2026

### A mini code agent with Docker Agent + Docker Model Runner - Part 1

Build a mini local code agent with Docker Agent and Docker Model Runner: a single shell tool on the small Mellum2 model, the agent loop, and running it safely in an sbx sandbox.

Aug 1, 2026

### Zed Editor + Docker Agent + ACP, but in a sandbox: running the agent with sbx

Run Docker Agent inside a Docker sandbox (sbx) and connect Zed to it over ACP, keeping llmman local while isolating the agent from your machine.

Oct 4, 2026
From other blogs
