Today I asked my LLM box what it was doing, and it told me it had a 27-billion-parameter model loaded. Total footprint 18.3 GB. Amount of that on the graphics card: 0.5 GB.
So the 27B was "running on the GPU" in the same sense that I am running a marathon when I walk to the shop. The other 17.8 GB sat in system RAM and did its arithmetic on the CPU, one patient token at a time.
That seemed like a good moment for a proper census: what is actually running, on what, and which AI jobs go to local hardware, a paid cloud API or a flat subscription. The short version: 41 Docker containers and 9 LXC containers across four boxes, and a single 6 GB card makes most of the routing decisions for me.
(Background, already written up: the small-company org chart, the six-job agent team and why the coding agent moved to OpenCode and OpenRouter.)
Measured on 24 September.
| Box | Hardware | What it runs |
|---|---|---|
| Proxmox node 1 | 4 cores, ~8 GB RAM | 5 LXC containers: three single-purpose service boxes, a Docker host (12 containers) and a Docker "proxy node" (4 containers) |
| Proxmox node 2 | 4 cores, 16 GB RAM | 4 LXC containers: the Hermes agent box (native, no Docker), Obico (4 containers), Vaultwarden (1), Langfuse (6) |
| ZimaBlade | 2 cores, 16 GB RAM, fanless | 14 Docker containers on ZimaOS |
| LLM box | i7-10875H (16 threads), 31 GB RAM, RTX 2060 with 6 GB VRAM | Ollama with 21 model tags, and ComfyUI sharing the same GPU. No containers. |
The arithmetic, since I got it wrong the first time: node 1 has 12 + 4 = 16 Docker containers, node 2 has 4 + 1 + 6 = 11, so 27 on the Proxmox side, plus 14 on the ZimaBlade. That's 41 Docker containers and 9 LXC containers, plus three native services: Ollama, ComfyUI and Hermes itself.
All of that sits on 26 CPU threads and 71 GB of RAM, if you add up four boxes that were never designed to be added up.
The node 1 Docker host is the toolbox: crawl4ai for page fetching, SearXNG for search, ESPHome, Wyoming Whisper and Piper for speech-to-text and text-to-speech, a small web app plus its public demo, and Honcho, which is four containers on its own (API, database, deriver, Redis). The proxy node carries Home Assistant, Mosquitto, and the NetBird server and dashboard.
On node 2, Langfuse, which traces the agent's model calls, is six containers (ClickHouse, web, worker, MinIO, Postgres, Redis). That's more containers than the thing it observes.
The ZimaBlade is the photo library and the networking plumbing: Immich (server, machine learning, Postgres, Redis), the tunnel connector and reverse proxy, a Homepage dashboard, and omp, a terminal coding agent that points back at the LLM box, plus a few small utilities.
Node 2's guests all live on a single 5400rpm hard drive. The agent moved there for the extra RAM, and it was the right call, but every reboot is a queue of cold-starting databases fighting over one spindle. The first time it happened the agent's dashboard answered 502 for about 15 minutes. The fix was stopping Langfuse, the heaviest starter and the one thing nobody misses for ten minutes. The dashboard came up 21 seconds later.
The move itself left a twin behind. The original agent box on node 1 was never stopped: same MAC, same IP, same VPN identity. What I called "flaky networking" for days was two containers taking turns answering for one address. A census would have caught it on day one.
Here's the rule the whole routing falls out of: the agent framework I run refuses any model with less than a 64K context window, and the LLM box has 6 GB of VRAM.
Those two facts don't get along. The best local model I've measured for the job is qwen3.5 9B, built as a 64K variant. At 64K context it's 7.66 GB in total, of which 3.93 GB lands on the card. That's 51% on the GPU and the rest on the CPU, and no amount of context tuning fixes it: at 8K context it's still only 64% on the GPU, because Ollama pins roughly the same slice of VRAM and spills the remainder.
I spent much of August trying to shop my way out of that. Every candidate was measured and rejected:
qwen2.5 buildsqwen2.5-coder
So the local model is genuinely the best option on this hardware, and a real agent turn on it still took around a minute and a half when I timed it in August. The same work on a cheap cloud model answers a tool-call probe in 1.5 to 2 seconds.
That gap decides most of the routing.
| Job | Runs on | Why |
|---|---|---|
| Coding agent | OpenCode Go subscription ( deepseek-v4-flash ) |
Flat $10 a month. Usage caps, not a meter |
| Chief of staff, morning brief, research, writing, travel | OpenRouter qwen/qwen3.7-flash |
Pay per token, about $1.39 a month at measured volume |
| Security audit and code review | Local Ollama only | Deliberately no third-party egress |
| Embeddings, session titles, compression | Local Ollama | The framework's auxiliary calls stay local even when the main model is remote |
| Honcho user modelling | Local Ollama | Background work where latency doesn't matter |
| Images | ComfyUI (SD1.5) on the same card | Free, and small enough to share |
| Speech | Whisper and Piper containers, CPU | Never touches the GPU at all |
| Last resort | An OpenRouter free model | Rate-limited, so it only ever goes last |
The cloud model became the primary for the general profiles in late August. The number that justified it came from measuring 30 days of real usage: 44.35M input tokens against 421K output. Input outweighs output by about 105 to 1, which means input price sets the bill, and headline "per million" pricing is close to useless for ranking models. On a model at $0.03 per million input tokens, that volume comes to roughly $1.39 a month.
The chain is vendor-diverse on purpose: qwen/qwen3.7-flash, then deepseek/deepseek-v4-flash-0731 as a second paid model from a different vendor, then the local 64K model, then a free model last. A fallback that shares its primary's failure mode isn't a fallback.
A watchdog probes the primary every three hours for two failure modes: the model dying, and the money running out. A healthy model you can't pay for is still an outage, so below a credit floor it moves everything back to local and says so.
It also got something badly wrong. On 7 September every paid model "failed" its health check at once, so the watchdog moved four agents (chief of staff, research, writing and travel) back to the local model. Nothing had died: an upstream PII-redaction filter had turned the city in the probe prompt into [ADDRESS], so no model could answer it properly. Worse, the watchdog only restored automatically after a credit revert, not a failure revert, so those four agents stayed on the slow local model for 17 days until I caught it on 24 September. It now re-probes after any failover and moves back when a model passes. A probe that fails the same way on every vendor is a problem with the probe, not the models.
The coding agent is the heaviest user by a distance, so it's on OpenCode's Go plan instead: $10 a month, with rolling caps of $12 per five hours, $30 a week and $60 a month. Every response reports a cost of zero. On a subscription the risk isn't a bill, it's hitting a cap halfway through a task, so the watchdog reads Go's usage endpoint and tells a cap from a death. A capped profile gets parked on OpenRouter and put back automatically when the window resets.
One trap worth passing on: Go lives at /zen/go/v1, not /zen/v1. The second is pay-as-you-go and returns "insufficient balance" on a perfectly good Go key, which cost me an hour and a wrong conclusion.
Local isn't the cheap option here, it's the private one. The security and review agents see the most sensitive estate detail and never leave the LAN. Everything else local is background work, or the fallback for when the internet or a credit balance lets me down.
Running it taught me more than the cloud side did:
Count first. Take the census from the hosts, not from memory. Mine turned up a twin container and a tracing stack bigger than the thing it traces.
Let VRAM make the first cut. Write down the smallest context your agent needs, then check what fraction of a model at that context actually lands on your card. If the answer is half, you have your routing decision.
Rank cloud models on input price. Agent workloads are overwhelmingly prompt. Output price is mostly a distraction.
Keep local for the jobs that must stay local, and as the fallback. Don't make it carry interactive work it can't do quickly.
Assume silence is a failure. A pinned model, an empty ledger, a probe that swallows exceptions: none of them raise an error. Build the check that notices nothing happening.
A 27B model on 0.5 GB of VRAM is a fun number. It's also a fair summary of the whole exercise: the hardware will happily let you do something slow and call it running. The census is how you find out.
🤖 Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.