cd /news/ai-infrastructure/the-universal-llm-server · home › topics › ai-infrastructure › article
[ARTICLE · art-149309] src=zml.ai ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

The Universal LLM Server

ZML released LLMD, a self-contained inference server for LLaMa, Gemma, Qwen and Mistral models that runs across NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI and Apple Metal. LLMD ships as Docker images (CUDA 1.7 GB, ROCm 3.9 GB, TPU 280 MB, oneAPI 350 MB) plus a 140 MB Homebrew install for Metal, and supports continuous batching, multi-device sharding, shared prompt-prefix reuse and a built-in /metrics endpoint. Native DFlash support ships for Gemma 4 series models, with Qwen support planned.

read1 min views4 publishedOct 11, 2026

Continuous batching

Serve concurrent requests efficiently without hand-building a scheduler around the model.

ZML/LLMD

A self-contained inference server for LLaMa, Gemma, Qwen and Mistral models, running across NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI and Apple Metal.

docker run -p 8000:8000 \
  --shm-size=256GB --gpus=all \
  -e HF_TOKEN -it zmlai/llmd:cuda \
  --model=hf://Qwen/Qwen3-8B

Serving

Run long-context workloads with the memory behavior expected from a production server.

Shard model execution across multiple devices while LLMD handles communication.

Reuse shared prompt prefixes across requests to reduce repeated work.

Expose modern assistant workflows without leaving the LLMD serving path.

                    Monitor the server directly through the built-in
                    `/metrics` endpoint.

Native DFlash support ships for Gemma 4 series models, with Qwen support planned.

Try it

docker run -p 8000:8000 --shm-size=256GB --gpus=all -e HF_TOKEN -it zmlai/llmd:cuda \
  --model=hf://Qwen/Qwen3.6-27B
docker run -p 8000:8000 --device=/dev/kfd --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:rocm \
  --model=hf://Qwen/Qwen3.6-27B
docker run -p 8000:8000 --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:oneapi \
  --model=hf://Qwen/Qwen3.6-27B
docker run --net=host --privileged -e HF_TOKEN -it zmlai/llmd:tpu \
  --model=hf://Qwen/Qwen3.6-27B
brew install zml/zml/llmd
llmd --model=hf://Qwen/Qwen3.6-27B

Storage

                LLMD uses ZML's VFS subsystem to load from Hugging Face, S3
                and GCS without a separate download step. Use
                `hf://`, `s3://` or
                `gs://` anywhere a model path is expected.

Platforms

Platform Image Size
CUDA zmlai/llmd:cuda 1.7 GB
ROCm zmlai/llmd:rocm 3.9 GB
TPU zmlai/llmd:tpu 280 MB
OneAPI zmlai/llmd:oneapi 350 MB
Metal brew install zml/zml/llmd 140 MB
── more in #ai-infrastructure 4 stories · sorted by recency
── more on @zml 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-universal-llm-se…] indexed:0 read:1min 2026-10-11 · —