{"slug": "the-universal-llm-server", "title": "The Universal LLM Server", "summary": "ZML released LLMD, a self-contained inference server for LLaMa, Gemma, Qwen and Mistral models that runs across NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI and Apple Metal. LLMD ships as Docker images (CUDA 1.7 GB, ROCm 3.9 GB, TPU 280 MB, oneAPI 350 MB) plus a 140 MB Homebrew install for Metal, and supports continuous batching, multi-device sharding, shared prompt-prefix reuse and a built-in /metrics endpoint. Native DFlash support ships for Gemma 4 series models, with Qwen support planned.", "body_md": "### Continuous batching\n\nServe concurrent requests efficiently without hand-building a scheduler around the model.\n\nZML/LLMD\n\nA self-contained inference server for LLaMa, Gemma, Qwen and Mistral models, running across NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI and Apple Metal.\n\n```\ndocker run -p 8000:8000 \\\n  --shm-size=256GB --gpus=all \\\n  -e HF_TOKEN -it zmlai/llmd:cuda \\\n  --model=hf://Qwen/Qwen3-8B\n```\n\nServing\n\nRun long-context workloads with the memory behavior expected from a production server.\n\nShard model execution across multiple devices while LLMD handles communication.\n\nReuse shared prompt prefixes across requests to reduce repeated work.\n\nExpose modern assistant workflows without leaving the LLMD serving path.\n\n                        Monitor the server directly through the built-in\n                        `/metrics` endpoint.\n                    \n\nNative DFlash support ships for Gemma 4 series models, with Qwen support planned.\n\nTry it\n\n```\ndocker run -p 8000:8000 --shm-size=256GB --gpus=all -e HF_TOKEN -it zmlai/llmd:cuda \\\n  --model=hf://Qwen/Qwen3.6-27B\ndocker run -p 8000:8000 --device=/dev/kfd --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:rocm \\\n  --model=hf://Qwen/Qwen3.6-27B\ndocker run -p 8000:8000 --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:oneapi \\\n  --model=hf://Qwen/Qwen3.6-27B\ndocker run --net=host --privileged -e HF_TOKEN -it zmlai/llmd:tpu \\\n  --model=hf://Qwen/Qwen3.6-27B\nbrew install zml/zml/llmd\nllmd --model=hf://Qwen/Qwen3.6-27B\n```\n\nStorage\n\n                    LLMD uses ZML's VFS subsystem to load from Hugging Face, S3\n                    and GCS without a separate download step. Use\n                    `hf://`, `s3://` or\n                    `gs://` anywhere a model path is expected.\n                \n\nPlatforms\n\n| Platform | Image | Size | \n|---|---|---|\n| CUDA | `zmlai/llmd:cuda` | 1.7 GB | \n| ROCm | `zmlai/llmd:rocm` | 3.9 GB | \n| TPU | `zmlai/llmd:tpu` | 280 MB | \n| OneAPI | `zmlai/llmd:oneapi` | 350 MB | \n| Metal | `brew install zml/zml/llmd` | 140 MB |", "url": "https://wpnews.pro/news/the-universal-llm-server", "canonical_source": "https://zml.ai/llmd/", "published_at": "2026-10-11 21:03:35+00:00", "updated_at": "2026-10-11 21:34:16.338028+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-tools", "mlops", "developer-tools"], "entities": ["ZML", "LLMD", "LLaMa", "Gemma", "Qwen", "Mistral", "NVIDIA CUDA", "AMD ROCm"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-universal-llm-server", "markdown": "https://wpnews.pro/news/the-universal-llm-server.md", "text": "https://wpnews.pro/news/the-universal-llm-server.txt", "jsonld": "https://wpnews.pro/news/the-universal-llm-server.jsonld"}}